Recent Searches

You haven't searched anything yet.

1 Site Reliability Engineer, Observability Job in Roseland, NJ

SET JOB ALERT
Details...
CoreWeave
Roseland, NJ | Full Time
$102k-120k (estimate)
2 Months Ago
Site Reliability Engineer, Observability
CoreWeave Roseland, NJ
$102k-120k (estimate)
Full Time 2 Months Ago
Save

CoreWeave is Hiring a Remote Site Reliability Engineer, Observability

About the role:

The Observability Team performs a critical role in enabling CoreWeave to understand, troubleshoot, and optimize complex systems by providing comprehensive insights into their behavior and performance. This team is responsible for the development, integration, and operation of observability platforms with the ultimate objective of enabling engineers across CoreWeave to do more, better. Central to the Observability Teams mission is the operation of our observability stack which leverages CoreWeave’s deep investment in the Kubernetes ecosystem.

We are seeking a Site Reliability Engineer with specialization in the observability stack who can help us execute on the mission of providing a comprehensive logging and metrics ecosystem. Integrating logging, metrics, tracing, and monitoring tools for proactive insights into system performance. This individual will work with a team of 6-8 engineers and have the opportunity to work on the full gamut of rewarding challenges that come with the business of building a cloud in a communicative, supportive, and high-performing environment. As a member of the Observability Team you will have the opportunity to:

  • Design and implement the platform that improves visibility into how the services are performing and operating.
  • Improve the performance, security, reliability, and scalability of our observability, and related services and participate in the teams on-call rotation.
  • Assist engineers in maximizing the observability stack to gain insights into the service's functionality and operation.
  • Develop dashboards, alerts, and insights into the customer experience using Grafana-ecosystem tools such as Mimir and Loki.
  • Develop meaningful insights by analyzing the gathered data.
  • Enable and evangelize the best practices around alerting. Collaborate with teams to establish observability standards.
  • Grow, change, invest in your teammates, be invested-in, share your ideas, listen to others, be curious, have fun, and, above all, be yourself.

Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are some qualities we’ve found compatible with our team. If a portion of this resonates with you, we’d love to talk. 

  • You have one or more years of experience in a software or infrastructure engineering industry.
  • You enjoy helping your colleagues achieve more with less effort.
  • You have experience operating services in production and at scale and are versed in reliability engineering concepts such as the different types of testing, progressive deployments, error budgets, the role observability, and fault-tolerant design.
  • You’re familiar with various logging and metrics systems like ELK, Victoria Metrics, Thanos or Grafana. 
  • You enjoy understanding the data model for observability systems. 
  • You’re familiar with Kubernetes and have interest or experience with using it for event-driven and/or stateful orchestration.
  • You’re comfortable with the idea of using Go as your primary programming language.
  • You know your way around a Linux distro, shell scripting, and/or the Linux storage and networking stacks.
  • You’re excited about being part of a team of diverse perspectives and backgrounds that believe in tackling challenges, growing hand in hand, and winning together.

Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $160,000/year in our lowest geographic market up to $185,000/year in our highest geographic market. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience. 

Hybrid Workplace

Successful candidates will be expected to attend onboarding training at our NJ Headquarters within their first several weeks of employment, with subsequent quarterly travel requirements of 1 week duration.

If you reside within a 30-mile radius of our New Jersey, New York, or Philadelphia offices, we're excited for you to join us at the office at least three times a week, recognizing the significance we place on fostering connections, collaboration, and creativity within our office culture. Our commitment to operating as a hybrid workplace underscores our dedication to enabling our employees to tailor their work-life balance to their individual preferences

Job Summary

JOB TYPE

Full Time

SALARY

$102k-120k (estimate)

POST DATE

04/12/2024

EXPIRATION DATE

07/10/2024

WEBSITE

coreweave.com

HEADQUARTERS

New York, NY

Show more

CoreWeave
Remote | Full Time
$130k-165k (estimate)
1 Day Ago
CoreWeave
Remote | Full Time
$90k-112k (estimate)
1 Day Ago
CoreWeave
Remote | Full Time
$148k-188k (estimate)
1 Day Ago