What are the responsibilities and job description for the Lead Site Reliability Engineer position at Eleven Recruiting?
About Eleven Recruiting
We are a specialized technology staffing agency supporting professional and financial services companies. Why do we stand out in technology staffing? We listen and act as advisors for our candidates on how they can best add value, find interesting projects, and pave a path for career advancement. We advocate for the best pay, diversity in tech, and the best job fit for every candidate we place.
Our client, an investment firm, is seeking a Lead Site Reliability Engineer to join their team in New York, NY!
The ideal candidate will help drive one of the firm's most impactful engineering modernization initiatives. This individual will play a critical role in scaling and operationalizing the next-generation observability and platform engineering ecosystem while partnering directly with engineering teams across a global organization.
Responsibilities
Lead Enterprise Observability Transformation
- Help drive strategic migration from more than 10 legacy monitoring and observability platforms into a unified Datadog ecosystem.
- Build, deploy, and maintain observability solutions across cloud, infrastructure, and application environments.
- Partner with engineering teams to onboard services, create dashboards, establish alerting standards, and improve operational visibility.
- Act as a technical advisor and subject matter expert for observability best practices across the organization.
Build and Scale Infrastructure as Code
- Design and maintain infrastructure using Terraform as the primary Infrastructure-as-Code platform.
- Develop and manage scalable cloud infrastructure, monitoring configurations, alerts, and platform services through code.
- Contribute heavily to Git-based workflows, pull requests, CI/CD pipelines, and automation initiatives.
- Leverage AI-assisted development tools responsibly while maintaining high standards for quality, reliability, and security.
Support Kubernetes and Platform Engineering Initiatives
- Work closely with platform engineering teams supporting Kubernetes-based environments.
- Assist engineering teams with deployments, troubleshooting, performance optimization, and operational excellence.
- Help drive adoption of modern platform tooling and cloud-native practices across the organization.
Enable Engineering Teams
- Serve as a technical resource for hundreds of engineers adopting new platform capabilities.
- Guide teams through onboarding, implementation, troubleshooting, and operational best practices.
- Create documentation, standards, and repeatable processes that improve developer experience and platform adoption.
Drive Reliability at Scale
- Help establish the next generation of incident response, alerting, and operational processes.
- Participate in designing future PagerDuty integrations and follow-the-sun support models across a globally distributed team.
- Continuously improve system reliability, observability, and operational maturity.
Required Qualifications
- 8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles.
- Strong hands-on expertise with Terraform and Infrastructure as Code.
- Experience implementing or managing enterprise observability solutions such as Datadog, New Relic, Dynatrace, Splunk Observability, OpenTelemetry, or similar platforms.
- Deep understanding of monitoring, alerting, logging, tracing, and observability concepts.
- Strong Kubernetes experience, including deployments, troubleshooting, and production operations.
- Experience with CI/CD pipelines and Git-based development workflows.
- Experience supporting large-scale cloud environments.
- Strong scripting and automation capabilities.
Preferred Qualifications
- Experience leading large-scale platform migrations or modernization initiatives.
- Experience supporting engineering organizations with hundreds of developers.
- Exposure to Azure environments and virtualized infrastructure.
- Experience implementing operational standards, incident management processes, and reliability practices.
- Familiarity with AI-assisted software development tools and engineering productivity platforms.
Who Will Thrive Here
- Highly technical and hands-on builders.
- Comfortable operating in environments with significant change and evolving priorities.
- Strong problem solvers who can navigate complex organizations and work across multiple teams. Effective communicators capable of influencing engineers, leadership, and stakeholders.
- Energized by seeing the direct impact of their work across a large engineering organization.
- Self-directed individuals who can balance strategic initiatives with day-to-day execution.
Pay Rate: $90 - $120/hr
Salary : $90 - $120