What are the responsibilities and job description for the Lead Site Reliability Engineer -- FTE-- Onsite at FL position at REDLEO SOFTWARE INC.?
Lead Site Reliability Engineer – Infrastructure & DevOps
Location: Onsite – Orlando, FL (Open to Glendale/Anaheim, Seattle, Orlando)
Category: DevOps / SRE
CLIENT NOTE:
Candidates should have some exposure to AI systems and infrastructure. They should have experience supporting the infrastructure behind AI systems or implementing visibility, observability, and monitoring to assess AI system health and performance.
MUST HAVE:
- 7 years SRE / Platform / DevOps experience
- Expert Kubernetes, Terraform, Helm
- Multi‑cloud production experience: AWS, Google Cloud Platform, Azure
- Proven technical leadership in high‑availability environments
- Strong CI/CD automation (Harness or equivalent)
- Deep observability experience (monitoring, logging, alerting)
- Experience supporting mission‑critical, highly available production systems
Required Skills
- AWS, Google Cloud Platform cloud networking
- Docker, Kubernetes
- Terraform, Helm
- CI/CD pipelines (Harness preferred)
- Multi‑cloud operations
- Infrastructure automation & reliability engineering
Role Overview
- Lead SRE initiatives across multi‑cloud environments
- Architect, automate, and optimize infrastructure for reliability and scalability
- Build and maintain CI/CD pipelines and deployment automation
- Enhance observability, monitoring, and alerting systems
- Support production workloads, troubleshoot incidents, and ensure uptime
- Collaborate with engineering, platform, and security teams
- Drive best practices for infrastructure, DevOps, and cloud operations
Salary : $110,000 - $120,000