What are the responsibilities and job description for the Senior Site Reliability Engineer (SRE) / Senior Software Engineer position at Cube Hub Inc.?
Job Title: Senior Site Reliability Engineer (Incident Lead & Cloud Resiliency)
Location: Urbandale, IA 50322 (100% Onsite)
Contract Duration: 12 Months Contract with possibility of Extension
Shift Structure: First Shift | 1st Shift Open Schedule, Monday through Friday
Pay Rate: $75.00 - $77.00/hr on W2
Assessment Requirement: A live technical screening will occur during the interview process.
Job Description: Our client, a global leader in precision machinery and advanced agricultural technology, is seeking a high-caliber Senior Site Reliability Engineer (SRE) to join their Digital Operations team in Urbandale, IA. In this senior technical leadership role, you will act as a primary Incident Commander and technical advisor—designing creative automation solutions for application resiliency, system availability, and cloud security across complex enterprise ecosystems. You will define Service Level Objectives (SLOs), lead failure recovery frameworks, conduct post-mortems, and eliminate single points of failure before they impact production.
Key Responsibilities:
- Incident Command & Recovery: Lead as Incident Commander during complex system outages—coordinating real-time recovery, analyzing failure modes, and automating self-healing capabilities.
- System Resiliency & Reliability: Design, build, and optimize division-wide solutions for application availability, scalability, performance monitoring, and disaster recovery.
- SLO & Observability Leadership: Establish and maintain Service Level Objectives (SLOs) across sub-systems while implementing metric-gathering pipelines in Datadog and cloud observability suites.
- Post-Mortem & Automation: Lead root-cause post-mortems, document operational playbooks, and write Infrastructure as Code (Terraform) scripts to prevent recurring failure modes.
- Cross-Functional Mentorship: Collaborate with product teams and junior SREs to enforce site reliability best practices, agile software delivery, and security standards.
Required Qualifications & Skills:
- Experience: 4 to 7 years of dedicated experience in Site Reliability Engineering (SRE), Cloud Infrastructure, or System Architecture in enterprise environments.
- Incident Response Leadership: Proven track record acting as an Incident Commander or senior escalation point for large-scale application failures.
- Cloud & Orchestration Infrastructure: Strong, hands-on operational experience with AWS Cloud Services and Kubernetes / EKS cluster management.
- Infrastructure as Code & Observability: Proficiency with Terraform alongside deep monitoring/tracing experience using Datadog.
- Polyglot / Multi-Language Familiarity: Broad exposure to enterprise programming languages and frameworks (including Java, Scala, Go, Python, JavaScript, or .NET).
- Communication: Exceptional verbal and written communication skills with the ability to articulate complex technical recovery plans to non-technical stakeholders under high pressure.
- Education: Bachelor’s degree in Computer Science, Information Technology, Software Engineering, or equivalent practical experience.
Stand-Out Skills (Preferred):
- Practical experience establishing SLOs/SLIs from scratch across multi-tenant cloud microservices.
- Active background working within a centralized Developer Experience, DevOps, or Platform Engineering organization.
Actual Skills Required:
- Site Reliability Engineering (SRE) & Incident Commander Leadership
- AWS Cloud Architecture & Kubernetes Cluster Management
- Datadog Observability & Metrics Automation
- Terraform Infrastructure as Code (IaC)
- Polyglot Development & Automated Root Cause Analysis (Java/Go/Python)
Salary : $75 - $77