What are the responsibilities and job description for the Site Reliability Engineer lead position at Stellerx IT Solutions?
ROLE OVERVIEW
We are seeking a seasoned Site Reliability Engineering Lead to take end-to-end ownership of our retail eCommerce platform's operational excellence. You will drive SRE culture across engineering teams, champion observability, and lead a cross-functional onsite/offshore team to deliver exceptional uptime, performance, and release velocity. This role blends deep technical expertise with strategic leadership — you'll be equally comfortable reviewing Kubernetes manifests, tuning Akamai rules, and presenting SLO dashboards to senior leadership.
KEY RESPONSIBILITIES
Platform Reliability & Availability
▸ Define, monitor, and enforce SLIs, SLOs, and Error Budgets for all eCommerce services
▸ Lead incident response (ICS model): detection, triage, remediation, post-mortem, and blameless RCA
▸ Drive chaos engineering and game-day exercises to validate fault tolerance
▸ Architect resilient multi-AZ, multi-region AWS deployments for zero-downtime operations
CI/CD & Automation
▸ Own and evolve Jenkins-based CI/CD pipelines for microservices deployments on Kubernetes/EKS
▸ Implement GitOps workflows; enforce trunk-based development and deployment gates
▸ Automate infrastructure provisioning using CloudFormation, Terraform, and AWS CDK
▸ Drive shift-left reliability practices: load testing, chaos gates, and SLO checks in pipeline
Cloud Infrastructure & CDN
▸ Manage Akamai CDN configuration: edge rules, caching policies, TLS, WAF, and traffic shaping
▸ Optimize AWS Auto Scaling Groups, ALBs, and CloudFront for peak eCommerce traffic events
▸ Design and govern caching strategy across CDN, Redis, and Memcached layers
▸ Manage EKS/K8s cluster operations: node pools, Helm releases, resource quotas, and HPA/VPA
Observability & Performance
▸ Build and maintain Splunk APM and Splunk Cloud dashboards, alerts, and runbooks
▸ Implement distributed tracing (OpenTelemetry) across all microservices
▸ Conduct performance profiling and capacity planning aligned to traffic growth projections
▸ Establish DORA metrics tracking: deployment frequency, lead time, MTTR, change failure rate
Security, Compliance & Collaboration
▸ Partner with InfoSec on PCI-DSS compliance, vulnerability management, and secrets governance
▸ Enforce least-privilege IAM, network segmentation, and container security scanning in pipelines
▸ Lead cross-functional reviews with Dev, QA, and InfoSec for new service onboarding
▸ Coordinate onsite and offshore SRE teams; mentor junior engineers and foster SRE culture
MUST-HAVE SKILLS & TECHNOLOGIES ⚙ Core Technical Skills
🔧 SRE Practices & Leadership
▸ Akamai CDN (edge config, WAF, TLS, traffic mgmt)
▸ SLI / SLO / Error Budget management
▸ AWS Core Services (EC2, EKS, RDS, S3, CloudWatch)
▸ Incident Management & Blameless Post-Mortems
▸ Kubernetes / Amazon EKS (production operations)
▸ Performance tuning & capacity planning
▸ Jenkins CI/CD (pipeline design & optimization)
▸ AWS IAM, VPC, Security Groups, KMS
▸ Splunk APM & Splunk Cloud (dashboards, alerts)
▸ Container security & image scanning
▸ Redis & Memcached (distributed caching strategy)
▸ Monitoring stack design (metrics, logs, traces)
▸ CloudFormation / Auto Scaling Groups
▸ Distributed systems troubleshooting at scale
▸ Microservices architecture & service mesh
▸ Cross-team technical leadership & mentoring