What are the responsibilities and job description for the Site Reliability Engineer (FedRamp Exp) position at Reqroute, Inc?
REMOTE ROLE
Position Title: Lead SRE Observability (FedRAMP Exp)
Location: Remote Role (PST hours)
Duration: 12 Months Contract
Position Type- W2 Only
Exp Level- 10 Years
Req Skills- Site Reliability Engineering/Platform Engineering/DevOps, (Splunk Enterprise or Splunk Cloud), Splunk SPL, Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka, Terraform and Infrastructure as Code, Python, Go, Ruby, or Bash, UNIX, Kafka, Terraform, Kubernetes, Docker,
About the Role
- Join our Observability team responsible for designing, building, and operating enterprise platforms for logging, metrics, tracing, and alerting across large-scale cloud infrastructure.
- You'll lead initiatives that improve reliability, scalability, and operational excellence.
Key Responsibilities
- Design, deploy, and operate enterprise observability platforms.
- Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
- Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
- Design, deploy, and support distributed tracing platforms using Grafana Tempo and Open Telemetry.
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
- Scale Prometheus, Grafana, Kafka, Tempo, and Open Telemetry-based monitoring solutions.
- Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
- Automate infrastructure using Terraform and configuration management tools.
Required Qualifications
- 7 years in Site Reliability Engineering, Platform Engineering, or DevOps.
- Hands-on experience administering Splunk Enterprise or Splunk Cloud.
- Strong knowledge of Splunk SPL.
- Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
- Experience implementing metrics, logs, and traces as part of a modern observability strategy.
- Experience with Terraform and Infrastructure as Code.
- Programming experience in Python, Go, Ruby, or Bash.
Preferred Qualifications
- Splunk certification.
- Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
- Experience supporting FedRAMP or regulated environments.