What are the responsibilities and job description for the Observability Operations Engineer position at Decision Six Inc.?
Title: Observability Operations Engineer
Location: Phoenix, AZ (Onsite) - Need locals only
Duration: Long Term Contract
Primary Skill:
Observability, Splunk Enterprise Administration, Kubernetes
Job Description:
- We are seeking a highly skilled Senior Observability Operations Engineer to manage and enhance our enterprise observability platform.
- The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions.
- Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.
- The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.
- Required Technical Skills Observability Platforms Dynatrace Administration Splunk Enterprise Administration OpenSearch Administration Elasticsearch Administration Grafana Prometheus Kibana Jaeger Open Telemetry Kafka (preferred)Administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
- Configure Dynatrace One Agent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).
- Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
- Develop dashboards, alerts, reports, and executive operational metrics.
- Support Linux-based infrastructure and Kubernetes environments (Docker/OpenShift/Rancher preferred).
- Implement observability best practices using Open Telemetry, distributed tracing, metrics, logs, and events.
- Perform root cause analysis for production incidents using observability platforms.
- Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams. Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible..
- Improve platform reliability through automation, self-healing, and AI-assisted operations Dynatrace Administration Splunk, Open Search Administration ELK, Prometheus, Grefana Kibana Kafka Docker, Kubernetes Any Cloud experience.
Salary : $40 - $50