What are the responsibilities and job description for the Observability Engineer position at METRIX IT SOLUTIONS INC?
JD:-
Role : Observability Engineer (Open Telemetry / LGTM / Ansible)
Location: Waltham MA
Contract
Overview
Customer is embarking a strategic initiative to establish a modern, observability capability, as part of a broader platform transformation. This role will be instrumental in designing and implementing an Open Telemetry based observability framework and enabling a successful platform go-live.
The engagement focuses on delivering scalable, enterprise-grade observability aligned with best practices, while enabling enhanced operational visibility and engineering productivity across distributed teams.
Scope of Services
- Design and implement an Open Telemetry driven observability solution using LGTM stack (Grafana, Loki, Tempo, Mimir)
- Deploy and manage services on Azure Kubernetes Service (AKS) and VMWare Kubernetes Services (VKS)
- Support Azure Landing Zone configuration aligned to governance models
- Deliver Infrastructure as Code (IaC) using Terraform
- Implement automation using Ansible
- Create and validate YAML configurations
- Support platform go-live activities including validation, remediation, and stabilisation
- Collaborate with multi-region engineering teams.
Key Deliverables
- Functional Open Telemetry PoC
- Scalable observability framework
- Successful go-live support
- Improved monitoring and visibility
- Support documentation and knowledge transfer
Required Capability
- Core Technical Expertise
- Kubernetes (AKS)
- Open Telemetry
- LGTM stack (Grafana, Loki, Tempo, Mimir)
- Azure platform and Landing Zones
- Terraform and Ansible
- Exposure to SolarWinds
- Exposure to Splunk
- Exposure to ServiceNow ITSM
- Programming & Scripting
- Python
- YAML
- Bash
- Exposure to Go (desirable)
Value Proposition
- Accelerate adoption of modern observability
- Improve incident detection and resolution
- Enable scalable cloud-native monitoring
- Enhance engineering productivity
Preferred Experience
- Transformation programmes
- DevOps / SRE practices
- Enterprise observability implementations.