What are the responsibilities and job description for the Senior observability engineer position at wellsfargo?
Description
Title:Senior Observability Engineer
Location: Irving, TX
Duration: 12 months
Work Engagement: W2
Work Schedule: Hybrid 3 days in office/2 days remote
Benefits on offer for this contract position: Health Insurance, Life insurance, 401K and Voluntary Benefits
Summary:
In this contingent resource assignment, you may: Consult on or participate in moderately complex initiatives and deliverables within Infrastructure Engineering and contribute to large-scale planning related to Infrastructure Engineering deliverables. Review and analyze moderately complex Infrastructure Engineering challenges that require an in-depth evaluation of variable factors. Contribute to the resolution of moderately complex issues and consult with others to meet Infrastructure Engineering deliverables while leveraging solid understanding of the function, policies, procedures, and compliance requirements. Collaborate with client personnel in Infrastructure Engineering. Required Qualifications: Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work or consulting experience, training, military experience, education.
Key Responsibilities:
Observability Engineering
• Design, implement, and enhance observability solutions across enterprise infrastructure and platform services.
• Identify and onboard new telemetry sources including metrics, logs, events, API data, and platform health information.
• Establish standardized approaches for collecting, parsing, normalizing, enriching, and visualizing operational data.
• Evaluate monitoring gaps and implement solutions that improve visibility, reliability, performance, and operational awareness.
• Support observability initiatives for both traditional infrastructure and cloud-native platforms.
Data Collection & Integration
• Integrate telemetry from multiple technology platforms into Grafana, Splunk, and other analytics platforms.
• Work with APIs, exporters, collectors, agents, and endpoints to gather operational data.
• Develop data ingestion pipelines and data transformation processes.
• Design telemetry schemas, metadata standards, tagging strategies, and data models to improve consistency and usability.
• Validate data quality, accuracy, and completeness.
Dashboard & Reporting Development
• Design dashboards tailored to multiple audiences including:
○ Executive Leadership
○ Platform & Infrastructure Operations
○ Site Reliability Engineering (SRE)
○ Application Development Teams
○ Platform Consumers and Tenants
• Translate complex technical telemetry into meaningful KPIs, operational indicators, and business-oriented insights.
• Create drill-down experiences that allow users to move from executive summaries to detailed operational diagnostics.
• Develop reporting solutions using Grafana, Splunk, Power BI, Tableau, and related platforms.
Analytics & Query Development
• Build and optimize queries across multiple data platforms.
• Develop metrics calculations, aggregations, and correlation logic.
• Analyze trends, capacity utilization, availability, performance, reliability, and platform health.
• Create reusable dashboards, reports, and operational analytics.
Platform Collaboration
• Partner with engineering teams to identify telemetry requirements for new services and platforms.
• Assist teams in exposing operational metrics and logs for observability consumption.
• Provide guidance on monitoring standards, instrumentation, and operational readiness.
• Participate in incident investigations and root cause analysis activities.
Continuous Improvement
• Identify opportunities for automation and self-service observability capabilities.
• Improve operational processes through proactive monitoring and analytics.
• Help establish best practices, governance, standards, and observability patterns across the organization.
Key Requirements:
· Applicants must be authorized to work for ANY employer in the U.S. This position is not eligible for visa sponsorship.
Experience working with metrics, logs, events, traces, and infrastructure telemetry.
• Experience identifying operational data sources and integrating them into observability platforms.
• Experience working with APIs and service endpoints.
• Experience parsing structured and semi-structured data including:
○ JSON
○ XML
○ CSV
○ Log formats
• Experience creating actionable dashboards and reporting solutions.
• Experience supporting enterprise infrastructure, platform engineering, SRE, or operations teams.
• Strong troubleshooting and analytical skills.Monitoring & Observability Platforms
• Grafana
• Splunk Enterprise / Splunk Cloud
• Prometheus
• OpenTelemetry concepts
• Visualization and analytics tools
Strong experience with:
• PromQL
• Splunk SPL
• SQL
• REST API querying
• JSON parsing and manipulation
Experience with one or more of the following is beneficial:
• KQL
• Elasticsearch/OpenSearch Query DSL
• Power Query (Power BI)
• Python data processing