What are the responsibilities and job description for the AI-Enabled Platform/SRE Engineer position at Intracruit Solutions?
Role: AI-Enabled Platform/SRE Engineer
Location: Hybrid in Richardson - need local from TX !
Duration: 12 months extensions
Client can call for f2f interview for last round - they should be fine with that
About the Rol
eWe are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale
.
Key Responsibilit
ies• Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operatio
ns.• Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook executi
on.• Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategi
es.• Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooti
ng.• Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environmen
ts.• Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectiv
es.• Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improveme
nt.
Core Technical S
kills• Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excell
ence.• Kubernetes Platform Engineering – 5 years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimiza
tion.• Cloud & Infrastructure Automation – Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automa
tion.• Software Development – 5 years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workfl
ows).• Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monito
ring.• API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strate
gies.• AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workf