What are the responsibilities and job description for the AI-Enabled Platform/SRE Engineer- Local TX position at TechCafeHub LLC?
AI-Enabled Platform/SRE Engineer- Local TX (f2f)
Introduction:
As an AI-Enabled Platform/SRE Engineer in our team based in Texas, you will play a crucial role in building and maintaining automation tools, leveraging AI technologies, managing Kubernetes platforms, ensuring high availability, and driving operational excellence. You will work closely with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Responsibilities:
- Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
- Leverage AI and Generative AI technologies to automate alert analysis, incident response, operational workflows, and runbook execution.
- Implement API and microservices reliability solutions using various tools and strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- Ensure platform reliability and high availability through active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
- Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics.
- Drive SRE best practices and operational excellence by collaborating with cross-functional teams.
Requirements:
Required Skills:
- Site Reliability Engineering (SRE) experience
- 5 years of hands-on experience with GKE and Rancher RKE2
- Strong knowledge of Cloud & Infrastructure Automation, including Google Cloud Platform and Terraform
- Advanced programming skills in Python and Java
- Experience with observability & monitoring tools
- Proficiency in API & Microservices Engineering
- Knowledge of AI-Driven Operations (AIOps)
Preferred Skills:
- Experience with Continuous Integration/Continuous Delivery (CI/CD)
- Familiarity with Incident Management and Disaster Recovery practices
- Understanding of High Availability concepts
- Knowledge of Kubernetes, Routing, and Operational Excellence
Salary : $50 - $60