What are the responsibilities and job description for the AWS AgentCore Platform Engineer position at Galent?
AWS AgentCore Platform Engineer
Location: Reading, PA (Hybrid - 2 to 3 days onsite per week)
Duration: Contract
Job Description
We are seeking an experienced AWS AgentCore Platform Engineer to design, implement, and optimize observability and reliability solutions for next-generation AI agent platforms built on AWS technologies.
Key Responsibilities
AI Platform Observability & Reliability
- Design and implement enterprise-grade observability solutions for AI agent ecosystems built on AWS Bedrock, AWS AgentCore, and MCP (Model Context Protocol) Servers.
- Evaluate and optimize AWS CloudWatch, AWS X-Ray, Bedrock logging, and AgentCore tracing capabilities to support complex agentic workflows.
- Conduct gap analyses and implement observability frameworks using Dynatrace and other monitoring platforms.
- Develop and enhance distributed tracing solutions across AI agent platforms and microservices architectures.
- Establish telemetry, monitoring, alerting, and reliability best practices for AI-driven applications.
- Analyze system performance, identify bottlenecks, and drive continuous reliability improvements.
- Collaborate with engineering and platform teams to ensure end-to-end observability across agent ecosystems.
- Implement proactive monitoring strategies and automated incident detection mechanisms.
- Support root cause analysis and remediation of production issues impacting AI platform performance.
Required Skills
- Strong experience with AWS Bedrock and AWS AgentCore.
- Hands-on expertise with CloudWatch, AWS X-Ray, and distributed tracing technologies.
- Experience implementing observability solutions using Dynatrace.
- Knowledge of MCP (Model Context Protocol) and AI agent architectures.
- Experience with telemetry, monitoring, logging, and performance optimization.
- Strong understanding of cloud-native architectures and microservices.
- Familiarity with OpenTelemetry (OTel) and observability best practices.
- Experience in Site Reliability Engineering (SRE), Platform Engineering, or DevOps environments.
- Excellent troubleshooting, analytical, and problem-solving skills.
Preferred Qualifications
- Experience supporting Generative AI and Agentic AI platforms.
- Knowledge of enterprise observability frameworks and monitoring standards.
- AWS certifications are a plus.
- Experience with large-scale distributed systems and cloud infrastructure.