What are the responsibilities and job description for the AWS DevOps Engineer position at Reqroute, Inc?
Remote Role
Position Title: Senior AWS DevOps/Platform Engineer
Location: Denver CO/ Remote
Duration: 12 Months Contract
Position Type- W2 Only
Exp Level- 10 Years
Req Skills- Platform Engineer, AWS, EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch, (infrastructure-as-code, Terraform preferred, CloudFormation), (AI agent deployments using GitLab CI/CD, Docker, Artifactory), AWS architectures, VPC peering, PrivateLink, transit gateway configurations, Docker, CI/CD pipelines (GitLab CI preferred), (Kafka, SNS/SQS, EventBridge), (CloudWatch, Prometheus, Grafana, or similar), (DNS, CIDR, NAT, VPC endpoints, firewalls), (REST APIs, Cloud SDKs, AWS SDK / boto3), (IIA Data Lake, S3 storage, Glue data catalog, Athena query engine, EMR compute clusters)
JOB SUMMARY
- Platform Engineer designs, builds, and maintains the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems, and develops the application code, automation, and internal tooling for those systems.
- This role ensures production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.
MAJOR DUTIES AND RESPONSIBILITIES
- Responsibilities span the following areas. Individual focus areas will be determined based on team needs and candidate strengths:
Infrastructure and Data Lake
- Design and manage AWS infrastructure for the IIA Data Lake including S3 storage, Glue data catalog, Athena query engine, and EMR compute clusters.
- Manage cross-account connectivity, VPC networking, security groups, and IAM roles/policies to enable secure data flow between IIA, upstream data providers, and downstream consumers.
- Build and maintain infrastructure for AI agent runtime environments, including compute resources for LangGraph agents deployed via LangSmith Deployments.
- Design and implement event-driven infrastructure including triggers, message queues, and pub/sub frameworks that orchestrate agent execution and data pipeline workflows
- Support deployment and operation of AWS Neptune for the network topology graph (digital twin), including capacity planning, schema design support, and performance tuning.
- Implement and manage infrastructure-as-code (Terraform, CloudFormation) for repeatable, auditable environment provisioning.
- Manage IAM access key rotations, secrets management (AWS Secrets Manager, Delinea), and security compliance for on-premises and cloud integrations (e.g., Splunk Edge Processor).
CI/CD and Agent Deployments
- Build and maintain CI/CD pipelines for AI agent deployments using GitLab CI/CD, Docker, and Artifactory.
- Manage container lifecycle for agents deployed via LangSmith Deployments, including image builds, versioning, and rollback procedures.
- Automate deployment workflows to enable rapid, reliable promotion of agents from development through production.
- Coordinate with SpecGPT platform team on AI Gateway integration, cross-account deployment, and connectivity requirements.
Application Development and Tooling
- Support and maintain existing ETL pipelines written in Scala/Spark, including troubleshooting, enhancements, and onboarding new data sources.
- Develop and maintain application code: scripts, CLIs, small services, and automation utilities (primarily Python) that support data ingestion, deployment, environment provisioning, and operational workflows.
- Write integration code and glue services that connect IIA systems with upstream data providers, downstream consumers, and external platforms.
Production Operations
- Ensure production environment stability through monitoring, alerting, and incident response. Maintain SLAs for data pipeline availability and agent uptime.
- Implement production monitoring and alerting for deployed agents (health checks, error rates, latency, resource utilization).
- Coordinate with upstream data teams and platform teams (SpecGPT, Splunk, Public Cloud) on connectivity, firewall requests, and integration requirements.
- Support data engineering team with infrastructure needs for new data source onboarding (storage provisioning, access controls, pipeline compute) and coordinate on scheduler maintenance and development (Airflow, event-driven triggers).
- Perform other duties as required.
REQUIRED QUALIFICATIONS
Skills/Abilities and Knowledge
- Strong communication skills with ability to explain infrastructure decisions to non-infrastructure stakeholders
- Expert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch
- Strong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable)
- Experience managing cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway configurations
- Experience with IAM policy design, least-privilege access patterns, and service account management
- Experience with containerization (Docker) and container orchestration
- Experience with CI/CD pipelines (GitLab CI preferred)
- Experience with event-driven architectures and messaging systems (e.g., Kafka, SNS/SQS, EventBridge)
- Proficiency with Linux-based operating systems and shell scripting
- Experience with monitoring and alerting tools (CloudWatch, Prometheus, Grafana, or similar)
- Understanding of networking fundamentals: DNS, CIDR, NAT, VPC endpoints, firewalls, security groups
- Demonstrated ability to work across teams and coordinate with external platform owners on connectivity and access requirements
- Proficiency in Python (or a comparable general-purpose language) for building automation, tooling, and applications
- Solid software engineering fundamentals: Git-based workflows, code review, modular and reusable design, dependency management, and writing maintainable, documented code
- Experience writing automated tests (unit/integration) for application and infrastructure code, and integrating those tests into CI/CD
- Ability to write integration code against REST APIs and cloud SDKs (e.g., AWS SDK / boto3)
PREFERRED QUALIFICATIONS
Skills/Abilities and Knowledge
- Experience with graph databases (AWS Neptune, Neo4j) including deployment, scaling, and operational management
- Experience with Apache Kafka or similar streaming platforms
- Experience with Apache Spark (Scala preferred) for distributed and streaming data processing such as Spark Streaming or structured streaming.
- Experience with Airflow or similar workflow orchestration platforms
- Experience in the telecommunications industry or other large-scale network operations environments
- Familiarity with AI/ML infrastructure requirements (model serving, GPU/CPU compute, artifact management via MLflow or similar)
- Experience with Splunk integration, particularly Edge Processor and MCP connectivity
- AWS certifications (Solutions Architect, DevOps Engineer, or similar)
- Experience developing and operating small services or APIs (e.g., FastAPI/Flask) in a production environment