What are the responsibilities and job description for the Site Reliability Engineer (SRE) for Google Cloud Platform (GCP) position at Matlen Silver?
Site Reliability Engineer (SRE) for Google Cloud Platform (GCP), focused on building, maintaining, and improving the reliability, scalability, and performance of cloud infrastructure using Infrastructure as Code (IaC) and Terraform Enterprise. Supports the delivery of secure, compliant, and highly available cloud environments aligned with enterprise standards and regulatory requirements.
Works closely with engineering and platform teams to develop and maintain reusable IaC modules, Terraform configurations, and automated cloud services, enabling consistent and efficient infrastructure provisioning. Contributes to the implementation of standardized platform patterns, including networking, identity, logging, and monitoring capabilities.
Participates in the end-to-end lifecycle of cloud infrastructure, including deployment, monitoring, incident response, and continuous improvement. Helps implement and maintain CI/CD pipelines, policy-as-code frameworks, and automation solutions to ensure reliable and repeatable deployments.
Applies SRE principles and practices, including monitoring, alerting, incident management, and root cause analysis, to improve system reliability and reduce operational risk. Supports the definition and tracking of service performance through metrics such as availability and latency.
Collaborates with architecture, security, and engineering teams to ensure infrastructure is secure, compliant, and operationally resilient. Contributes to DevSecOps practices by integrating security and compliance controls into automated workflows.
Continuously identifies opportunities to improve system reliability, reduce manual effort, and enhance automation. Leverages emerging tools and technologies, including AI/ML where applicable, to support proactive operations, observability, and platform stability.
Key Responsibilities
• Design, develop, and maintain Google Cloud Platform (GCP) infrastructure using Infrastructure as Code (IaC) with Terraform Enterprise
• Contribute to the implementation of scalable, secure, and compliant cloud solutions aligned with enterprise standards
• Develop and maintain reusable Terraform modules and standardized infrastructure patterns to enable consistent and automated provisioning of GCP resources
• Follow and contribute to code quality standards, design patterns, and peer review practices to ensure reliable and maintainable infrastructure code
• Support the adoption and use of Terraform Enterprise for automated provisioning, policy enforcement, and infrastructure governance
• Implement and maintain cloud automation workflows, including provisioning, configuration management, and environment setup
• Build and enhance CI/CD pipelines for infrastructure delivery, ensuring automated testing, validation, and compliance checks
• Implement policy-as-code and security controls, ensuring infrastructure meets regulatory and enterprise compliance requirements
• Participate in the end-to-end lifecycle of infrastructure delivery, including deployment, monitoring, and continuous improvement
• Collaborate with architecture, security, and engineering teams to ensure secure, resilient, and compliant cloud configurations
• Apply DevSecOps and cloud-native practices to improve automation, security, and deployment efficiency
• Contribute to observability, logging, and monitoring solutions to support proactive incident detection and response
• Execute testing and validation of IaC modules, including integration and deployment verification
• Identify opportunities to automate manual processes and improve operational efficiency
• Support reliability, scalability, and performance of cloud platforms through automation and standardization
• Troubleshoot and resolve infrastructure and platform issues, contributing to root cause analysis and continuous improvement
• Work with stakeholders to implement infrastructure solutions that meet technical and business requirements
• Evaluate and adopt emerging tools and technologies to enhance automation, reliability, and platform performance
• Conduct performance testing and capacity planning to ensure systems scale reliably under load.
• Optimize system performance, latency, and resource utilization across cloud environments
• Design and implement observability solutions, including metrics, logs, traces, and alerting strategies.
• Reduce alert fatigue and improve signal quality through meaningful alert design and tuning.
• Develop dashboards and monitoring frameworks aligned to SLOs.
Required Qualifications
• 4-8 years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
• Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise for infrastructure provisioning
• Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code
• Experience developing and maintaining Terraform modules and infrastructure configurations to support automated cloud environments
• Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes
• Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows
• Good understanding of GCP services and cloud architecture fundamentals, including networking (VPCs, IAM, load balancing)
• Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
• Experience supporting automation and standardization efforts to improve consistency and efficiency in cloud deployments
• Understanding of monitoring, logging, and observability tools to support system reliability and performance
• Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
• Ability to collaborate effectively with engineering, architecture, and security teams to support reliable and secure platform operations
• Strong problem-solving and analytical skills, with the ability to diagnose and resolve infrastructure issues
• Effective communication skills, with the ability to work within cross-functional teams and document technical solutions clearly
• Interest in learning and applying emerging technologies and automation techniques (including AI/ML where applicable) to improve platform reliability