What are the responsibilities and job description for the Site Reliability Engineer position at QualiTest?
Are you interested in working with the World’s leading AI-first Quality Engineering Company? Ready to advance your career, team up with global thought leaders across industries and make a difference every day? Join us at QualityAI!
We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States!
Responsibilities:
- Partner with Application Development teams to build resiliency for Payment application.
- Partner with our Application Develop teams to implement service level objectives.
- Partner with our Application Development teams and other SREs to build out end to end observability.
- Implement monitoring, alerting and dashboards needed for our apps.
- Automated operational processes.
- Help to develop our capacity management and performance management tools.
- Help to define the DR plan needed for our critical apps.
- Help to develop a chaos testing process.
- Participate in an on-call rotation and support production Incidents.
SRE Skillsets - Expectations from Pricing & Settlements team:
- Good understanding of hybrid infrastructure.
- Expertise with AWS.
- Expertise in one or more general purpose programming languages: Python, Go, shell scripting (Unix/Linux), Java.
- Experience in CI/CD pipelines preferably Jenkins expertise.
- Experience in container technology (OpenShift, Kubernetes).
- Expertise in automation tools experience (preferably Ansible).
- Expertise in observability tools including APM (Datadog), synthetic monitoring and log aggregation (Elk).
- Experience in dashboarding tools such as Grafana and Kibana.
- Understating of Agile concepts and experience in JIRA.
- Basic understating of Release Management.
- Hands-on experience on SNOW.
SRE Skillsets - Expectations from Data Platform team:
- Expertise in Message Broker (preferably Rabbit MQ, Kafka).
- Expertise on Hadoop, spark commands JSON formatting.
Qualifications:
- 6-12 years of overall experience.
- Professional experience as a Site Reliability Engineer (SRE).
- Software development “hands on” engineer with excellent understanding of SDLC Application delivery.
- Ability to translate functional and non-functional requirements into appropriate NFT Automation tests.
- Experience with DevOps, CI/CD tools.
- Good experience of Linux, AWS Cloud and on Prem deployments.
- Good experience in Systems Observability and APM tools, preferably Datadog.
- Strong ability to track and contribute to technical discussions around application integration and high-availability, resilience and observability.
- Strong JIRA knowledge.
Skills:
- Expertise in AWS-Lambda Services 4/5.
- Strong programming skills (Java, Python, Shell Optional Java Script) 4/5.
- Proficiency in Database concepts, Strong knowledge of SQL and experience with MySQL 4/5.
- APM tools - Datadog 4/5.
- Hands-on experience on SNOW.
Must have:
- Professional experience as a Site Reliability Engineer (SRE).
- Experience in performance testing, Ability to translate functional and non-functional requirements into appropriate NFT Automation tests.
- Experience of AWS Cloud Application (Must for sure).
- Experience of Linux, AWS Cloud(Must) and on Prem deployments.
- Good experience in Systems Observability and APM tools, preferably Datadog.
- Experience in dashboarding tools such as Grafana and Kibana.
- Strong ability to track and contribute to technical discussions around application integration and high-availability, resilience and observability.
- Expertise in one or more programming languages: Python, shell scripting (Unix/Linux), Java.
Nice to have:
- Hands-on experience on SNOW.
- Experience in container technology (OpenShift, Kubernetes).
- Strong JIRA knowledge.
- Basic understating of Release Management.
- Experience in CI/CD pipelines preferably Jenkins expertise.
Salary : $110,000 - $130,000