What are the responsibilities and job description for the Site Reliability Engineer position at Ascyndent?
At Ascyndent we believe Digital transformation is all about leveraging information and technology to enhance the human experience. It is also about transformative execution. Our teams bring business visions and strategies into reality leveraging technology capabilities.
Our consultants leverage innovative capabilities and deep domain expertise to develop short and long-term strategies aligned with your enterprise goals, skill requirements and technology platforms. As an experienced member of our team, we look first and foremost for people who are passionate around solving business problems through innovation and engineering practices. We embrace a culture of experimentation and constantly strive for improvement and learning. You'll work in a collaborative, trusting, thought-provoking environment-one that encourages diversity of thought and creative solutions that are in the best interests of our customers globally.
This job is for you if
- You’re interested in a meaningful job at a growing technology company where you’ll work with a variety of technology solutions and be appreciated for doing your job well
- You want to help shape the future of emerging technology
- You enjoy collaborating with highly skilled professionals with a passion for selling
- You want to join a culture that embraces dynamic individuals with a focus on innovation
.
A day in the lif
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
- Supports the adoption of site reliability engineering best practices within your team
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirement
What we are looking for…
Required
- Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar
- Proficient in at least one programming language (e.g., Python, Java, Ruby, AWS) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines
- Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production