Site Reliability Engineer
Listed on 2026-07-20
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Disaster Recovery IT
Overview
The Infosys Cloud unit is dedicated to empowering enterprises with innovative cloud solutions that drive digital transformation and operational excellence. We specialize in leveraging advanced cloud technologies and AI‑driven insights to create scalable, secure, and resilient infrastructures. Our solutions enable organizations to achieve agility, efficiency, and sustainable growth in a hyperconnected world. Join us to be part of a pioneering team at the forefront of cloud and AI innovation.
You'll have the opportunity to work with cutting‑edge technologies, collaborate with industry experts, and contribute to transformative projects that shape the future of business. We are committed to fostering a culture of continuous learning and growth, ensuring that our team members thrive in a dynamic and supportive environment. If you're passionate about cloud and AI, and eager to make a significant impact, the Infosys Cloud unit is the perfect place for you to grow and excel.
- Collaborate with internal and client teams to resolve complex incidents, conduct root cause analyses, and document findings with preventive recommendations
- Participate in evaluation of client IT infrastructure, prepare actionable assessment reports, and support due diligence to document infrastructure maturity and improvement opportunities
- Contribute to the design of scalable, cost‑effective IT infrastructure solutions, review reusable components, and develop technical documentation for deployed systems
- Align release schedules and environment readiness, execute deployments as per protocols, perform post‑deployment testing, and manage version control to track changes
- Co‑ordinate maintenance schedules, emergency fixes, and technology upgrades while ensuring uninterrupted integration into existing systems and processes
- Facilitate performance data analysis across systems, coordinate insights on system behavior, and support capacity planning to optimize performance
- Conduct security checks, recovery drills, and compliance audits, implement security measures, and coordinate continuity plans to maintain adherence to standards
- Gather feedback to identify automation opportunities, analyze existing infrastructure processes, and propose enhancements for efficiency gains
- Act as liaison with onsite, offshore, and vendor teams to document project requirements, ensuring effective collaboration
- Develop a centralized repository of technical and procedural knowledge, leveraging insights from other projects to drive efficiency and retain organizational expertise
- A collaborative spirit and excellent communication skills.
- Ability to handle complex incidents and implement resolutions.
- A knack for conducting IT infrastructure assessment and identifying key optimization opportunities.
- Focused approach towards deployment management, system optimization, and process automation initiatives including sector specific focus.
- The ability to work with cross‑functional teams.
- Support SLIs, SLOs, error budgets, and reliability KPIs.
- Drive service availability, resiliency, scalability, and performance improvements.
- Establish proactive operational models and reliability governance.
- Engage in reliability reviews and continuous improvement programs.
- Drive Infrastructure as Code adoption using Terraform.
- Develop reusable modules and platform standards.
- Implement policy‑as‑code and automation frameworks.
- Reduce manual infrastructure management through automation.
- Define observability standards and monitoring frameworks.
- Manage Datadog implementation including dashboards, alerts, APM, logs, and tracing.
- Establish observability‑as‑code practices.
- Improve alert quality and operational visibility.
- Define DR strategy and recovery objectives.
- Engage in DR testing, failover planning, and resiliency reviews.
Ensure business continuity readiness and compliance. - Manage recovery runbooks and operational procedures.
- Engage in periodic DR exercises, failover validations, recovery evidence collection, post‑DR…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).