Site Relablity Engineer
Listed on 2026-08-22
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Project Manager
Wipro Limited (NYSE: WIT, BSE: 507685, NSE: WIPRO) is a leading technology services and consulting company focused on building innovative solutions that address clients’ most complex digital transformation needs. Leveraging our holistic portfolio of capabilities in consulting, design, engineering, and operations, we help clients realize their boldest ambitions and build future-ready, sustainable businesses. With over 230,000 employees and business partners across 65 countries, we deliver on the promise of helping our customers, colleagues, and communities thrive in an ever-changing world.
For additional information, visit us at
This role is primarily focused on Site Reliability Engineering and Production Operations, with strong emphasis on Dev Ops automation, incident management, platform reliability, security, and identity management. The ideal candidate should have hands‑on experience in production support, CI/CD pipelines, infrastructure automation, NGINX administration, AWS, authentication technologies (SAML/SSO/OAuth), certificate management, and operational excellence initiatives.
- Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operations, and refinement.
- Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
- Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
- Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in Dev Ops automation and best practices.
- Practice sustainable incident response and blameless postmortems.
- Take a holistic approach to problem solving, by connecting the dots during a production event thru the various technology stack that makes up the platform, to optimize mean time to recover
- Collaborate with a global team spread across tech hubs in multiple geographies and time zones
- Share knowledge and mentor junior resources.
- Develop and maintain automation pipelines for certificate renewal, traffic routing, alerting, and compliance reporting using tools like Ansible, Venafi.
- Drive improvements in ITSM and DQ SLOs, ensuring timely CRQ status updates and incident closure.
- Lead initiatives for Safety & Soundness and Operational Excellence across quarterly EPICs, covering areas such as PCI compliance, threat/toil management, self-healing, and ITSM defect resolution
- Background in operational resiliency and self-healing systems.
- Understanding of two factor authentication.
- Strong documentation and communication skills.
- Strong understanding and experience in implementing NGINX configuration.
- Intermediate understanding of Active Directory (Users / Groups), SAML, LTPA, SSO, Oauth.
- Understanding of DEVOPS technologies like Chef, Jenkins, Groovy, shell scripting, bitbucket, GIT.
- Experience in working with or implementing automation workflows and/or scripting development.
- Understanding of:
- Client-server relationships
- Network concepts (Layer 1 to Layer
3) - Stack trace analysis (TCP dumps, heap dumps, CPU/memory analysis, thread dumps).
- Load balancers and application firewalls.
- Operating System navigation.
- Logging and monitoring methods, standards, and tools.
- High availability and business continuity planning
- Caching concepts
- Configuration management
- Awareness of security implementations, certificate management lifecycle, mutual TLS, SSL handshake, SSH keys, symmetric and asymmetric encrypt ions.
- Experience with AWS infrastructure and secure access practices.
- Familiarity with ITSM processes, compliance frameworks, and incident management.
- Excellent communication and collaboration skills across cross-functional teams
Mandatory
Skills:
Site Reliability Engineering (SRE).
Experience:
3-5 Years.
The expected compensation for this role ranges from…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).