SRE Engineer
Listed on 2026-07-21
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description
- Develop and maintain tooling used for environment monitoring and task automation
- Identify application reliability and availability improvements and build solutions to drive an improved experience
- Analyze and establish efficient configurations for software and servers, database connections, indexes, drivers, etc.
- Coordinate with development teams, technical and non‑technical partners and clients to maintain wide knowledge on dependencies of the critical business transaction, including platform, services and tools
- Monitor internal and vendor service level objectives (SLOs) and agreements (SLAs); identify and resolve SLO/SLA gaps
- Serve as technical subject‑matter expert (SME) for cross‑functional engineering teams
- Assist with and troubleshoot systems‑related issues and maintenance
- Collaborate on maintaining services once they are live; measure and monitor availability, latency, and overall system health
- Develop run book and build automation
- Develop and maintain end‑to‑end monitoring dashboards to support critical business transaction
- Develop and maintain synthetic monitoring for critical business transaction using tools such as Thousand Eyes
- Practice sustainable incident response and blameless post‑mortems
- Document and promote SRE standards and procedures
- Develop and assist in deployment and rollback automation
- Review release and deployment requirements
- Build and set up automation tests
- Incident communication to impacted stakeholders
- Coach and mentor junior engineers and fellow practitioners
Experienced SRE Engineer in designing, managing and supporting distributed systems across multi‑cloud environments.
Key Skills & Expertise- CI/CD:
Git Hub, Harness - Cloud Platforms: GCP, PCF, AWS
- Monitoring & Observability:
Splunk, Grafana, App Dynamics, Thousand Eyes - Containers & Orchestration:
Docker, Kubernetes, Cloud Foundry - Messaging & Streaming:
Kafka, MQ - Protocols & Web Services: HTTP, DNS, TCP/UDP, REST, SOAP, JSON
- Strong troubleshooting and debugging in microservices architecture
- Incident management, issue resolution and RCA creation
- Multi‑cloud platform management (SRE practices)
- Enterprise cloud infrastructure handling
- Agile development practices with tools like Git, Jira, Confluence
- Site Reliability Engineering (SRE)
Experience:
5–8 years.
Expected compensation: $60,000 to $135,000, variable based on location, minimum wage obligations, skills and relevant experience.
Equal Opportunity EmployerWe are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, caste, creed, religion, gender, marital status, age, ethnic and national origin, gender identity, gender expression, sexual orientation, political orientation, disability status, protected veteran status, or any other characteristic protected by law.
Wipro is committed to creating an accessible, supportive, and inclusive workplace. Reasonable accommodation will be provided to all applicants, including persons with disabilities, throughout the recruitment and selection process. Accommodations must be communicated in advance of the application, where possible, and will be reviewed on an individual basis. Wipro provides equal opportunities to all and values diversity.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).