Incident Analyst - In Office
Listed on 2026-08-21
-
IT/Tech
SRE/Site Reliability
Incident Analyst
Analyst III – Reliability Operations Role Summary
The Analyst III is a senior operational leader who acts as Incident Commander during major outages, leads problem management efforts, and reviews/approves complex changes for operational readiness. This role partners closely with SRE and engineering teams on reliability strategy.
Key Responsibilities
- Incident Management
- Lead high severity P1/P0 incidents as Incident Commander.
- Coordinate cross functional engineering teams in real time.
- Drive rapid troubleshooting, impact assessment, and resolution decisions.
- Ensure high-quality incident documentation, executive-ready summaries, and follow through
- Apply Technical knowledge of Application architecture flows in driving the incident towards mitigation
- Problem Management
- Lead problem investigations for major or recurring incidents.
- Perform deep root cause analysis with engineering teams.
- Validate corrective actions and track long-term problem remediation.
- Present problem findings and preventive strategies to leadership.
- Change Management
- Review and approve high-risk or complex changes for operational readiness.
- Participate in Change Advisory Board (CAB) when needed.
- Validate rollback strategies and operational safety measures.
- Lead change execution for major maintenance or reliability events.
- Reliability Leadership
- Drive reliability initiatives to reduce MTTA, MTTR, and incident volume.
- Mentor junior analysts on incident handling and operational maturity.
- Partner with SRE teams to expand observability, automation, and resilience.
- Partner with SRE/Engineering teams on service reliability initiatives.
- Lead maintenance events, failover tests, and resilience validation exercises.
- Review and enhance runbooks, automation workflows, and monitoring strategies
- Contribute to Automation & AI Ideas for improving efficiency & reduce MTTM /MTTR
- Reliability Engineering Contributions
- Perform deep post incident analysis to identify systemic issues.
- Contribute to automation solutions and self healing systems.
- Own service reliability dashboards and operational KPIs.
- Leadership & Mentoring
- Provide technical leadership to Analyst I/II team members.
- Serve as a point of escalation for complex incidents.
- Lead operational readiness reviews and training sessions.
Required Qualifications
- 3–5 years in SRE, Operations, Incident Management, Dev Ops, or related fields.
- Expert knowledge of monitoring tools, logging systems, and incident response.
- Strong troubleshooting skills: networking, Linux/Windows servers, cloud services.
- Strong communication and leadership skills in high-pressure situations.
Preferred Qualifications
- Hands-on experience with cloud platforms (AWS, Azure, GCP).
- Automation experience (Python, Go, Bash, or Power Shell).
- Familiarity with microservices, containers, and distributed architectures.
Mandatory
Skills:
Network configuration.
Experience:
3-5 Years.
The expected compensation for this role ranges from $45,000 to $110,000. Final compensation will depend on various factors, including your geographical location, minimum wage obligations, skills, and relevant experience. Based on the position, the role is also eligible for Wipro's standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options.
Wipro provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. Applications from veterans and people with disabilities are explicitly welcome.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).