×
Register Here to Apply for Jobs or Post Jobs. X

Customer Reliability Engineer

Job in Chicago, Cook County, Illinois, 60601, USA
Listing for: iManage
Full Time position
Listed on 2026-07-03
Job specializations:
  • IT/Tech
    IT Support, Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Job Description & How to Apply Below

Customer Reliability Engineer

We offer a flexible working policy that supports a healthy balance between personal and professional well-being. This role requires in-office presence on Tuesdays & Thursdays to collaborate, connect, and learn from peers - while also maintaining the flexibility for meaningful work-life balance. Being a Customer Reliability Engineer at iManage means…

You're a data driven problem solver with experience in triaging issues  are an expert at explaining complex technical problems to a variety of audiences. You are a passionate customer advocate and work proactively to anticipate and address or raise visibility to problems in a data driven way.

You will be obsessed with uptime. When production breaks, you lead the charge. If it hasn't broken yet, you're already building the monitors, queries, and automation that make sure it doesn't.

The Customer Reliability team is expert in consumption and reliability of the iManage Cloud platform across services and interfaces (first party & partner). They guide the way we anticipate, communicate, and react to reliability issues while driving iterative improvements across products & services both internally and externally. When there are multi-faceted reliability problems which don't fit existing constructs, they engage as a customer advocate to guide data-driven decision making and holistic improvements across the platform based on lessons learned.

iM

Responsible For…
  • Developing and maintaining a deep technical knowledge of iManage platform services, with the ability to personally deep dive into logs, queries, and infrastructure to unblock different teams.
  • Collaborating with customer-facing, product, and infrastructure teams on the development and deployment of scalable, reliable software.
  • Drive a shift from reactive to proactive. Identify what is breaking before it breaks. Use telemetry, anomaly detection, and trend analysis to surface systemic issues then partner with technical stakeholders to eliminate them at the source.
  • Serving as incident commander on P1/P2 incidents: leading communication, coordinating engineering stakeholders to get services back up and running again.
  • Continuously refining the observability stack — building and tuning dashboards, alerts, and synthetic monitoring that give real-time visibility into system and end-user experience health.
  • Serving as the escalation point for the Platform Support team, acting as Subject Matter Expert (SME) on the consumption and observability of our services; sharing knowledge to level-up expertise across the team.
  • Engaging with Engineering and SRE teams to improve supportability through internal tooling and observability; leveraging large-scale data sets to troubleshoot complex emerging problems.
  • Assisting with remediation, tooling, and communication for Customer Advisories; driving critical technical escalations with engineering teams through problem isolation, data gathering, and validation of resolution.
  • Personally leading post-incident root cause analysis; owning postmortems end-to-end, running blameless postmortems reviews, and driving systemic fixes that prevent recurrence.
  • Advocating for users and stakeholders by exposing friction and reliability concerns within the products.
  • Driving automation and self-service that eliminates repeat tickets and matures our knowledge management strategy; implement AI-powered operational improvements to improve quality and speed.
  • Engaging with iManage partners proactively to build agreements and shared understanding of reliability, levelling up knowledge of reliability principles across the iManage ecosystem.
iM Qualified Because I Have…
  • 3–5 years' experience in a technical escalation role in a Support, CRE, Development, or SRE function.
  • Proven ability to understand and communicate highly complex issues for non-technical audiences, including executive stakeholders.
  • Seasoned incident responder with hands-on experience leading P1/P2 response, running postmortems, and driving systemic fixes that prevent recurrence.
  • Experience troubleshooting and supporting distributed cloud services.
  • Strong technical proficiency across SQL, Python, Bash/Shell, Power Shell, and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary