Site Reliability Engineer; Hybrid
Listed on 2026-07-19
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
The Company
Unlimited Systems authors the category‑leading Unlimited Financials practice management system focused on the unique revenue cycle requirements of specialty healthcare providers. Unlimited Systems customers enjoy streamlined business office workflows, reduced claim denial rates, and accelerated and amplified revenue streams. Unlimited Systems is committed to ensuring that specialty healthcare providers thrive in a dynamic reimbursement environment.
Unlimited Systems is a portfolio company of Francisco Partners, a leading technology investment firm with deep sector focus and a track record of delivering outstanding returns. Through private equity and credit funds, they provide flexible capital and partnership to growth‑aspiring technology companies.
The ChallengeUnlimited Financials operates across a sophisticated cloud‑native environment — Azure Kubernetes Service, event‑sourced microservices, healthcare integrations, and a growing network of downstream data pipelines — all of which requires dedicated technical support to remain stable, observable, and performant for our customers. We are building the reliability practice this platform deserves. As we scale into new specialty healthcare markets and take on greater complexity, we need someone who thinks proactively — not just responding to incidents but engineering the systems and signals that prevent them.
This is not a run‑and‑maintain role. It is a “build‑and‑shape‑the‑future” role.
Our Site Reliability Engineer will own the observability strategy across our observability applications, define and drive meaningful SLOs and SLIs for our critical services, and work shoulder‑to‑shoulder with Dev Ops and Software Engineering to make reliability a first‑class concern — not an afterthought. You will bring structure to chaos, clarity to ambiguity, and measurable improvement to customer experience.
The PositionA full‑time, hybrid role reporting to our Technical Support Manager, our Site Reliability Engineer will work frequently from our headquarters in Cincinnati, Ohio. Residency in the state of Ohio is preferred.
As Site Reliability Engineer, you will own the reliability, scalability, and performance of our production systems hosted on Azure Kubernetes Service (AKS), working at the intersection of software engineering and operations to build the tooling, processes, and culture that keep our services running will be a key contributor to our observability practice — using Splunk for log analytics and alerting, Instana for APM and distributed tracing, and native Azure tools including Azure Monitor, Log Analytics, and Application Insights to provide a comprehensive, real‑time view of system health.
ResponsibilitiesReliability & Incident Management
- Defining, tracking, and reporting on SLIs, SLOs, and error budgets for all critical services.
- Designing and maintaining runbooks, escalation paths, and on‑call rotation schedules.
- Designing chaos engineering practices to proactively surface reliability weaknesses before they impact customers.
Observability — Splunk, Instana & Azure
- Building and maintaining Splunk searches, dashboards, and alert policies covering application and infrastructure logs.
- Developing KPIs and unified service‑health views for engineering and leadership.
- Configuring and extending instrumentation across microservices for distributed tracing and real‑time baselining.
- Creating smart alerts integrated with on‑call and ticketing systems for automated incident routing.
- Maintaining Azure Monitor alert rules, action groups, and workbooks across Azure subscriptions.
- Utilizing Log Analytics work spaces including data‑retention policies and ingestion cost governance using KQL.
- Leveraging Application Insights for APM, availability testing.
- Driving convergence of Splunk, Instana, and Azure signals into a unified observability strategy.
- Building internal tooling in Python, Go, or Bash to eliminate toil and accelerate incident response.
- Participating in security reviews and threat‑modeling sessions for new platform capabilities.
The opportunity to cast a vision for success, to blend art and science with proactive strategy and tactical execution,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).