×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Northern, Floyd County, Kentucky, USA
Listing for: Uncover
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
Location: Northern

About Plenful

About Plenful Plenful is on a mission to transform healthcare operations from the inside out. Fresh off our $50M Series B and backed by Notable Capital, Bessemer Venture Partners, TQ Ventures, Susa/Kivu Ventures, and other leading investors, we’re building the category-defining AI workflow automation platform that healthcare teams rely on to operate smarter, faster, and more efficiently. Our technology empowers healthcare operators across hospital and health systems, pharmacies and payors to eliminate manual work, reduce administrative burden, and improve compliance, all while unlocking critical revenue to fund programs for their in-need patient populations.

Built by healthcare operators for healthcare operators, Plenful is driven by a deep understanding of the challenges facing today’s care teams. We’re passionate about equipping healthcare workers with world-class tools that deliver real, measurable impact, and we’re proud to serve 90+ leading health systems across the country.

About the Role

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.

You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

What You’ll Do Reliability Engineering & System Ownership
  • Define and implement SLIs, SLOs, and error budgets across core services.
  • Own production system health: uptime, latency, and availability targets.
  • Improve system resilience through proactive reliability work.
  • Find and mitigate single points of failure across distributed systems.
Production Operations & Incident Response
  • Take part in and improve on-call rotations and incident response.
  • Lead incident triage, mitigation, and resolution in real time.
  • Run blameless postmortems and follow through on action items.
  • Build tooling and automation to cut MTTR (Mean Time to Recovery).
Observability & System Insight
  • Design and evolve observability across metrics, logs, and distributed tracing (Open Telemetry), using tools like Datadog, Cloud Watch, Grafana, and Sentry.
  • Improve signal quality to cut noise and alert fatigue.
  • Build dashboards and alerts that reflect real system health and user impact.
  • Use observability data to drive performance and reliability improvements.
Performance & Scalability
  • Analyze system performance under load and find bottlenecks.
  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, Click House).
  • Partner with engineering teams to improve system efficiency and scaling behavior.
Automation & Reliability Tooling
  • Build automation that eliminates repetitive operational work.
  • Improve deployment safety through reliability checks and safeguards.
  • Contribute to CI/CD pipelines (Git Hub Actions) with a focus on stability.
  • Build tools for incident response, debugging, and capacity planning.
Security, Compliance & Operational Maturity
  • Partner with security and compliance to keep systems meeting operational standards.
  • Support audit readiness and reliability-related compliance requirements (Vanta).
  • Integrate monitoring and alerting into security and SIEM workflows.
  • Help mature operational practices across engineering.

You’ll know it’s working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.

You May Be a Fit If
  • You’ve spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary