×
Register Here to Apply for Jobs or Post Jobs. X

Senior Inference Reliability Engineer

Job in San Mateo, San Mateo County, California, 94403, USA
Listing for: Parasail
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Senior/Staff Inference Reliability Engineer

Parasail is redefining AI infrastructure by enabling seamless deployment across a distributed network of GPUs, optimizing for cost, performance, and flexibility. Our mission is to empower AI developers with a fast, cost-efficient, and scalable cloud experience—free from vendor lock-in and designed for the next generation of AI workloads.

The Senior/Staff Inference Reliability Engineer will own the end-to-end reliability and production performance of customer inference workloads. This role sits at the intersection of inference platform engineering, LLM performance, and infrastructure reliability.

You will ensure that customer endpoints meet expectations for availability, latency, throughput, quality, and cost. When an endpoint degrades, you will follow the problem across the entire serving path—from APIs, routing, scheduling, and autoscaling through model servers, GPUs, networking, and underlying infrastructure—and drive it through resolution.

This is not a traditional Dev Ops role focused only on clusters and deployments. It is a production systems role for an engineer who enjoys investigating ambiguous performance problems, building diagnostic tooling, and turning recurring incidents into durable platform improvements.

Prior LLM-inference experience is valuable but not required. We are looking for someone with deep production systems experience who can quickly learn inference-specific technologies and metrics.

What You Will Own End-to-End Inference Reliability

Own the production health of customer inference workloads, including availability, request success, time to first token, inter-token latency, throughput, and operational efficiency.

Establish clear service-level indicators, objectives, performance baselines, and escalation paths for production endpoints.

Detection and Observability

Build the telemetry, dashboards, alerts, and automated diagnostics needed to detect meaningful endpoint degradation before customers report it.

Create visibility across the full inference-serving path, including request queues, routing, scheduling, model servers, GPU utilization, networking, storage, and provider infrastructure.

Production Investigation

Lead the investigation of complex latency, throughput, capacity, and reliability regressions.

Determine whether an issue originates in customer traffic patterns, platform services, inference-engine configuration, GPU hardware, networking, storage, or an external infrastructure provider.

Remain accountable for the customer outcome while partnering with the appropriate engineering teams to implement the fix.

Incident Response and Prevention

Help lead customer-impacting incidents and establish effective operational practices for acknowledgement, diagnosis, recovery, and communication.

Convert significant incidents into automated tests, safeguards, runbooks, capacity controls, anomaly detection, and platform improvements.

Performance and Capacity

Partner with the LLM Performance team to validate that engine-level optimizations deliver measurable improvements in production.

Analyze workload behavior, capacity requirements, utilization, tail latency, and cost efficiency across heterogeneous GPU providers and hardware.

Help ensure that customer performance requirements are met without consuming unnecessary infrastructure capacity.

Production Feedback Loop

Identify recurring patterns across incidents, workloads, and customer escalations.

Translate those findings into improvements to the inference platform, reliability architecture, deployment processes, observability, and product roadmap.

Technical Abilities Production Systems

Deep experience operating critical, customer-facing or business-critical production systems.

Ability to reason about service health across multiple layers rather than treating infrastructure availability as the complete customer outcome.

Reliability Engineering

Experience defining and operating service-level indicators and objectives, building actionable observability, leading incidents, performing failure analysis, and reducing mean time to detection and recovery.

Performance Diagnosis

Strong understanding of latency, throughput, queueing, resource contention, capacity, workload distribution, and tail-performance behavior.

Demonstrated ability to diagnose difficult production regressions and isolate bottlenecks across applications and infrastructure.

Distributed Systems

Knowledge of distributed-systems principles, including fault tolerance, scheduling, routing, load balancing, capacity management, consistency, and failure recovery.

Cloud and Infrastructure

Strong experience with Kubernetes, Linux, networking, storage, cloud infrastructure, and containerized production environments.

Experience operating across multiple cloud providers, regions, hardware configurations, or infrastructure suppliers is especially valuable.

Software Engineering

Ability to write production-quality software and build internal tooling, instrumentation, automation, and…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary