×
Register Here to Apply for Jobs or Post Jobs. X

Inference Lead

Job in Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listing for: Compunnel, Inc.
Full Time position
Listed on 2026-09-07
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 170000 - 210000 USD Yearly USD 170000.00 210000.00 YEAR
Job Description & How to Apply Below

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments. This position requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, production operations, CI/CD, and high-availability engineering, along with the ability to provide technical leadership and mentorship to inference and platform engineering teams.

Key Responsibilities
  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and operational cost.
  • Define monitoring, alerting, Service Level Objectives (SLOs), runbooks, and incident-response practices.
  • Build CI/CD pipelines to support repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and provide mentorship to inference and platform engineers.
  • Establish operational and performance standards for real-time model-serving services across cloud and on-premises environments.
Required Qualifications
  • 8–12 years of professional experience in software engineering, inference engineering, platform engineering, ML engineering, SRE, or related technical disciplines.
  • Strong experience with online inference and real-time model-serving architecture.
  • Hands-on experience designing REST and/or gRPC APIs and distributed microservices.
  • Strong experience with Kubernetes, containers, autoscaling, and traffic management.
  • Experience with performance engineering, latency optimization, benchmarking, and load testing.
  • Experience with monitoring, SLOs, capacity planning, and production operations.
  • Strong experience building CI/CD pipelines and implementing progressive-deployment approaches.
  • Experience designing resilient, highly available systems, including failover, rollback, disaster recovery, and graceful degradation.
  • Experience deploying and operating solutions across both cloud and on-premises environments.
  • Strong technical leadership skills with experience conducting technical reviews and mentoring engineering teams.
Preferred Qualifications
  • Degree in Computer Science, Engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary