×
Register Here to Apply for Jobs or Post Jobs. X

Inference Lead

Job in Charlotte, Mecklenburg County, North Carolina, 28202, USA
Listing for: Kasmo Global
Full Time position
Listed on 2026-09-07
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Description & How to Apply Below

Real-Time Inference Engineering Lead

Hybrid Onsite - local candidates preferred. Real-Time Services – Real-Time Inference Engineering Lead

Experience:

8–12 years

Role

Summary:

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.

Key Responsibilities:

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.

Required Skills:

  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience.

Preferred Qualifications:

  • Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary