Inference Lead
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-09-07
Listing for:
Compunnel, Inc.
Full Time
position Listed on 2026-09-07
Job specializations:
-
Software Development
Job Description & How to Apply Below
The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments. This position requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, production operations, CI/CD, and high-availability engineering, along with the ability to provide technical leadership and mentorship to inference and platform engineering teams.
Key Responsibilities- Architect low-latency online inference and real-time model-serving solutions.
- Develop scalable APIs, microservices, and deployment patterns for predictive models.
- Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
- Conduct benchmarking, performance tuning, capacity planning, and load testing.
- Optimize latency, throughput, resource consumption, availability, and operational cost.
- Define monitoring, alerting, Service Level Objectives (SLOs), runbooks, and incident-response practices.
- Build CI/CD pipelines to support repeatable model and service releases.
- Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
- Lead technical reviews and provide mentorship to inference and platform engineers.
- Establish operational and performance standards for real-time model-serving services across cloud and on-premises environments.
- 8–12 years of professional experience in software engineering, inference engineering, platform engineering, ML engineering, SRE, or related technical disciplines.
- Strong experience with online inference and real-time model-serving architecture.
- Hands-on experience designing REST and/or gRPC APIs and distributed microservices.
- Strong experience with Kubernetes, containers, autoscaling, and traffic management.
- Experience with performance engineering, latency optimization, benchmarking, and load testing.
- Experience with monitoring, SLOs, capacity planning, and production operations.
- Strong experience building CI/CD pipelines and implementing progressive-deployment approaches.
- Experience designing resilient, highly available systems, including failover, rollback, disaster recovery, and graceful degradation.
- Experience deploying and operating solutions across both cloud and on-premises environments.
- Strong technical leadership skills with experience conducting technical reviews and mentoring engineering teams.
- Degree in Computer Science, Engineering, or a related discipline.
- Experience with enterprise model-serving platforms and inference runtimes.
- Cloud, Kubernetes, SRE, or ML engineering certification.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×