More jobs:
Inference Engineering Lead
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-08-31
Listing for:
Intone Networks Inc.
Full Time
position Listed on 2026-08-31
Job specializations:
-
Software Development
DevOps, Backend Developer, Cloud Engineer - Software
Job Description & How to Apply Below
We are seeking a Real-Time Inference Engineering Lead for a contract position based in Charlotte, North Carolina. This role offers the opportunity to lead technical design and delivery for highly available, production-grade real-time inference systems over a six-month or longer engagement.
Responsibilities- Lead technical design and engineering delivery for production inference services with high availability and performance requirements
- Design and operate online inference architectures, model-serving frameworks, and predictive-model deployment patterns
- Design and operate RESTful, gRPC, or event-driven APIs supporting real-time inference workloads
- Implement and optimize autoscaling, resource management, capacity planning, performance testing, and load testing strategies for distributed services
- Establish and maintain observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices
- Ensure resiliency, high availability, fault tolerance, and disaster recovery for critical services
- Collaborate with data science, ML engineering, platform engineering, application, security, and business teams to align technical solutions with organizational goals
- Oversee CI/CD processes, Git-based development practices, automated testing, deployment automation, and production-release procedures
- Required
- 8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience
- 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real‑time data and ML workloads
- 4+ years of hands‑on experience with Kubernetes and container platforms in production, including GKE, Open Shift, or comparable environments
- Demonstrated experience leading technical design and engineering delivery for highly available, performance‑sensitive production services
- Strong experience with online inference architecture, model‑serving frameworks, or predictive‑model deployment patterns
- Hands‑on experience designing and operating RESTful, gRPC, or event‑driven APIs
- Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services
- Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident‑management practices
- Experience with CI/CD, Git‑based development, automated testing, deployment automation, and production‑release practices
- Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services
- Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×