×
Register Here to Apply for Jobs or Post Jobs. X

Inference Engineering Lead

Job in Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listing for: Intone Networks Inc.
Full Time position
Listed on 2026-08-31
Job specializations:
  • Software Development
    DevOps, Backend Developer, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 120 - 180 USD Hourly USD 120.00 180.00 HOUR
Job Description & How to Apply Below
Position: Real-Time Inference Engineering Lead

We are seeking a Real-Time Inference Engineering Lead for a contract position based in Charlotte, North Carolina. This role offers the opportunity to lead technical design and delivery for highly available, production-grade real-time inference systems over a six-month or longer engagement.

Responsibilities
  • Lead technical design and engineering delivery for production inference services with high availability and performance requirements
  • Design and operate online inference architectures, model-serving frameworks, and predictive-model deployment patterns
  • Design and operate RESTful, gRPC, or event-driven APIs supporting real-time inference workloads
  • Implement and optimize autoscaling, resource management, capacity planning, performance testing, and load testing strategies for distributed services
  • Establish and maintain observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident-management practices
  • Ensure resiliency, high availability, fault tolerance, and disaster recovery for critical services
  • Collaborate with data science, ML engineering, platform engineering, application, security, and business teams to align technical solutions with organizational goals
  • Oversee CI/CD processes, Git-based development practices, automated testing, deployment automation, and production-release procedures
Qualifications
  • Required
    • 8+ years of software engineering, platform engineering, cloud engineering, SRE, or infrastructure engineering experience
    • 4+ years of experience designing, building, or operating production APIs, distributed systems, platform services, or real‑time data and ML workloads
    • 4+ years of hands‑on experience with Kubernetes and container platforms in production, including GKE, Open Shift, or comparable environments
    • Demonstrated experience leading technical design and engineering delivery for highly available, performance‑sensitive production services
    • Strong experience with online inference architecture, model‑serving frameworks, or predictive‑model deployment patterns
    • Hands‑on experience designing and operating RESTful, gRPC, or event‑driven APIs
    • Experience with autoscaling, resource management, capacity planning, performance testing, and load testing for distributed services
    • Experience implementing observability, monitoring, dashboards, alerts, SLIs, SLOs, and incident‑management practices
    • Experience with CI/CD, Git‑based development, automated testing, deployment automation, and production‑release practices
    • Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical services
    • Ability to work effectively with data science, ML engineering, platform engineering, application teams, security, and business stakeholders
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary