Manager - AI Inference Engineer
Job in
San Antonio, Bexar County, Texas, 78208, USA
Listed on 2026-08-03
Listing for:
EY
Full Time
position Listed on 2026-08-03
Job specializations:
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Location:
Anywhere in Country
. This role is focused on the engineering of a production-grade inference platform: architecture, serving stack design, benchmarking, integration, reliability, and operational readiness.
Your
Key Responsibilities Private AI inference architecture
- Design and deploy secure private inference architectures for enterprise AI inference.
- Translate workload requirements into technical architecture decisions across compute, caching, interconnect, storage & networking.
- Define reference architectures that support high-throughput, low-latency, enterprise inference at scale.
- Design the model-serving stack, including inference runtime, orchestration, model registry, artifact management, API layer, routing, observability, deployment and fallback.
- Evaluate and recommend inference technologies and platform components for production use.
- Configure and harden the target environment for secure, reliable inference workloads.
- Build and run benchmark suites for enterprise workloads.
- Measure latency, throughput, concurrency, utilization, reliability, and cost efficiency.
- Tune serving configurations and system parameters to improve production performance.
- Produce evaluation results and recommendations based on objective testing.
- Support integration of inference platforms with client IT systems.
- Work with infrastructure and security stakeholders to validate technical requirements and remediate issues.
- Onboard models into the target inference platform.
- Configure serving endpoints, routing behavior, monitoring, and operational controls.
- Support deployment readiness, runbooks, and operational handoff for production use.
- Ability to turn business and workload needs into concrete infrastructure and software design decisions.
- Clear technical communication and strong architecture documentation skills.
- Bachelor’s Degree in a relevant field
- Strong experience designing and operating AI inference systems in production.
- Deep familiarity with multi-GPU serving and high-performance deployment patterns.
- Hands‑on experience with Kubernetes and cloud-native platform technologies.
- Strong understanding of scaling, routing, autoscaling, observability, and reliability for production AI workloads.
- Experience integrating AI platforms into enterprise security and access environments.
- Experience with private, sovereign, or enterprise AI infrastructure.
- Experience benchmarking large-model inference on GPU clusters.
- Familiarity with service mesh, policy enforcement, secrets management, and RBAC.
- Experience working across infrastructure, security, and application teams in a regulated environment.
- Experience supporting pilot-to-production AI platform rollouts.
- We offer a comprehensive compensation and benefits package where you’ll be rewarded based on your performance and recognized for the value you bring to the business. The base salary range for this job in all geographic locations in the US is $125,500 to $230,200. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $150,700 to $261,600.
Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography. In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options. - Join us in our team-led and leader-enabled hybrid model. Our expectation is for most people in external, client serving roles to work together in person…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×