More jobs:
Platform Engineer (AI/LLM Infrastructure
Job in
Santa Clara, Santa Clara County, California, 95050, USA
Listed on 2026-08-05
Listing for:
United Software Group
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Infrastructure
Job Description & How to Apply Below
Platform Engineer (AI/LLM Infrastructure)
Day to Day
Job Duties:
- Lead the design, implementation, and operation of scalable infrastructure platforms supporting AI/LLM-based solutions for enterprise clients
- Act as a hands-on technical lead (player-coach), contributing to development while guiding a team of engineers
- Own end-to-end infrastructure architecture below the application layer, including compute, container orchestration, CI/CD, observability, and security
- Partner directly with clients and stakeholders to design, present, and deliver robust AI infrastructure solutions
- Architect and manage production-grade Kubernetes environments (AKS/EKS), including cluster operations and RBAC
- Design and operationalize RAG pipelines, including ingestion, chunking, embedding workflows, and vector database management
- Lead GPU infrastructure provisioning and optimization (NVIDIA A100/H100 or similar)
- Drive Infrastructure-as-Code adoption using Terraform and Git Ops (ArgoCD/Flux)
- Build and maintain CI/CD pipelines using Git Hub Actions and Azure Dev Ops
- Establish observability standards using Datadog, Open Telemetry, and ELK/Open Search
- Lead incident response, on-call processes, and post-mortem analysis
- Ensure strong security posture and lead Info Sec review processes
- Coordinate delivery across multiple teams and client engagements
Basic Qualifications:
- 5–8 years of experience in Platform Engineering, SRE, or Infrastructure Engineering
- 3+ years of proven experience delivering and leading infrastructure for AI/LLM-based production systems
- Strong hands-on expertise in Kubernetes, Docker, Helm
- 3+ years of experience with Terraform and Git Ops (ArgoCD/Flux)
- 3+ years of experience with Azure (Key Vault, Monitor, Dev Ops Pipelines)
- 3+ years of experience leading client-facing technical engagements
- 3+ years of experience managing multiple concurrent projects or teams
- 3+ years of hands-on experience with incident management and SLA-driven environments
- 3+ years of experience leading security/Info Sec reviews
- Strong understanding of vector databases, RAG pipelines, and LLM inference systems
- 3+ years of experience with CI/CD and container registry management
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience
Nice to Have (But Not Required):
- Experience with AWS in addition to Azure
- Familiarity with Azure API Management and AKS
- Experience with Pulumi (Python/Type Script)
- Knowledge of NIM deployment and lifecycle management
- Python scripting for infrastructure automation
- Experience with load testing tools (k6, Locust, JMeter)
- Exposure to Fin Ops and cost optimization practices
Location:
Santa Clara, CA (3 days onsite in a week)
Duration: 13+ months
USA
• Canada
• Mexico Costa Rica
• India - Thanks & Regards, Sudheer Senior US IT Recruiter | United Software Group Inc.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×