×
Register Here to Apply for Jobs or Post Jobs. X

Platform Engineer (AI​/LLM Infrastructure

Job in Santa Clara, Santa Clara County, California, 95050, USA
Listing for: United Software Group
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Infrastructure
Job Description & How to Apply Below
Position: Platform Engineer (AI/LLM Infrastructure)

Platform Engineer (AI/LLM Infrastructure)

Day to Day

Job Duties:

  • Lead the design, implementation, and operation of scalable infrastructure platforms supporting AI/LLM-based solutions for enterprise clients
  • Act as a hands-on technical lead (player-coach), contributing to development while guiding a team of engineers
  • Own end-to-end infrastructure architecture below the application layer, including compute, container orchestration, CI/CD, observability, and security
  • Partner directly with clients and stakeholders to design, present, and deliver robust AI infrastructure solutions
  • Architect and manage production-grade Kubernetes environments (AKS/EKS), including cluster operations and RBAC
  • Design and operationalize RAG pipelines, including ingestion, chunking, embedding workflows, and vector database management
  • Lead GPU infrastructure provisioning and optimization (NVIDIA A100/H100 or similar)
  • Drive Infrastructure-as-Code adoption using Terraform and Git Ops (ArgoCD/Flux)
  • Build and maintain CI/CD pipelines using Git Hub Actions and Azure Dev Ops
  • Establish observability standards using Datadog, Open Telemetry, and ELK/Open Search
  • Lead incident response, on-call processes, and post-mortem analysis
  • Ensure strong security posture and lead Info Sec review processes
  • Coordinate delivery across multiple teams and client engagements

Basic Qualifications:

  • 5–8 years of experience in Platform Engineering, SRE, or Infrastructure Engineering
  • 3+ years of proven experience delivering and leading infrastructure for AI/LLM-based production systems
  • Strong hands-on expertise in Kubernetes, Docker, Helm
  • 3+ years of experience with Terraform and Git Ops (ArgoCD/Flux)
  • 3+ years of experience with Azure (Key Vault, Monitor, Dev Ops Pipelines)
  • 3+ years of experience leading client-facing technical engagements
  • 3+ years of experience managing multiple concurrent projects or teams
  • 3+ years of hands-on experience with incident management and SLA-driven environments
  • 3+ years of experience leading security/Info Sec reviews
  • Strong understanding of vector databases, RAG pipelines, and LLM inference systems
  • 3+ years of experience with CI/CD and container registry management
  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience

Nice to Have (But Not Required):

  • Experience with AWS in addition to Azure
  • Familiarity with Azure API Management and AKS
  • Experience with Pulumi (Python/Type Script)
  • Knowledge of NIM deployment and lifecycle management
  • Python scripting for infrastructure automation
  • Experience with load testing tools (k6, Locust, JMeter)
  • Exposure to Fin Ops and cost optimization practices

Location:

Santa Clara, CA (3 days onsite in a week)

Duration: 13+ months

USA
• Canada
• Mexico Costa Rica
• India - Thanks & Regards, Sudheer Senior US IT Recruiter | United Software Group Inc.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary