AI Platform Engineer
Listed on 2026-09-30
-
Software Development
AI Engineer (Applied/Software), Cloud Engineer - Software, Machine Learning/ ML Engineer, DevOps
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job TitleAI Platform Engineer
Location: 100% Remote (Continental United States)
Position Type: Full-time, Direct W2
Salary Range: $130,000–$180,000 Annually (based on experience)
Experience
Required:
10+ Years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job SummaryBright Vision Technologies is seeking a highly experienced AI Platform Engineer with 10+ years of experience in distributed systems, cloud-native infrastructure, and AI platform engineering to design, build, and operate enterprise-scale AI inference and machine learning platforms. The ideal candidate will possess deep expertise in LLM serving, GPU optimization, Kubernetes, cloud infrastructure, distributed systems, and MLOps , with a proven ability to deliver highly scalable, reliable, secure, and cost-efficient AI platforms supporting production machine learning workloads.
Key Responsibilities- Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
- Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
- Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
- Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
- Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
- Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
- Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
- Collaborate with AI researchers, ML engineers, Dev Ops teams, and software engineers to deploy and support production AI models.
- Drive cloud infrastructure optimization, resource utilization, Fin Ops initiatives, and operational excellence.
- Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development.
- Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
- Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
- Extensive experience with Large Language Model (LLM) serving
, model inference optimization, and production AI infrastructure. - Hands‑on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks
. - Strong expertise in Kubernetes
, container orchestration, Docker, and cloud-native application architectures. - Experience optimizing GPU workloads using CUDA
, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure. - Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).