×
Register Here to Apply for Jobs or Post Jobs. X

AI Platform Engineer

Job in Washington, District of Columbia, 20022, USA
Listing for: Bright Vision Technologies
Full Time position
Listed on 2026-09-30
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Cloud Engineer - Software, Machine Learning/ ML Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 130000 - 180000 USD Yearly USD 130000.00 180000.00 YEAR
Job Description & How to Apply Below

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title

AI Platform Engineer

Location: 100% Remote (Continental United States)
Position Type: Full-time, Direct W2
Salary Range: $130,000–$180,000 Annually (based on experience)
Experience

Required:


10+ Years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

Bright Vision Technologies is seeking a highly experienced AI Platform Engineer with 10+ years of experience in distributed systems, cloud-native infrastructure, and AI platform engineering to design, build, and operate enterprise-scale AI inference and machine learning platforms. The ideal candidate will possess deep expertise in LLM serving, GPU optimization, Kubernetes, cloud infrastructure, distributed systems, and MLOps , with a proven ability to deliver highly scalable, reliable, secure, and cost-efficient AI platforms supporting production machine learning workloads.

Key Responsibilities
  • Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
  • Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
  • Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
  • Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
  • Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
  • Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
  • Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
  • Collaborate with AI researchers, ML engineers, Dev Ops teams, and software engineers to deploy and support production AI models.
  • Drive cloud infrastructure optimization, resource utilization, Fin Ops initiatives, and operational excellence.
  • Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development.
  • Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
  • 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
  • Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
  • Extensive experience with Large Language Model (LLM) serving
    , model inference optimization, and production AI infrastructure.
  • Hands‑on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks
    .
  • Strong expertise in Kubernetes
    , container orchestration, Docker, and cloud-native application architectures.
  • Experience optimizing GPU workloads using CUDA
    , NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure.
  • Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Strong understanding of distributed systems, networking,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary