AI Platform Engineer
Job in
New York, New York County, New York, 10261, USA
Listed on 2026-07-18
Listing for:
Spencer Duncan
Full Time
position Listed on 2026-07-18
Job specializations:
-
Software Development
AI Engineer (Applied/Software), DevOps, Cloud Engineer - Software
Job Description & How to Apply Below
AI Platform Engineer Position Overview
We are seeking an experienced AI Platform Engineer to build and operate the foundational infrastructure that powers AI-driven game development, live-service operations, player experiences, and studio-wide AI initiatives. This role focuses on creating scalable AI platforms, deployment pipelines, model serving infrastructure, monitoring systems, and developer tooling that enable game teams to efficiently build, deploy, and manage AI-powered applications.
Key Responsibilities AI Platform Architecture- Design, build, and maintain enterprise-scale AI platforms that support model development, deployment, monitoring, evaluation, and lifecycle management.
- Develop self-service infrastructure and tooling that enable game studios to rapidly deploy and manage AI-powered applications.
- Establish platform standards, architecture patterns, and operational best practices for AI workloads.
- Build scalable multi-tenant environments supporting multiple game projects and development teams.
- Drive continuous improvements in platform reliability, performance, security, and developer productivity.
- Design and implement model deployment pipelines supporting machine learning models, Large Language Models (LLMs), recommendation systems, and AI agents.
- Build scalable model serving infrastructure capable of supporting real-time and batch inference workloads.
- Develop automated deployment workflows, rollback mechanisms, and version management systems.
- Implement canary deployments, blue-green deployments, and progressive rollout strategies for AI services.
- Optimize model serving performance, latency, throughput, and resource utilization across production environments.
- Establish end-to-end MLOps workflows covering training, testing, deployment, monitoring, retraining, and governance.
- Develop automated CI/CD pipelines for machine learning and AI applications.
- Implement model registry solutions, artifact management systems, and reproducible deployment processes.
- Support experiment tracking, feature management, and model lifecycle automation.
- Collaborate with AI teams to improve development workflows and reduce time-to-production for AI solutions.
- Design and implement comprehensive monitoring systems for AI models, inference services, and platform infrastructure.
- Track model accuracy, latency, throughput, drift, hallucination rates, retrieval quality, and business performance metrics.
- Develop evaluation frameworks for Large Language Models, retrieval systems, recommendation engines, and AI agents.
- Create dashboards, alerting systems, and operational analytics supporting proactive issue detection.
- Perform root-cause analysis and implement improvements to increase platform reliability and service quality.
- Develop APIs, SDKs, and service layers that expose AI capabilities to game clients, backend systems, and studio development teams.
- Build reusable platform services supporting embeddings, vector search, retrieval pipelines, inference routing, and model orchestration.
- Create developer tools that simplify integration of AI services into game applications.
- Design authentication, authorization, rate limiting, and service governance mechanisms.
- Ensure platform APIs are scalable, secure, and easy to adopt across engineering teams.
- Build and manage cloud-native infrastructure across AWS, Google Cloud Platform (GCP), and/or Microsoft Azure.
- Design highly available, fault-tolerant systems supporting mission-critical AI workloads.
- Implement Infrastructure as Code (IaC) using Terraform and related automation tools.
- Manage Kubernetes clusters, container orchestration platforms, networking, and service meshes.
- Optimize infrastructure costs while maintaining performance and operational reliability.
- Implement security controls protecting AI services, model assets, datasets, and infrastructure resources.
- Develop governance frameworks supporting responsible AI deployment and operational compliance.
- Establish access management, secrets management, encryption, and auditing mechanisms.
- Ensure platform architectures comply with organizational security policies and industry best practices.
- Conduct infrastructure reviews and risk assessments for AI systems operating in production.
- Partner with AI engineers, machine learning teams, game developers, Dev Ops engineers, and architects to deliver platform capabilities.
- Support game studios in adopting AI technologies and deploying AI-powered services.
- Participate in architecture reviews, technical planning sessions, and platform roadmap discussions.
- Mentor engineers on cloud-native development, MLOps practices, and AI infrastructure technologies.
- Stay current with advancements in AI platforms, cloud infrastructure, and distributed systems.
- Bachelor's or Master's degree in…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×