AI Platform Engineer
Listed on 2026-07-27
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software), Systems Engineer
Are you passionate about cloud platforms, AI infrastructure, Kubernetes, and building the foundational systems that enable enterprise-scale AI innovation?
We are seeking an AI Platform Engineer to help design, build, and operate the platform powering enterprise AI capabilities. This role sits at the intersection of cloud engineering, platform engineering, AI/ML infrastructure, and developer experience, enabling teams to build, deploy, and scale AI solutions efficiently and securely.
The ideal candidate combines strong cloud-native engineering skills, infrastructure automation expertise, and a passion for creating reliable, scalable platforms that support AI applications, model serving, and large language model (LLM) workloads.
Key Responsibilities
- Design, develop, and maintain platform capabilities that support enterprise AI initiatives.
- Build foundational services for:
- AI Applications
- Model Serving
- Inference Workloads
- LLM Integrations
- Developer Self-Service
- Support the evolution of the enterprise AI platform architecture.
- Deliver scalable and reusable infrastructure patterns.
- Develop and manage cloud-native infrastructure across AWS, Azure, or GCP environments.
- Build and maintain secure, scalable, and resilient platform services.
- Support containerized deployments and platform automation.
- Implement modern infrastructure patterns that enable efficient AI deployment and operations.
AI/ML Infrastructure
- Contribute to model serving and inference infrastructure.
- Support deployment and operationalization of AI and machine learning workloads.
- Assist with enterprise LLM integration patterns and AI gateway capabilities.
- Collaborate with AI engineering teams to optimize platform capabilities for model deployment and execution.
Infrastructure as Code & Git Ops
- Develop and manage infrastructure using:
- Terraform
- Git Ops Practices
- ArgoCD
- Automate infrastructure provisioning and deployment workflows.
- Ensure environments remain consistent, secure, and reproducible.
- Improve deployment efficiency through automation and standardization.
Platform Reliability & Operations
- Support platform reliability initiatives and operational excellence activities.
- Contribute to:
- Monitoring & Alerting
- Availability Improvements
- Performance Optimization
- Implement observability standards and telemetry solutions.
- Participate in troubleshooting and root-cause analysis activities.
Governance, Security & Compliance
- Implement security and compliance controls across platform services.
- Support:
- Identity & Access Management
- Audit Logging
- Data Residency Requirements
- Infrastructure Security Controls
- Ensure platform solutions align with enterprise governance standards.
Technical Design & Documentation
- Participate in architecture discussions, design reviews, and technical planning activities.
- Create and maintain:
- Technical Documentation
- Platform Runbooks
- Communicate technical trade-offs and design decisions clearly.
Collaboration & Cross-Functional Partnership
- Partner with:
- Product Teams
- Platform Teams
- Security Teams
- Translate business and technical requirements into platform capabilities.
- Collaborate on enterprise AI initiatives and platform improvements.
- Support shared engineering standards and best practices.
- Improve internal developer experience through platform automation and self-service tooling.
- Contribute to internal developer platform (IDP) initiatives.
- Promote engineering excellence, code quality, and knowledge sharing.
- Continuously evaluate emerging technologies and platform innovations.
Qualifications
Education
- Bachelor's Degree in:
- Computer Science
- Software Engineering
- Information Technology
- Mathematics
- Related Technical Discipline
Required Experience
- 2+ years of experience in:
- Site Reliability Engineering (SRE)
- Experience owning and operating production platform services.
- Experience building and deploying cloud-native applications.
- Experience delivering solutions from development through production support.
Required Technical Skills
- AWS, Azure, or Google Cloud Platform (GCP)
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).