AI Platform Engineer
Listed on 2026-07-26
-
Software Development
AI Engineer (Applied/Software), Cloud Engineer - Software, DevOps, Backend Developer
Overview
A career with us means you’ll work alongside exceptional people and be empowered to reach your professional and personal goals. Our employees are at the foundation of what enables Mass Mutual to deliver on our purpose to help people secure their futures and protect the ones they love.
We embrace the idea that we all are stronger and better through our support for one another. We strive to create a culture where employees feel valued and are celebrated for who they are.
Job DescriptionAI Platform Engineer | AI Platform Engineering
Full-Time Hybrid Onsite (3 days/week in office)
The OpportunityMass Mutual’s AI Platform Engineering team is seeking a skilled AI Platform Engineer to contribute to the design, development, and operation of our growing AI platform. You will work on meaningful platform challenges, collaborate closely with senior engineers, and take ownership of platform components that power Mass Mutual’s AI initiatives.
The TeamThis is a unique opportunity to work on the team that builds and operates the platform powering Mass Mutual’s AI initiatives. The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience—delivering foundational capabilities that shape how the entire organization builds and deploys AI. We partner closely with AI engineering, product, and cloud engineering teams across the enterprise, and we invest in growth through a culture of peer learning, candid feedback, and shared technical standards.
This team is defined by a shared commitment to engineering excellence, clear documentation, and a drive to make hard problems tractable.
- Contribute to the design and implementation of platform components—cloud infrastructure, AI serving layers, and developer tooling, with guidance from senior engineers.
- Take ownership of discrete features or modules within critical platform systems such as the LLM gateway, model serving infrastructure, and enterprise integration patterns. Participate in writing ADRs and technical documentation.
- Participate in design reviews, contribute constructive feedback on pull requests, and collaborate with teammates on complex engineering challenges.
- Execute platform initiatives end to end within your scope—breaking down tasks, managing your work, and delivering quality outcomes in production.
- Support platform reliability efforts: contribute to Service Level Objective definitions, implement observability instrumentation, participate in incident response, and help improve platform stability.
- Implement governance and compliance controls—data residency, access management, audit logging—in alignment with enterprise requirements and established patterns.
- Collaborate with AI engineering, product, and cloud engineering teams; communicate technical context and trade‑offs clearly in cross‑functional settings.
- Contribute to team documentation habits, code quality, and shared engineering standards.
Minimum Qualifications
- 2+ years of experience in platform, infrastructure, or SRE roles, with demonstrated ownership of platform components or services.
- 2+ years of experience with cloud-native architecture across AWS, GCP, or Azure—containerized deployments, managed services, and basic multi‑tenancy patterns.
- 2+ years of experience delivering platform features from development through production, with the ability to manage ambiguity within a defined scope.
- 2+ years of experience working with IaC and Git Ops tools:
Terraform or Pulumi, ArgoCD, and standard deployment patterns.
- Familiarity with Kubernetes (CKA, CKAD, or equivalent AWS certifications a plus); working knowledge of cloud‑native concepts including managed services, networking, and identity.
- Exposure to AI/ML infrastructure: model serving, inference pipelines, or LLM integration patterns.
- Clear written communication: ability to produce useful technical documentation and explain implementation decisions to teammates.
- Hands‑on experience with LLM serving frameworks—vLLM, Triton, Ray Serve—or familiarity with AI gateway patterns.
- Exposure to Fin Ops principles or GPU cost considerations in inference workloads.
- Experience contributing to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).