×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer

Job in Burbank, Los Angeles County, California, 91520, USA
Listing for: Paramount Pictures
Full Time position
Listed on 2026-07-24
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 130200 - 195300 USD Yearly USD 130200.00 195300.00 YEAR
Job Description & How to Apply Below

Select how often (in days) to receive an alert:

#We Are Paramount  on a mission to unleash the power of content… you in?

We’ve got the brands, we’ve got the stars, we’ve got the power to achieve our mission to entertain the planet – now all we’re missing is… YOU! Becoming a part of Paramount means joining a team of passionate people who not only recognize the power of content but also enjoy a touch of fun and uniqueness. Together, we co-create moments that matter – both for our audiences and our employees – and aim to leave a positive mark on culture.

Overview

We are seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure. This is a high-leverage role designed for an expert who understands the deep technical stack required to shift ML models from research to global production. You will be responsible for the "engine room" of the AMLG, ensuring that our MLEs can train massive models efficiently and serve them with sub-millisecond reliability.

This role requires a unique blend of expertise in distributed systems and hardware acceleration. You will lead the adoption and optimization of Any Scale (Ray) for distributed training and manage a high-performance Kubernetes-based inference environment. You aren't just managing clusters; you are building a seamless, scalable platform that abstracts the complexity of GPUs and distributed compute for the entire organization.

Why

This Role Matters

The ML Platform Lead is the force-multiplier for every other ML pod. In this role, you will directly shape:

  • The Training Foundation:
    Establishing Any Scale/Ray as the standard for distributed compute, enabling MLEs to train models on petabytes of data without managing infrastructure.
  • Inference at Scale:
    Architecting the serving layer that handles billions of requests per day, optimizing for both p99 latency and GPU utilization.
  • Operational Excellence:
    Setting the organizational standards for how ML models are deployed, monitored, and scaled across the enterprise.
Key Responsibilities
  • Technical Roadmap & Strategy:
    Own the long-term architectural direction for the Training and Inference domains, ensuring the platform scales 10x over a 1–3 year horizon.
  • Distributed Training Leadership:
    Lead the implementation and optimization of Ray/Any Scale, providing a unified compute layer for batch processing, model training, and reinforcement learning.
  • High-Performance Inference:
    Design and maintain K8s-based inference servers (e.g., Triton, Torch Serve, or vLLM) optimized for GPU memory management and high throughput.
  • Hardware & Cost Optimization:
    Navigate the trade-offs between different GPU instances (A100s, H100s, T4s), optimizing for cost, availability, and performance.
  • Cross-Team Standardization:
    Solve high-leverage problems that affect multiple pods (e.g., Entry, Session, Presentation), establishing reusable patterns for CI/CD, model versioning, and canary deployments.
  • Reliability Engineering:
    Define and enforce SLIs/SLOs for the platform, ensuring that infrastructure failures never interrupt the user-facing personalization experience.
  • Mentorship &

    Coaching:

    Act as a technical mentor to senior engineers across the ML Platform and Applied ML pods, raising the bar for system design and operational rigor.
Basic Qualifications
  • 6-8+ years of experience in ML Infrastructure, Platform Engineering, or high-scale Backend Engineering.
  • Orchestration & Serving:
    Extensive experience with Kubernetes (K8s) and serving frameworks for large-scale ML models.
  • Hardware Proficiency:
    Strong knowledge of GPU architecture, CUDA, and optimizing ML workloads for hardware acceleration.
  • Leadership (IC4/5):
    Proven track record of owning the technical direction for a major domain and driving impact across multiple teams.
Preferred Qualifications
  • Experience with Infra-as-Code (Terraform/Pulumi) and building automated MLOps pipelines.
  • Distributed Systems Mastery:
    Deep expertise with Ray (Any Scale) or similar distributed compute frameworks.
  • Familiarity with ML observability tools (Prometheus, Grafana, Weights & Biases, or MLFlow).
  • Experience managing multi-cloud or hybrid-cloud ML…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary