ML Engineer, Inference & Optimization
Job in
Palo Alto, Santa Clara County, California, 94302, USA
Listed on 2026-10-02
Listing for:
Pika
Full Time
position Listed on 2026-10-02
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below
Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
What You’ll DoAccelerate Inference:
Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
Maximize GPU Parallelism:
Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
Programming for Performance:
Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.Advance AI Deployment:
Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
Technical Excellence:
Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
What We’re Looking For
Experience:
5+ years engineering experience, with a strong track record in inference acceleration and model deployment erence Mastery:
Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
GPU & Parallelism:
Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
AI Domain Knowledge:
Familiarity with video generation (videogen) models and large language models (LLMs).
Collaboration:
Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
Ownership Mindset:
Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
Bonus:
Experience in enhancing training efficiency, stability, or resource optimization for large models.
Nice to Have
Experience with high-throughput video or real-time streaming model deployment
Familiarity with distributed training and optimization toolkits
Contributions to open source projects in AI infrastructure or deep learning compilers
Startup or rapid prototyping experience
What We Offer Competitive salary in the AI industry
Equity in a fast-growing startup shaping the future of AI Comprehensive health benefits, monthly stipends, company retreatsA supportive and collaborative office culture—we’re all building and launching together
About Pika At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.
We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.
Compensation Range: $250K - $350K Location Palo Alto HQ
Employment Type
Full time Location Type On-site
Department Research Compensation base salary $250K – $350K
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×