Member of Technical Staff, AI Training Infrastructure
Listed on 2026-07-18
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Member of Technical Staff, AI Training Infrastructure
San Mateo, CA
About UsFireworks is building the future of generative AI infrastructure. Our platform delivers high-quality models with scalable inference and fast performance. We are a Series C company valued at $4 billion, backed by investors including Benchmark, Sequoia, Lightspeed, Index, and Evantic. Our team includes veterans of Meta PyTorch and Google Vertex AI.
The RoleAs a Training Infrastructure Engineer, you ll design, build, and optimize the infrastructure that powers our large-scale model training operations. You will collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development.
Key Responsibilities- Design and implement scalable infrastructure for large-scale model training workloads
- Develop and maintain distributed training pipelines for LLMs and multimodal models
- Optimize training performance across multiple GPUs, nodes, and data centers
- Implement monitoring, logging, and debugging tools for training operations
- Architect and maintain data storage solutions for large-scale training datasets
- Automate infrastructure provisioning, scaling, and orchestration for model training
- Collaborate with researchers to implement and optimize training methodologies
- Analyze and improve efficiency, scalability, and cost-effectiveness of training systems
- Troubleshoot complex performance issues in distributed training environments
- Bachelor s degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience
- 3+ years of experience with distributed systems and ML infrastructure
- Experience with Py Torch
- Proficiency in cloud platforms (AWS, GCP, Azure)
- Experience with containerization, orchestration (Kubernetes, Docker)
- Knowledge of distributed training techniques (data parallelism, model parallelism, FSDP)
- Master s or PhD in Computer Science or related field
- Experience training large language models or multimodal AI systems
- Experience with ML workflow orchestration tools
- Background in optimizing high-performance distributed computing systems
- Familiarity with ML Dev Ops practices
- Contributions to open-source ML infrastructure or related projects
Total compensation for this role also includes meaningful equity in a fast-growing startup, along with a competitive salary and comprehensive benefits package. Base salary is determined by a range of factors including individual qualifications, experience, skills, interview performance, market data, and work location. The listed salary range is intended as a guideline and may be adjusted.
$175,000 - $220,000 USD
Why Fireworks AI?- Solve Hard Problems:
Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving. - Build What Next:
Work with bleeding-edge technology that impacts how businesses and developers harness AI globally. - Ownership & Impact:
Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results. - Learn from the Best:
Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).