Machine Learning Engineer – World Model
Listed on 2026-08-02
-
Software Development
Cloud Engineer - Software, Machine Learning/ ML Engineer, DevOps
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
The TeamWe are the All World Team under the Institute of Foundation Model (IFM) All World, we are pioneering the development of the PAN (Physical, Agentic, and Networked) world models—next‑generation foundation models to unlock machine intelligence beyond lingual.
Our mission is to tackle the fundamental challenges of world modeling and establish a new paradigm for next‑generation machine reasoning.
Role OverviewWe’re looking for a Machine Learning Engineer focused on ML infrastructure and MLOps to design and operate the systems that power our research environment. You’ll build scalable, reliable, and observable cloud infrastructure, working closely with researchers to support data pipelines, experimentation, and evaluation workflows.
This role balances fast‑moving research needs with production‑grade systems, ensuring that experimental work can scale reliably when needed.
Key Responsibilities- Design, build, and operate scalable ML infrastructure on AWS (e.g., compute, storage, networking, access control).
- Develop and maintain MLOps workflows for data versioning.
- Build and manage distributed systems for large‑scale data processing (filtering, captioning, etc.) and model evaluation.
- Own architecture decisions for ML infrastructure and drive best practices in reliability, scalability, and cost efficiency.
- Implement observability across systems, including monitoring, logging, and alerting.
- Integrate OpenWebUI, Gradio, or similar UIs for data quality assurance.
- Build and maintain dashboards for experiment tracking and system health.
- Partner closely with researchers to translate experimental workflows into robust, scalable systems.
- 3+ years of experience in MLOps, ML infrastructure, or related backend/platform engineering roles.
- Strong experience with cloud platforms (preferably AWS) and core services for compute, storage, and access control.
- Experience designing and operating distributed systems (e.g., Kubernetes, Ray, or similar frameworks).
- Solid software engineering skills, including system design, debugging, and testing (Python, Docker, Git).
- Familiarity with data processing and pipeline orchestration tools (e.g., Spark, Kafka, or similar).
- Experience with observability practices (monitoring, logging, alerting).
- Ability to work closely with researchers and translate ambiguous requirements into production‑ready systems.
- Experience in fast‑paced or research‑driven environments.
- Experience with large‑scale video or multimodal data pipelines.
- Experience building automated model evaluation or benchmarking systems.
- Knowledge of cost optimization, security, and networking in multi‑tenant environments.
- Familiarity with modern developer and AI‑assisted coding (e.g., Codex, Cursor, Claude Code).
This position is eligible for visa sponsorship.
Benefits- Comprehensive medical, dental, and vision benefits
- Bonus
- 401k Plan
- Generous paid time off, sick leave and holidays
- Paid Parental Leave
- Employee Assistance Program
- Life insurance and disability
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).