More jobs:
Platform Engineer, Model Shaping
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-09-09
Listing for:
Together AI
Full Time
position Listed on 2026-09-09
Job specializations:
-
Software Development
Cloud Engineer - Software, DevOps, Backend Developer
Job Description & How to Apply Below
- As a Platform Engineer at Model Shaping, you will work on the foundational layers of Together’s platform for model customization and evaluation
- You will design the infrastructure and backend services that will allow us to sustainably and reliably scale the systems powering production workflows launched by our users, as well as internal research experiments
- You will operate in a cross‑functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on
- You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams
- Design and build Together’s systems and infrastructure for model customization, including user‑facing features and internal improvements
- Contribute to reliability improvements for the platform, participating in an on‑call rotation and improving processes for incident response
- Create and improve internal tooling for deployment, continuous integration, and observability
- Build a job orchestration platform spanning multiple data centers, supporting a highly heterogeneous hardware landscape
- Partner with teams developing internal services, co‑designing these services and incorporating them in systems built by Model Shaping
- Competitive health insurance plans
- Pre‑tax flexible spending accounts
- Dental and vision insurance
- Income protection & retirement
- Mental health support and services
- Life insurance
- AD&D insurance
- 401(k) plan
- STD & LTD insurance
- Monthly commuting stipend + pre‑tax bene
- Flexible time off policy
- Monthly team lunches
- Team‑driven celebrations and events
- Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (Git Hub Actions, ArgoCD)
- 3+ years of experience in building infrastructure or backend components of production services
- Skilled with analyzing non‑trivial issues of complex software systems and documenting your findings
- Strong software engineering background in Python or Go
- Strong communication skills, willing to document systems and processes and collaborate with peers of varying technical expertise
- Have cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare‑metal/cloud environment
- Comfortable with the fundamentals of Linux environments and modern container/orchestration stacks (e.g., Docker and Kubernetes)
- Developing large‑scale production systems with high reliability requirements
- Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte)
- Deployment of services for AI training or inference
- Managing GPU workloads on HPC clusters, ideally with hands‑on experience in operating NVIDIA’s networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA)
- Maintaining or contributing to open‑source projects
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×