×
Register Here to Apply for Jobs or Post Jobs. X

Infra Ops Engineer — Automation & First‑Line Support

Job in New York, New York County, New York, 10261, USA
Listing for: Weights & Biases
Full Time position
Listed on 2026-07-19
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 153000 - 204000 USD Yearly USD 153000.00 204000.00 YEAR
Job Description & How to Apply Below
Location: New York

Core Weave, the AI Hyperscaler™, acquired Weights & Biases to create the most powerful end-to-end platform to develop, deploy, and iterate AI faster. Since 2017, Core Weave has operated a growing footprint of data centers covering every region of the US and across Europe, and was ranked as one of the TIME
100 most influential companies of 2024. By bringing together Core Weave’s industry-leading cloud infrastructure with the best-in-class tools AI practitioners know and love from Weights & Biases, we’re setting a new standard for how AI is built, trained, and scaled.

The integration of our teams and technologies is accelerating our shared mission: to empower developers with the tools and infrastructure they need to push the boundaries of what AI can do. From experiment tracking and model optimization to high-performance training clusters, agent building, and inference at scale, we’re combining forces to serve the full AI lifecycle — all in one seamless platform.

Weights & Biases has long been trusted by over 1,500 organizations — including AstraZeneca, Canva, Cohere, OpenAI, Meta, Snowflake, Square,Toyota, and Wayve — to build better models, AI agents and applications. Now, as part of Core Weave, that impact is amplified across a broader ecosystem of AI innovators, researchers, and enterprises.

As we unite under one vision, we’re looking for bold thinkers and agile builders who are excited to shape the future of AI alongside us. If you're passionate about solving complex problems at the intersection of software, hardware, and AI, there's never been a more exciting time to join our team.

What You'll Do

The Infra Ops team is a new team in the Execute pillar of Weights & Biases infrastructure. Our mission is to keep our platform engineers building: we absorb, reroute, and automate incoming infrastructure requests — incidents, how-to questions, troubleshooting requests, and one-off asks — so that other pillars can focus on their deliverables. Build and release engineering is part of the Execute pillar — an organized, safe, and robust release process is the natural way to reduce post-release incidents, rollbacks, and toil.

We work across the full W&B/Core Weave infra stack — Kubernetes, Go, Terraform, Click House, CircleCI, ArgoCD, Argo Rollouts, Git Hub Actions, and others — running on GCP, AWS, and Azure, across multi-tenant SaaS, dedicated cloud, and on-prem deployments.

About

The Role

We are seeking an Infrastructure Operations Engineer to be part of the first line of support for the infrastructure org. This is a support-oriented, interrupt-driven role: your days are shaped by incoming requests from our internal customers — Solutions Engineers, security, compliance, and product developers on other teams — rather than by a single long-running project. With guidance from senior teammates, you'll triage and resolve many requests end to end, and help convert recurring issues into documentation and automation to prevent repeats.

It's a fast way to learn the entire infrastructure surface, and a natural "landing pad" into infrastructure engineering. You'll write code most days — automation and scripts in Python, Bash, and Go — and take part in a first-responder rotation covering the daily request peak in Slack and office hours. Your customers are primarily internal, though you'll occasionally work a customer escalation when our merchant-support and SA teams need help.

In This Role, You Will

  • Triage and resolve incoming infrastructure requests across Slack, Jira, and office hours — resolving well-scoped ones yourself and escalating others with a clear, reproducible hand-off.
  • Troubleshoot problems across the stack — dig through logs, systems, and configuration to work out what is actually happening, asking for help when a problem runs deep.
  • Write scripts and other automation in Python, Bash, and Go that reduce repetitive work.
  • Help maintain and improve build and release pipelines, along with the runbooks and self-service docs the team depends on.
  • Partner and collaborate closely with our SA, security, and product engineers on one side, and with the other infra pillars on the other.
  • Take part in a 24/7 escalation…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary