×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer, Production Engineering; Cloud Prem - W&B

Job in Livingston, Essex County, New Jersey, 07039, USA
Listing for: Weights & Biases
Full Time position
Listed on 2026-07-27
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 139000 - 185000 USD Yearly USD 139000.00 185000.00 YEAR
Job Description & How to Apply Below
Position: Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B

Core Weave, the AI Hyperscaler™, acquired Weights & Biases to create the most powerful end-to-end platform to develop, deploy, and iterate AI faster. Since 2017, Core Weave has operated a growing footprint of data centers covering every region of the US and across Europe, and was ranked as one of the TIME
100 most influential companies of 2024. By bringing together Core Weaveor cloud infrastructure with the best-in-class tools AI practitioners know and love from Weights & Biases, were setting a new standard for how AI is built, trained, and scaled. Weights & Biases has long been trusted by over 1,500 organizations — including AstraZeneca, Canva, Cohere, OpenAI, Meta, Snowflake, Square,Toyota, and Wayve — to build better models, AI agents and applications.

Now, as part of Core Weave, that impact is amplified across a broader ecosystem of AI innovators, researchers, and enterprises. As we unite under one vision, were looking for bold thinkers and agile builders who are excited to shape the future of AI alongside us. If youre passionate about solving complex problems at the intersection of software, hardware, and AI, theres never been a more exciting time to join our team.

What Youll Do The Production Engineering team builds and operates the platform that lets Core Weave's engineers ship software quickly, reliably, and safely. We own the observability systems, reliability tooling, incident management systems, and the infrastructure-as-code that underpins our engineering organizations ability to deliver software to our enterprise customers across GCP, AWS, Azure, and self-hosted, on-premises environments. Our mission is simple: you build it, you deploy it, you run it and we build the tools and processes that make owning your service in production straightforward and safe.

About The Role We are seeking a Senior Production Engineer with deep expertise across both cloud and on-premises environments to design, build, and operate the reliability platform at the core of Core Weave's engineering organization. This is a hands-on Dev Ops/SRE-flavored role: youll work across infrastructure-as-code, CI/CD, observability, and incident systems – automating away toil and giving service teams the visibility and guardrails they need to run their own code in production.

Youll evolve our systems to meet increasingly challenging scale, reliability, and performance demands, and own large, ambiguous operational problems end to end – from error attribution and alert routing to release safety and capacity planning. Youll also provide technical leadership across the team, partnering with senior and principal engineers to influence direction and drive outcomes for cross-functional stakeholders. We care deeply about a sustainable on-call culture, so youll participate in on-call rotations while championing the architecture and automation that reduce on-call burden over time.

In This Role, You Will Design, build, deploy, and operate critical reliability and infrastructure services across AWS, GCP, Azure, and on-premises / hybrid environments. Improve error attribution and alert routing, automatically routing errors and pages to the team that owns the affected service, so engineers are only on-call for what they own. Own observability patterns, SLI/SLO frameworks, service catalog metadata, and dashboards that give teams visibility into their services from code change through production.

Build release-safety systems – canary deployments, smoke tests, staged rollouts, and reliable roll-back / roll-forward – so shipping to production is fast and safe. Advance the incident and on-call program – incident tooling, on-call rotations, runbooks, and operational readiness reviews; measure incident volume by team to focus reliability investment. Reduce on-call burden through better architecture, automation, and observability – and participate in on-call rotations yourself.

Provision and manage infrastructure with Terraform and drive manual, incident-time operations toward automated, repeatable infrastructure-as-code. Break down and solve large, ambiguous operational problems, turning them into well-scoped, shippable engineering…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary