Software Engineer, Capacity Engineering
Listed on 2026-07-18
-
IT/Tech
SRE/Site Reliability, Data Engineering
Capacity Engineer
Anthropic’s Capacity Engineering team manages one of the largest and fastest‑growing infrastructure fleets, spanning multiple accelerator families, CPU families and cloud providers. The engineer builds production systems that ingest telemetry, provide observability, and measure utilization of compute resources.
Key Responsibilities- Data Platform
:
Build data pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve the resulting tables in Big Query for analysis by researchers, finance, and leadership. Ensure correctness, completeness and low latency of the data. - Planning
:
Create real‑time tooling for fleet health, capacity planning and alerting. Operate Kubernetes‑native infrastructure at scale and coordinate cross‑team scheduling efforts to reduce fragmentation and improve resource usage. - Efficiency
:
Measure and improve the utilization of training, inference and evaluation workloads. Develop benchmarking infrastructure, establish baseline metrics, and work with system‑owning teams to close gaps. - Attribution & Forecasting
:
Reconcile cloud provider billing exports with internal telemetry, attribute spend to workloads and teams, and produce defensible compute plans that survive finance review. - Own the planning and allocation stack used by leadership and teams for capacity allocation, adopting cross‑region and cross‑provider placement guardrails, queueing, and occupancy KPIs.
- Drive efficiency programs such as rightsizing, unused capacity recovery, and job‑level utilization improvements.
- Develop and maintain the Big Query data platform, ensuring SLOs for completeness, latency and detection of data gaps.
- Strong track record building and operating production systems in a dev‑ops environment.
- Proficiency in Python and SQL; code must be idiomatic, well‑tested and maintainable.
- Hands‑on experience with at least one major cloud provider (AWS, GCP, or Azure) and its operations.
- Experience with observability tooling (Prometheus, PromQL, Grafana) and building monitoring that teams rely on.
- Ability to gather requirements and work across organizational boundaries in ambiguous environments.
- Capacity planning, resource management or cost attribution experience at a hyperscaler or large‑scale ML environment.
- Product engineering experience focused on developer experience and internal data products.
- Experience with scheduling, packing efficiency or profiling‑driven optimization of distributed workloads.
- Multi‑cloud data ingestion expertise, including billing export normalization and reservation APIs.
- Knowledge of total cost of ownership and forecasting, including decomposing infrastructure growth drivers.
- Familiarity with accelerator infrastructure and GPU/TPU utilization metrics.
- Experience building internal data products with self‑service access, schema contracts and discoverability.
- Storage efficiency, retention, and lifecycle program expertise at scale.
Annual Salary: $320,000 – $485,000 USD
Qualifications and Application NotesMinimum education:
Bachelor’s degree or equivalent combination of education, training, and/or experience. Field of study pertinent to the role as demonstrated through coursework, training, or professional experience.
Minimum years of experience:
Years of experience required will correlate with the internal job level requirements for the position.
Location-based hybrid policy:
Staff are expected to be in an Anthropic office at least 25% of the time.
Visa Sponsorship:
Anthropic sponsors visas for qualified candidates. We will make every reasonable effort to obtain the appropriate visa.
Anthropic is an Equal Opportunity Employer and does not discriminate on the basis of race, color, religion, sex, national origin, disability, veteran status, gender identity, sexual orientation, or any other protected characteristic.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).