×
Register Here to Apply for Jobs or Post Jobs. X

Senior Staff+ Software Engineer, Node Infra

Job in London, Greater London, W1B, England, UK
Listing for: Humanloop
Full Time position
Listed on 2026-08-12
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 325000 GBP Yearly GBP 325000.00 YEAR
Job Description & How to Apply Below
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role Anthropic's Infrastructure organization is foundational to our mission of developing AI systems that are reliable, interpretable, and steerable. The systems we build determine how quickly we can train new models, how reliably we can run safety experiments, and how effectively we can scale Claude to millions of users — demonstrating that safe, reliable infrastructure and frontier capabilities can go hand in hand.

Node Infra owns the full lifecycle of accelerator capacity  ingest and provision compute from all major cloud providers and from datacenters custom-built for Anthropic, stand up and scale the clusters behind one of the industry's largest AI compute fleets, and build the health, diagnostics and repair automation that keep every GPU, TPU and Trainium node in the fleet usable and ready to power Anthropic's frontier AI research.

Key responsibilities

Own the technical strategy and roadmap for node lifecycle management - ingestion, bring-up, health checking, and automated repair

Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families

Design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBI and minimizing stranded capacity

Define infrastructure architecture, ensuring the hardest problems get solved - whether by you directly or by working through others

Work closely with cloud providers and internal research/inference/product teams to shape long-term compute, data, and infrastructure strategy

Establish and evolve operational excellence practices (incident response, postmortem culture, on-call)
Support the growth of engineers around you through technical mentorship and coaching

Minimum qualifications

Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform.

Hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium)
Track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems

Ability to build alignment across senior stakeholders and communicate effectively at all levels

Preferred qualifications
12+ years of software engineering experience, including time as a technical lead setting direction for a team Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency

Depth in one or more of:
Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines

Low-level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons

Familiarity with high-performance networking (EFA, RDMA, Infini Band) for distributed ML workloads.

Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems

Contributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.)Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems

The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:325,000—485,000 GBPLogistics

Minimum education:

Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience:
Years of experience required will correlate with the internal…
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary