×
Register Here to Apply for Jobs or Post Jobs. X

Senior DevOps​/Infrastructure Engineer

Job in New York, New York County, New York, 10261, USA
Listing for: Category Labs
Full Time position
Listed on 2026-09-18
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 180000 - 250000 USD Yearly USD 180000.00 250000.00 YEAR
Job Description & How to Apply Below

Category Labs (formerly known as Monad Labs) is a team of systems engineers and researchers on a mission to design and build at the frontier of decentralized technology. We strive to deliver significant improvements over existing blockchain solutions. After raising $225M in series A funding, led by Paradigm, we are growing our team.

We’re the team behind Monad, a high-performance, EVM-compatible Layer 1 whose public mainnet is now live. We write the core software that runs it: a parallel-execution EVM, a custom state database, and a BFT consensus client, all developed in the open.

The Role

We're looking for a Senior Dev Ops / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet.

As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads.

What You'll Do
  • Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.

  • Own our infrastructure-as-code:
    Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (Git Ops) for platform services.

  • Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.

  • Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).

  • Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.

  • Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.

  • Harden nodes and services, manage secrets, and continuously drive down manual toil.

Who You Are
  • You have 5+ years in Dev Ops, SRE, or Infrastructure Engineering, operating production systems at scale.

  • You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.

  • You have deep, hands‑on infrastructure-as-code experience with Ansible and Terraform.

  • You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).

  • You have hands‑on fluency with AI‑assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.

  • You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.

  • You have programming and scripting experience (e.g., Python, bash).

  • Experience with Kubernetes and Git Ops (Flux or Argo) is a plus.

  • Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.

  • Experience serving inference, either locally or as a service is a plus.

  • Previous experience with blockchain clients or node operations is a plus.

  • A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.

Why Work with Us
  • Challenging problems: You’ll work on extremely challenging problems with massive impact. See our Blogs and…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary