×
Register Here to Apply for Jobs or Post Jobs. X

Senior CloudOps Engineer

Job in Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: cloudzero
Full Time position
Listed on 2026-08-04
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

About the Role

Cloud Zero is growing fast. Our customer base is expanding, the data challenges we're solving are getting more complex, and the platform is scaling to match. We're standing up real-time ingestion on Kafka right now, spanning several engineering teams, and it's the most operationally demanding thing we've built. Nobody owns the reliability of that path end to end today. That's the first thing you'd own.

As a Senior Site Reliability Engineer you'll be a force multiplier for our engineering organization, owning the reliability, performance, and observability of the systems every team depends on, and empowering teams to ship features that help customers understand and optimize their cloud spend.

This is real infrastructure work at real scale, not a ticket-closing role and not a console-clicking job. Cloud Zero processes billions of events daily across AWS, Azure, and GCP. Our customers rely on real-time, accurate cost data to make business-critical decisions, and any instability in our system impacts their planning. Built entirely on a unique serverless architecture with no EC2s and no containers, our platform demands infrastructure that scales gracefully, fails predictably, and recovers automatically.

There are no Kubernetes clusters or broker fleets to tune here, so the reliability work is engineering, not firefighting. On-call is light:
Nimbus carries a weekly rotation for shared infrastructure, and feature teams respond for their own services.

If you thrive on hard operational problems, care deeply about reliability and performance, and want to see your work matter to customers in direct and measurable ways, this role was built for you.

What You'll Do

Reliability and Observability

  • Own the reliability practice for Cloud Zero's real-time ingestion path: SLOs that span team boundaries, the failure modes nobody owns because they live in the seams, and the architectural changes that come out of what you learn

  • Sign off on shared critical paths before they go live, and call the pause conversation when an error budget burns

  • Instrument systems so that failures surface quickly and debugging happens with data, not guesswork

  • Build observability into everything so you know about problems before customers do

Python and Infrastructure as Code

  • Build the reliability tooling, not just the recommendations: load generators, fault-injection harnesses, SLO instrumentation libraries, deployment safety checks

  • Write production Python across shared libraries, internal services, automation, and agents, and set the standards others build against

  • Design and maintain Cloud Formation and SAM modules that provision reliable, cost-efficient cloud resources

  • Own infrastructure end to end with no clicking through consoles

Automation

  • Automate deployments, scaling, backups, and limit changes; if humans are doing it repeatedly, build a system to do it instead

  • Balance automation intelligently, building solutions to real problems rather than automating for its own sake

  • Evaluate the autonomous agents we already run in production honestly, make the good ones excellent, and throw away the ones that aren't worth it

  • Make our systems legible to AI tooling as well as to people, starting with the service and ownership metadata in our developer portal

Partner with Product Engineering

  • Help teams design resilient services, review architectures for operational complexity, and build deployment pipelines that enable safe and fast shipping

  • Bake SLOs and instrumentation into shared templates so teams inherit good practice instead of reinventing it

  • Drive adoption across 40+ engineers by building the case, not by mandate

  • Optimize for cost and performance;
    Cloud Zero's business is helping others optimize cloud costs, and we should be exemplars of efficient cloud usage ourselves

What You Bring
  • Strong production Python as your primary language, owned, tested, and maintained at scale

  • An SLO you defined yourself, including what you deliberately chose not to alert on

  • Experience operating asynchronous, event-driven systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure (Kafka, Kinesis, SQS, Pulsar, or Step Functions; we…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary