×
Register Here to Apply for Jobs or Post Jobs. X

Senior MLOps Engineer

Remote / Online - Candidates ideally in
Atlanta, Fulton County, Georgia, 30383, USA
Listing for: Jobot
Remote/Work from Home position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 150000 - 175000 USD Yearly USD 150000.00 175000.00 YEAR
Job Description & How to Apply Below

Job details Own where changes land and what it costs to run

This Jobot Job is hosted by:
Charles Simmons

Salary: $150,000 - $175,000 per year

A bit about us:

Small, mid staged AI native SaaS startup helps state and local governments modernize paper-based processes into intelligent, AI-driven digital workflows. As we evolve into an AI-first platform, our development velocity, model iteration frequency, and cross-team complexity increase. A reliable, cost-disciplined platform is essential to scale safely and predictably.

Why join us?
  • 100% work from home (US based only)
  • Own the platform that build, validation, and release loops run on and deploy to: infrastructure, environments, Kubernetes, Networking, observability, and the AI serving and routing layer.
  • Competitive base, bonus, and equity options
  • Medical, dental, and vision insurance plans, with significant employer contributions for employees AND dependents (contributions based on base-level plan; buyup plans available at additional costs)
  • Company-sponsored life, short-term, and long-term disability insurance
  • 11 Paid holidays
  • Flexible time off
  • 401k plan with 4% employer match
  • Monthly stipend for home office expenses
  • Monthly wellness stipend
Job Details

Everything runs on Azure, and the platform is getting more interesting: an AI product suite heading toward general availability, self-hosted AI observability and telemetry inside a FedRAMP-conscious boundary, autonomous agents participating in delivery, and a microservices decomposition in flight. You will deploy, operate, and scale that platform, and you will own its cost discipline.

This is a production seat with production access, and we treat that as an engineering responsibility, not a badge: least privilege, audit trails, and environment integrity are part of the job, because our customers are governments.

You will work alongside our AI Operations Engineer, who owns the agentic delivery system (the loops that build, validate, and release code). You own the platform those loops run on and deploy to: infrastructure, environments, Kubernetes, networking, observability, and the AI serving and routing layer. The boundary is simple: they own how changes move; you own where changes land and what it costs to run.

Responsibilities
  • Deploy and operate our Azure platform: AKS, networking, identity, storage, and environments from development through production
  • Own infrastructure as code end to end: environments are reproducible, drift is detected, and nothing reaches an environment without platform visibility
  • Operate the AI infrastructure layer: self-hosted observability and evaluation tooling (Langfuse), product telemetry, model gateway and per-workload routing, and compliant Gov Cloud inference paths
  • Own cloud and AI cost: metering, budgets, unit economics, MACC drawdown strategy, and active remediation; cost is an engineering metric here, not a finance afterthought
  • Harden production access and controls: least privilege, secrets management, audit evidence, and a FedRAMP-conscious security posture
  • Partner with AI Operations on the deploy-and-release path:
    Octopus Deploy, environment promotion, progressive rollout, and rollback
  • Build platform reliability: monitoring, alerting, incident response, and capacity planning
  • Give the microservices decomposition the platform primitives it needs: service infrastructure, scaling patterns, and clean environment boundaries
Qualifications
  • 5+ years in Dev Ops, platform engineering, or site reliability engineering in SaaS environments
  • Deep Azure experience: AKS, networking, identity (Entra), and monitoring; you have run production Kubernetes
  • Infrastructure as code as your default (Terraform, Bicep, or similar), plus strong scripting; you automate before you document
  • MLOps experience: deploying and operating LLM or ML systems in production, including model gateways, inference infrastructure, or AI observability stacks
  • Demonstrated cost work: you can point to cloud spend you found, explained, and reduced
  • Experience in compliance-heavy environments (FedRAMP, StateRAMP, SOC 2, or similar) is a strong plus
  • Comfortable holding production access, with the discipline that implies
Key Competencies
  • Treats environment integrity as sacred: no invisible changes, no snowflake servers, no heroics that cannot be audited
  • Cost literacy: reads a cloud bill the way an engineer reads a stack trace
  • Automates first: your instinct is a pipeline or a policy, not a runbook step
  • thinks in the open: surfaces risk early and documents what you build
  • Calm in production incidents; rigorous in the postmortem
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary