×
Register Here to Apply for Jobs or Post Jobs. X

SRE AI; share AI​/ML profiles C2C

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: Tech Mirrors
Full Time position
Listed on 2026-09-19
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

SRE with AI

Location:

Mountain View, CA (Onsite)

Contract

Key Responsibilities
  • Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
  • Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
  • Extend tracing across agentic flows - planner → executor → retrieval → tool calls - spanning GenOS, AI Gateway, and MCP Gateway.
  • Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
  • On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events - blameless postmortems with tracked remediations.
  • Own the SRE side of the progressive-delivery seam:canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls - Dev Ops builds the pipeline; you define the gates.
  • Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
  • Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).
Must-Have Qualifications
  • 7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
  • Strong Go or Python - this is a build role: tooling, automation, instrumentation.
  • Hasbuiltobservability stacks, not just consumed them:
    Prometheus/Grafana, Open Telemetry, or equivalent at scale, including cardinality and cost control.
  • Production LLM/ML monitoring:
    Langfuse, Arize, Why Labs, or homegrown - token/cost tracking, drift and quality metrics.
  • Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
  • Kubernetes + AWS operational depth - debugs across cluster, mesh, and gateway layers.
  • Structured incident-command and postmortem experience.
Nice-to-Have
  • AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
  • Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
  • Fin Ops / cost engineering; fintech or regulated-industry reliability experience.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary