SRE AI; share AI/ML profiles C2C
Job in
Mountain View, Santa Clara County, California, 94039, USA
Listed on 2026-09-19
Listing for:
Tech Mirrors
Full Time
position Listed on 2026-09-19
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software), SRE/Site Reliability
Job Description & How to Apply Below
SRE with AI
Location:
Mountain View, CA (Onsite)
Contract
Key Responsibilities- Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
- Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
- Extend tracing across agentic flows - planner → executor → retrieval → tool calls - spanning GenOS, AI Gateway, and MCP Gateway.
- Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
- On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events - blameless postmortems with tracked remediations.
- Own the SRE side of the progressive-delivery seam:canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls - Dev Ops builds the pipeline; you define the gates.
- Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
- Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).
- 7+ years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
- Strong Go or Python - this is a build role: tooling, automation, instrumentation.
- Hasbuiltobservability stacks, not just consumed them:
Prometheus/Grafana, Open Telemetry, or equivalent at scale, including cardinality and cost control. - Production LLM/ML monitoring:
Langfuse, Arize, Why Labs, or homegrown - token/cost tracking, drift and quality metrics. - Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
- Kubernetes + AWS operational depth - debugs across cluster, mesh, and gateway layers.
- Structured incident-command and postmortem experience.
- AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
- Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
- Fin Ops / cost engineering; fintech or regulated-industry reliability experience.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×