AgenticOps SME
Job in
Santa Clara, Santa Clara County, California, 95053, USA
Listed on 2026-09-06
Listing for:
Noblesoft Technologies
Full Time
position Listed on 2026-09-06
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software), SRE/Site Reliability
Job Description & How to Apply Below
Agentic Ops SME LLMOps / MLOps / GCP
Santa Clara, CA
Role Summary
- Senior Agentic Ops / LLMOps / MLOps Subject Matter Expert responsible for operating, monitoring, and continuously improving production AI agents.
- Own the Agentic Ops framework across the AI Agent Factory, ensuring reliability, safety, observability, cost efficiency, governance, and business KPI performance
. - Lead the complete agent lifecycle from validation and deployment through production monitoring, evaluation, optimization, and continuous improvement.
- Define and operate Agentic Ops frameworks covering agent registry, versioning, controlled rollouts, rollback, and lifecycle governance
. - Establish continuous evaluation and monitoring for quality, autonomy, safety, latency, cost, reuse, and reliability
. - Implement observability and distributed tracing for multi-agent systems using Google Cloud Agent Engine, Cloud Monitoring, Cloud Logging, and Cloud Trace
. - Own the validation-to-production gate process and manage post-production issues and escape remediation.
- Design Human-in-the-Loop (HITL) supervision, feedback mechanisms, and automated pre-production simulations.
- Track agent business KPIs such as CSAT, TAT, MTTR, cost savings, and operational performance through dashboards and analytics.
- Drive LLM/agent cost optimization through model tiering, context caching, batch/flex inference, budget controls, and cost alerts
. - Partner with Dev Ops, AI, and Data teams to establish an effective build deploy operate improve lifecycle.
- Provide technical guidance on Agentic Ops operating models, ownership transition, governance, and enterprise adoption.
- Strong hands‑on experience with LLMOps, MLOps, Agent Ops, or AI platform operations in production.
- Extensive experience operating GenAI and agentic AI systems in enterprise environments.
- Hands‑on expertise with Google Cloud Vertex AI, Agent Engine, and GCP observability tools
. - Strong knowledge of AI/agent evaluation frameworks, guardrails, Model Armor, HITL, prompt testing, and robustness testing
. - Experience with monitoring, distributed tracing, reliability engineering, and SRE practices for AI workloads
. - Strong understanding of LLM/agent performance and cost optimization
. - Proficiency in Python
. - Strong understanding of agent lifecycle management, governance, reliability, and production operations.
- Experience with ADK, A2A, MCP, and multi‑agent orchestration in production.
- Experience using Big Query and Looker for agent analytics and KPI reporting.
- Knowledge of Responsible AI, model governance, audit, compliance, and AI risk frameworks
. - Experience supporting enterprise‑scale AI/ML platforms
.
- 9 12+ years of experience in ML/AI platform operations, SRE, MLOps, or LLMOps, with significant production GenAI/agentic AI experience.
- Google Cloud Professional certification in Machine Learning, Dev Ops, or a related discipline is preferred.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×