×
Register Here to Apply for Jobs or Post Jobs. X

AI DevOps. Engineer

Job in Pittsburgh, Allegheny County, Pennsylvania, 15201, USA
Listing for: Expedient
Full Time position
Listed on 2026-07-11
Job specializations:
  • Software Development
Job Description & How to Apply Below

AI Dev Ops Engineer

Join Expedient's AI CTRL product team as our AI Dev Ops Engineer — a senior, hands-on engineer who will build the framework that manages, configures and ships agentic workflows, tooling applications, and AI integrations to clients quickly, safely, and repeatably. You'll own the path from commit to production:
Git-driven CI/CD, infrastructure as code, release and config management, observability, and the LLMOps practices that keep model-powered systems reliable and cost-efficient. This is a build role — you won't be maintaining someone else's pipelines, you'll be creating the framework the AI Dev team builds on. Because AI CTRL runs on enterprise model APIs (Anthropic Claude, OpenAI, Google Gemini) with RAG and MCP integrations rather than training custom models, this role is LLMOps-focused : prompts, configs, and integrations are the primary code surface, and the operational challenges are deployment velocity, traceability, cost, and risk t You'll Do:

  • CI/CD Pipelines:
    Design and build Git-based pipelines that automate build → test → deploy for Retool apps, agentic workflows, MCP servers, and data connectors — turning manual client deployments into repeatable, gated releases.
  • Infrastructure as Code:
    Make the platform reproducible. Use Terraform, Helm, and Git Ops (ArgoCD/Flux) to provision and manage Kubernetes (Nutanix NKP) clusters and per-client environments as code.
  • Configuration Management:
    Manage environment and deployment configuration as code across a growing fleet of client deployments — eliminate config drift and one-off manual changes.
  • Release Management:
    Own versioning, environment promotion, release gates, and clean rollback. Maintain versioned, deployable artifacts so any release can be reproduced or reverted.
  • Observability & Tracing:
    Build the monitoring backbone — Elastic/ECK, APM, and telemetry distributed tracing — with deployment health, SLOs/SLIs, and usage/cost instrumentation across all client deployments. Strengthen alerting so issues surface before clients feel them.
  • LLMOps Practices:
    Stand up prompt and configuration versioning, model/prompt evaluation pipelines, A/B testing of prompts and models, multi-provider traffic routing and failover, and token/cost dashboards — the AI-specific discipline that keeps model-powered systems accurate, available, and affordable.
  • Change & Risk Management (incl. Compliance):
    Implement controlled-change processes — approvals, audit trails, and guardrails — with compliance-as-code for SOC 2 audit logging, secrets management (e.g., vaults/sealed-secrets), and SSO/OIDC configuration.
  • Automation Marketplace:
    Build an internal library of vetted, reusable workflows, connectors, and IaC modules that accelerate client delivery — and graduate proven items into a client-facing catalog aligned to the Agentic Workflow Engine (AWE).
  • Collaborate & Document:
    Partner with the AI Dev engineering team on platform standards; write the runbooks, release guides, and architecture docs that let the framework scale beyond

What We're Looking For:

  • Experience:

    3–5 years in Dev Ops, platform engineering, site reliability, or MLOps/LLMOps. Prior experience at a managed service provider, SaaS company, or enterprise technology team is a strong plus.
  • Git-based CI/CD: designing automated build/test/deploy pipelines from scratch
  • Infrastructure as Code:
    Terraform and Helm;
    Git Ops with ArgoCD or Flux
  • Kubernetes: operating and automating clusters (Nutanix NKP or equivalent); name spaces, workloads, container lifecycle
  • Observability:
    Elastic/ECK, APM, Open Telemetry tracing; defining alerts, SLOs/SLIs (Prometheus/Grafana experience transfers)
  • Scripting & data: strong Python and Bash ; SQL fundamentals
  • Secrets & identity: secrets management (Vault or equivalent), SSO/OIDC configuration (Entra , Okta, One Login)
  • Workflow orchestration:
    Argo Workflows, Airflow, or similar (a plus)
  • LLM APIs: working familiarity with Anthropic Claude, OpenAI, and/or Google Gemini — prompt construction, tool use/function calling, token management
  • RAG & MCP awareness: chunking, embedding, vector search, context-window management;
    Model Context Protocol integrations (a plus)
  • Compliance exposure:…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary