×
Register Here to Apply for Jobs or Post Jobs. X

Senior Platform AI Engineer

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Drata Inc
Full Time position
Listed on 2026-07-01
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, AI Reliability/ Performance Engineer, Backend Developer
Job Description & How to Apply Below

Drata AI Platform Team Role

Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance platform — from MCP servers that make Drata's data available to AI agents, to LLM workflow orchestration that automates SOC 2, TPRM, and policy analysis. You'll own the systems that sit between our AI models and our customers: tool definitions that agents actually understand, deployment pipelines that handle model upgrades without breaking output quality, and orchestration layers that manage multi-step agent workflows with persistent state.

This is not a traditional infrastructure role. You'll debug prompt templates alongside Terraform modules. You'll design API schemas optimized for LLM token budgets, not just HTTP throughput. When a model upgrade changes behavior across 15 workflows, you'll assess quality impact — not just confirm the containers are healthy.

You'll work closely with our agent developers, product engineers, and an embedded SRE partner, sitting at the intersection of AI development and production reliability.

Our north star is simple: minimize the time it takes to launch a new agent in production. You're someone who asks "are we solving the right problem?" before writing the first line of code, who builds systems that make five other engineers faster, not just yourself, and who's equally proud of what they chose not to build.

What you'll do:

MCP Server Development & AI-Optimized API Design
  • Design and build MCP (Model Context Protocol) servers that expose Drata's platform to AI agents. This means making architectural decisions about tool granularity, naming conventions for agent disambiguation, response compression for LLM context windows, and workspace isolation for multi-tenant access. You'll own the protocol layer that determines whether agents can reliably find and use the right tools — writing semantic parameter descriptions, contextual hints, and tool schemas that optimize for model comprehension, not just developer ergonomics.

Agent

Orchestration & Workflow Infrastructure
  • Build and operate the infrastructure for deploying multi-step agent workflows — state management across complex reasoning chains, tool routing and execution runtimes, and long-running agentic processes that persist over time. Own the orchestration layer that coordinates agent planning, tool calls, and human-in-the-loop patterns. Design systems that handle agent failure modes gracefully: retries on ambiguous tool outputs, fallback strategies when models produce unexpected results, and observability into multi-step execution traces.

LLM

Operations & Model Lifecycle Management
  • Own the operational side of our LLM workflows: model upgrades across production pipelines (assessing behavior changes, not just version bumps), prompt versioning and A/B testing, AI workflow deployment with custom container compatibility, and output quality monitoring.

  • Manage token capacity planning — understanding model costs, context limits, batching strategies, and rate governance across workflows. When an AI workflow fails, you'll investigate whether it's a prompt template issue, a model behavior change, or an infrastructure problem. Making that distinction requires understanding both systems.

Production AI Infrastructure & RAG Systems
  • Operate and evolve our production AI stack: vector storage and indexing (designing chunking strategies and metadata schemas for retrieval quality), document parsing pipelines, multi-region deployment, and cost optimization across LLM providers. You'll make RAG architecture decisions — embedding strategies, retrieval filtering, data model coordination — where the engineering challenge is search quality, not just system uptime. Implement caching layers and token-aware request routing to manage spend as AI workloads scale.

Platform

Enablement & Developer Experience
  • Build CI/CD patterns specific to AI workflows (reproducible deployments, SDK version compatibility, workflow rollback semantics). Own AI-specific observability — token usage dashboards, response quality metrics, agent execution traces, and cost-per-workflow tracking alongside traditional infrastructure monitoring.…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary