Principal Python + LangGraph Engineer
Listed on 2026-07-25
-
Software Development
AI Engineer (Applied/Software), Backend Developer
Engagement context
We are partnering with a US-based health-tech company on the takeover of a production AI-powered mobile coaching platform. The platform is built around a Python AI core (atlas-ai) which runs:
- A FastAPI chat surface with intent routing and tool-using agents
- A Lang Graph-based agent framework with multi-agent orchestration
- A Celery + Redis task queue for asynchronous agent flows
- MongoDB for fitness-plan storage, Redis for conversation state, Postgres for Lang Graph checkpoints
- Several agent personas: onboarding, chat, plan creation, in-workout smart adjust, plan smart adjust, habit formation
- Direct OpenAI integration for the LLM layer
This is the product. If the AI core breaks, the product breaks. If LLM token cost runs away, our margin runs away. If the agent flows behave unpredictably under load, real users get bad coaching.
Role summaryAs Principal AI Engineer you are the technical owner of the AI core. You lead the audit of the existing codebase, define the architecture we evolve toward, build the test and evaluation harness that lets us ship changes safely, and you are the engineer who is paged when the AI surface misbehaves in production.
You will be the deciding voice on whether we can responsibly take this service in our scope or whether we keep the client s original team embedded longer. The audit conclusions you write in the first 4 weeks will inform the contractual scope of the entire engagement.
What you ll doFirst 90 days
- Audit atlas-ai: agent flows, Lang Graph state machines, Celery topology, datastore usage, OpenAI integration patterns
- Produce a written assessment of operational risk: failure modes, race conditions, retry semantics, idempotency, checkpoint integrity
- Quantify token cost per agent flow and per user session
- Identify the highest-risk subsystems and propose stabilisation plans
- Build (or harden) an evaluation harness for the agent flows — golden cases, regression suites, hallucination/safety tests
- Lead the knowledge-transfer sessions from the client s AI team
Ongoing
- Set the technical direction for the AI core
- Lead design for new agent flows and major changes to existing ones
- Own the production health of the AI surface (with platform/SRE support)
- Hire and mentor the rest of the AI squad (~10 engineers at full scale)
- Represent the AI core in cross-team architecture conversations with the client
Must-have skills
- 7+ years Python in production at senior+ level
- Deep Lang Graph experience — state graphs, checkpoints, interrupts, multi-agent supervision, subgraphs
- Strong Lang Chain ecosystem knowledge (chains, tools, memory, output parsers, callbacks)
- Production FastAPI — streaming responses, dependency injection, middleware, async patterns
- Celery + Redis broker in production — task ordering, retries, idempotency, priority queues, dead-letter handling
- Concurrency in Python — asyncio (gather, structured concurrency, cancellation), threading boundaries, mixing sync and async code safely
- Multi-datastore operations — MongoDB + Redis + Postgres in a single service, transaction boundaries across them
- OpenAI API at scale — rate limits, retries with exponential backoff, fallback model routing, streaming, tool/function calling
- Prompt engineering with discipline — evaluation, A/B testing, version control of prompts, regression detection
- Token cost optimisation — prompt caching (Anthropic-style), model tiering, context window trimming, summary memory
- Production LLM observability — per-route token spend, prompt-level tracing, drift monitoring
- Testing discipline — pytest (including pytest-asyncio), property-based testing, snapshot tests for prompts, eval-based tests for agents
- Pydantic v2 fluency, type-hinted code throughout
Nice-to-haves
- Production incident command for LLM-powered systems
- ML engineering background (model serving, feature engineering)
- Anthropic / Claude API experience in addition to OpenAI
- Data pipeline experience (Airflow, Dagster, Prefect)
- Domain knowledge in fitness / health / wearables
- Experience working with cross-team JSI / native bridges (the Python core integrates with a mobile JSI layer)
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).