AI/LLM Engineer — Healthcare Integrity
Listed on 2026-09-16
-
Software Development
AI Engineer (Applied/Software)
CODOXO ISNOTABLE TO OFFER SPONSORSHIP OR ACCOMMODATE ANY CANDIDATES THAT ARE CURRENTLY BEING SPONSORED NOW OR IN THE FUTURE
The United States spends roughly $4.9 trillion on healthcare each year, and an estimated quarter of that is lost to waste, fraud, abuse, and error. Codoxo is the premier provider of AI-driven solutions that help healthcare companies and government agencies proactively detect and reduce those losses and ensure payment integrity.
We are purpose-driven, with the goal of making healthcare more affordable and accessible to all. If you are passionate about applied AI and driven by positive impact in healthcare, Codoxo is where you belong. We are venture backed by some of the top investors in the country, with strong financials, and remain one of the fastest growing healthcare AI companies in the industry.
PositionSummary
As an AI/LLM Engineer, you will design, ship, and operate LLM-powered features that accelerate claim audits and SIU investigations — retrieval and question answering over case documents, structured extraction from claims and medical records, and investigator-facing summarization. Just as importantly, you will build the evaluation systems that tell us whether any of it actually works.
Our production stack is Python and Django on AWS, with Amazon Bedrock for inference, Open Search for vector and hybrid search, and Celery for asynchronous document processing. You will work on a small team where you own features end to end and your work reaches fraud investigators at national health plans and government agencies.
As our LLM foundation matures, this role grows into multi-step, tool-using systems — MCP tools, Bedrock Agent Core, and multi-agent orchestration. You would help decide when that complexity is earned rather than inheriting the decision. We are looking for someone who reaches for the simplest thing that clears the bar and can explain why.
Key Responsibilities- Build and ship LLM features against real investigator workflows: document question answering, structured extraction, and summarization over claims, medical records, notes, and correspondence.
- Own our retrieval pipeline end to end — document ingestion and extraction quality, chunking, metadata filtering, hybrid lexical and semantic search, reranking, and citations traced back to source text.
- Build and maintain the evaluation systems that gate our releases: curated golden sets built with subject‑matter experts, regression suites that run on every prompt and model change, and rubric‑based grading validated against human reviewers.
- Treat prompts as engineering artifacts — versioned, code‑reviewed, documented with change notes, and never shipped without an evaluation run.
- Design reliable structured output: schema‑constrained generation, validation with bounded retries, and graceful handling of refusals, truncation, and partial results.
- Build on serverless AWS — Bedrock, Lambda, S3, API Gateway, Step Functions — keeping PHI inside our trust boundary and out of logs and traces.
- Keep cost and latency predictable as usage grows: token accounting, prompt caching, routing tasks to the right model tier, and asynchronous processing for non‑interactive work.
- Instrument LLM features for observability and audit — prompt and model versions, token counts, validation and retry rates, retrieved document IDs.
- Partner with investigators, clinical coders, and product to turn expert judgment into evaluation criteria, not just requirements.
- 3+ years of professional Python experience.
- 2+ years building with LLMs in production — not prototypes or demos. You have shipped something real, watched it fail in ways you did not expect, and fixed it.
- Hands‑on RAG experience: embeddings, vector or hybrid search (Open Search,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).