×
Register Here to Apply for Jobs or Post Jobs. X

Applied AI Software Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Canvas Construction
Full Time position
Listed on 2026-07-08
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 300000 - 400000 USD Yearly USD 300000.00 400000.00 YEAR
Job Description & How to Apply Below

Canvas Medical is the electronic medical records (EMR) and payments development platform for healthcare. We build modern, elegant front‑and‑back‑end tooling to enable new ways for developers and clinicians to collaborate to solve healthcare’s toughest challenges. Canvas is institutionally backed by some of the greatest technology investors in the world (funded notable health tech companies such as Good Rx, Oscar Health, and Hims & Hers Health).

The Role

We’re hiring an Applied AI Software Engineer to lead evaluations for agents in development and the post‑deployment fleet of agents operating in Canvas to automate work for our customers. You will help develop agents in Canvas using state‑of‑the‑art foundation model inference and fine‑tuning APIs along with our server‑side SDK. The server‑side SDK provides extensive tools and virtually all the context necessary for excellent agent performance.

You’ll be responsible for designing and running rigorous evaluation experiments that measure performance, safety, and reliability across a wide variety of clinical, operational, and financial use cases.

This role is ideal for someone with deep experience evaluating LLM‑based agents ’ll create high‑fidelity unit evals and end‑to‑end evaluations, define expert‑determined ground truth outcomes, and manage iterations across model variants, prompts, tool use, and context window configurations. Your work will directly inform model selection, fine‑tuning, and go/no‑go decisions for AI features used in production settings.

You’ll collaborate with product, ML engineering, and clinical informatics teams to ensure that Canvas’s AI agents are not only capable, but trustworthy and robust under real‑world healthcare constraints. You will also work with technical product marketers and developer advocates to help our broader developer community and the broader market understand the uniquely differentiated value of agents in Canvas.

Who You Are
  • You have extensive hands‑on experience evaluating LLM‑based systems, including multi‑agent architectures and prompt‑based pipelines.
  • You are deeply familiar with foundation model APIs (OpenAI, Claude, Gemini, etc.) and how to systematically benchmark agent performance using those models in applied settings.
  • You care about correctness and reproducibility and have built or contributed to frameworks for automated evals, annotation pipelines, and experiment tracking.
  • You bring structure to ambiguity and know how to define “correctness” in complex, nuanced domains.
  • You are comfortable collaborating across engineering, product, and clinical subject matter experts.
  • You are not afraid of complexity and are energized by the rigor required in healthcare deployments.
What You’ll Do
  • Design and execute large‑scale evaluation plans for LLM‑based agents performing clinical documentation, scheduling, billing, communications, and general workflow automation tasks.
  • Build end‑to‑end test harnesses that validate model behavior under different configurations (prompt templates, context sources, tool availability, etc.).
  • Partner with clinicians to define accurate expected outcomes (gold standard) for performance comparisons in domains of clinical consequence, and partner with other subject matter experts in other non‑clinical domains.
  • Run and replicate experiments across multiple models, parameters, and interaction types to determine optimal configurations.
  • Deploy and maintain ongoing sampling for post‑deployment governance of agent fleets.
  • Analyze results and summarize tradeoffs in clarity for product and engineering stakeholders, as well as for technical stakeholders among our customers and the broader market.
  • Take ownership over internal eval tooling and infrastructure, ensuring speed, rigor, and reproducibility.
  • Identify and recommend candidates for reinforcement fine‑tuning or retrieval augmentation based on gaps identified in evals.
What Success Looks Like at 90 Days
  • An expanded set of robust evaluation suites exists for all major AI features currently in development and in production.
  • We have well‑defined correctness criteria for each workflow and a reliable source of expert‑determined outcome objects.
  • Product and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary