×
Register Here to Apply for Jobs or Post Jobs. X

Applied AI Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Soulside AI
Full Time position
Listed on 2026-08-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 200000 USD Yearly USD 150000.00 200000.00 YEAR
Job Description & How to Apply Below

Applied AI Engineer

Soulside AI
· US On-Site
· Reports to the CTO

About Soulside

Soulside AI is the specialist AI platform for behavioral health documentation and compliance. We generate audit-ready clinical documentation across individual and group sessions, virtual and in-person care, admissions, and treatment planning-and we embed real-time chart audits and payer-aligned compliance checks into everyday workflows. The result is immediate and measurable: higher-quality charts, stronger medical necessity, and hours given back to clinicians every week.

We're backed by Counterpart Ventures, Grey Matter Capital, and One Mind, and we're a UCSF Rosenman Institute and One Mind Accelerator company. We've reached strong product-market fit and are scaling fast.

The Role

We're looking for an Applied AI Engineer to own the model layer that makes Soulside's documentation trustworthy. In behavioral health, a note isn't just text—it has to be clinically sound, defensible for medical necessity, and safe. Your job is to build the post-training pipelines and evaluation systems that get our models there, and keep them there as we scale.

This is a hands-on role for someone who lives at the intersection of applied ML and product. You'll fine-tune and adapt open-source models, stand up the infrastructure to serve them, and build the rigorous evaluation sets that tell us—objectively—whether a change made the product better or worse.

Why This Role Matters
  • Accuracy Isn't Optional: In behavioral health, a wrong or unsupported note has real clinical and financial consequences. The pipelines and evals you build are what let us ship model changes with confidence.

  • Own the Model Layer: You'll define how we post-train, evaluate, and deploy models end-to-end—not inherit someone else's stack.

  • Direct Clinical Impact: Every improvement in clinical reasoning or note quality directly reduces documentation burden and strengthens the charts clinicians and payers rely on.

What You'll Do
  • Build post-training pipelines on open-source models-supervised fine-tuning, preference optimization (DPO/RLHF), LoRA/adapters, and distillation-for domain-specific clinical tasks.

  • Fine-tune, deploy, and serve models across managed inference and fine-tuning platforms such as Fireworks AI, Baseten, and Together AI
    , and make pragmatic build-vs-buy calls on where each workload should run.

  • Design and maintain rigorous evaluation sets for high-stakes tasks like clinical reasoning and AI note generation
    -defining metrics, curating gold-standard data, and building automated and human-in-the-loop eval harnesses.

  • Turn eval results into a fast, trustworthy iteration loop: catch regressions before they ship, and quantify the impact of every model or prompt change.

  • Optimize the full LLM pipeline-prompting, retrieval, structured-output validation, latency, and cost.

  • Partner with clinical experts to translate documentation and compliance requirements into model behavior and evaluation criteria.

  • Monitor models in production for quality, drift, and failure modes, and close the loop back into training data and evals.

What We're Looking For
  • 3+ years in applied ML / AI engineering, or a Master's degree in a related field, with hands-on experience taking LLM-based systems into production.

  • Practical experience with post-training / fine-tuning open-source models (e.g., Llama, Qwen, Mistral) using SFT, LoRA/PEFT, or preference-based methods.

  • Experience serving or fine-tuning models on managed platforms such as Fireworks AI, Baseten, or Together AI (or comparable inference/training infra).

  • Demonstrated ability to build evaluation frameworks for LLM tasks—you think in terms of measurable quality, not vibes.

  • Strong Python and familiarity with the modern ML tooling ecosystem (PyTorch, Hugging Face, etc.).

  • Solid grounding in prompt engineering and structured-output validation.

  • Ability to thrive in a fast-paced, remote startup and communicate clearly with technical and clinical teammates.

  • We're willing to sponsor visas, including H-1B and O-1, for the right candidate.

Bonus Points
  • Experience with healthcare, clinical NLP, or other high-stakes / regulated domains.

  • Familiarity with HIPAA and handling sensitive clinical data.

  • RAG systems, retrieval quality tuning, or long-context document workflows.

  • Experience with LLM observability, monitoring, and drift detection in production.

  • Data pipeline and labeling workflow experience for curating high-quality training and eval sets.

  • Open-source contributions in the ML/LLM ecosystem.

What We Offer
  • Salary range of $150,000-$200,000, plus equity with significant upside potential as a founding team member

  • Comprehensive health, dental, and vision insurance

  • Flexible, remote-first culture

  • Direct access to founders and influence on technical direction

  • Professional development budget and conference attendance

  • The chance to build AI that measurably improves mental health care at scale

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary