Data Architect - Austin
Listed on 2026-08-22
-
Software Development
Data Engineering
About the company
Biorce is a pioneering Healthtech company dedicated to revolutionizing drug development through the power of AI. We are passionate about accelerating medical advancements and improving patient outcomes.
Our team comprises seasoned clinical research professionals, data scientists, and AI experts, working collaboratively to bridge the gap between cutting-edge technology and real-world clinical needs.
With an unwavering commitment to revolutionize healthcare, we envision a world where all patients benefit from accelerated and cost-effective access to treatments. Biorce is poised to redefine the landscape of healthcare, shaping a future where innovation and accessibility converge for the betterment of humanity.
About the roleFollowing our successful expansion into the U.S. and continued growth across Europe, we are seeking a Senior Data Engineer / Data Architect to lead a greenfield build-out of Biorce's agentic data platform from our Austin hub.
This is a rare opportunity to design a platform from a clean slate. You will architect and build a hybrid, open-source + Databricks Lakehouse platform on Google Cloud (GCP), with dbt as the transformation and modeling layer, deliberately favoring open standards and open table formats to stay flexible and avoid lock-in, while leaning on Databricks for scale where it earns its place.
At the center of the vision is a self-healing, agentic data warehouse: a platform where querying, pipeline construction, and the merging and mapping of heterogeneous sources into a common data model are all driven by AI agents, with humans setting direction, guardrails, and standards. You will be the technical anchor who turns that vision into a running system.
This is a high-impact, hands-on leadership role for someone who wants to design the data backbone of a next-generation clinical AI platform and rethink how data engineering itself gets done.
Who We're Looking ForA senior engineer or architect who is equally comfortable whiteboarding a Lakehouse from first principles and getting into the weeds of a Spark job or a dbt model. Someone who has strong opinions on data architecture, cares deeply about reliability and compliance in a regulated environment, and is genuinely excited about using AI agents to accelerate not replace rigorous engineering.
You will work closely with data scientists, AI engineers, MLOps, and Dev Ops to design and operationalize robust data flows that fuel advanced analytics, machine learning, and regulatory-grade insights while mentoring the team and raising the bar on how we build.
The Self-Healing, Agentic Data WarehouseThis is the defining concept of the role. We want to build a platform where the core data engineering loops are agentic and self-correcting:
Agentic querying, a natural-language and self-correcting query layer over a well-defined semantic model, where agents translate intent into validated SQL, explain results, and recover from errors without hand-holding.
Agentic pipeline building, agents that scaffold, build, test, and document ingestion and transformation pipelines (Spark / dbt) from specifications, with engineers reviewing and steering rather than hand-writing boilerplate.
Agentic mapping to a common data model, agents that profile, merge, and map heterogeneous clinical, research, and third-party sources into a canonical common data model, handling schema mapping, entity resolution, and normalization.
Self-healing operations, pipelines and tests that detect schema drift, data-quality failures, and anomalies, then diagnose and auto-remediate (or open a well-scoped fix for review), minimizing manual firefighting.
You will design the guardrails, evaluation criteria, and human-in-the-loop checkpoints that make this safe and trustworthy for clinical, regulated data.
Key ResponsibilitiesPlatform & Architecture (Greenfield)
- Architect and build, from a clean slate, a hybrid open-source + Databricks Lakehouse platform on GCP, with dbt as the transformation standard.
- Make deliberate build vs. buy and open-source vs. managed decisions, selecting and integrating open table formats, orchestration, and processing frameworks that keep us flexible and cost-aware.
- Design scalable, fault-tolerant batch and streaming pipelines using Spark/PySpark, Structured Streaming, and Databricks Workflows / declarative pipelines.
- Establish dbt project structure, layering (medallion / bronze-silver-gold), testing, documentation, and CI/CD via the dbt-Databricks adapter.
- Define a common/canonical data model for clinical and research data and the mapping strategy into it.
- Set standards for data quality, lineage, and observability across all systems through rigorous validation, testing, and monitoring.
AI-First / Agentic Data Engineering
- Design and build the self-healing, agentic data warehouse described above, agentic querying, pipeline building, source mapping, and self-healing operations.
- Pioneer and scale AI-assisted data engineering across the team using Claude Code, coding…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).