C-BRAIN Data Engineer; Remote; Neurology
Florissant, St. Louis city, Missouri, 63034, USA
Listed on 2026-08-26
-
IT/Tech
Data Engineering
Location: Florissant
Location Remote, US Scheduled Hours 40
The C-BRAIN Data Engineer is a key technical member of the C-BRAIN team responsible for designing, building, and maintaining the data infrastructure that powers C-BRAIN's AI tools. Reporting to the C-BRAIN Chief Technology Officer (CTO), this role is responsible for all aspects of data ingestion, pipeline development, data harmonization, and cloud infrastructure management — ensuring that high-quality, analysis-ready data is available to C-BRAIN's AI tools and research teams.
C-BRAIN is building an AI Biomedical Research Scientist platform that integrates diverse multi-institutional datasets (including NACC, ADNI, and consortium member data contributions). The Data Engineer will be central to building the technical infrastructure that makes this platform possible, working in close partnership with the CTO, the Senior Technical Product Manager, and external data science collaborators. This is not a standard data pipeline position.
The Data Engineer is building the technical backbone of an AI biomedical research platform — infrastructure that must ingest and harmonize multi-modal neurodegeneration datasets at consortium scale and serve as the data foundation for agentic AI tools including Insight Engine and Open Scientist. The ideal candidate brings software engineering discipline, strong cloud platform experience, and demonstrated knowledge of neurodegeneration or biomedical research data.
Domain knowledge is a prerequisite, not a nice-to-have; C-BRAIN-specific context will be provided, but neurodegeneration data experience and software engineering fundamentals will not.
Duties & Responsibilities
- Data Pipeline Development and Maintenance Designs, builds, tests, and maintains scalable data ingestion pipelines to ingest consortium member datasets from diverse sources and formats into the C-BRAIN data infrastructure. Develops and maintains ETL/ELT workflows using tools such as Apache Spark, dbt, Airflow, or equivalent; ensure pipelines are robust, well-documented, and auditable. Implements automated pipeline monitoring and alerting; troubleshoot and resolves pipeline failures in a timely manner.
Works collaboratively with the CTO and data science teams to understand data requirements for AI tool development and translates those requirements into technical pipeline specifications. Maintains version control for all pipeline code and infrastructure configurations; follows software engineering best practices including code review and documentation. Integrates and processes multi-modal data including omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), longitudinal clinical records, and digital pathology — reconciling differences in data type, format, spatial resolution, and dimensionality into unified analytical frameworks.
Identifies where cross-modal integration produces genuine signal versus where it introduces noise or artifact; establishes ground truth benchmarks for downstream AI use. - Data Infrastructure and Cloud Operations Manages and optimizes the C-BRAIN data infrastructure: storage accounts, computes resources, data lakes, and access controls. Implements and maintains data access controls and permissions aligned with DUA requirements and WashU data governance policies. Collaborates with the CTO on cloud architecture decisions; contributes to infrastructure planning for Phase 2 scale-up including foundation model compute requirements. Monitors infrastructure costs, resource utilization, and performance;
identifies and implements optimization opportunities. Supports the deployment of C-BRAIN AI tools on cloud-based platforms; coordinates with technical teams on infrastructure requirements. Ensures all data handling complies with DUA terms and applicable PHI de-identification requirements; implements, documents, and maintains de-identification workflows for each incoming dataset. Uploads curated datasets to ADDI/AD Workbench and other designated repositories (NIAGADS, GP2, or equivalent) as directed;
manages access controls within the platform to ensure data is accessible only by authorized users and tools. - Data Harmonization…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).