Lead AI Data Engineer; Hybrid
Bethesda, Montgomery County, Maryland, 20811, USA
Listed on 2026-09-12
-
IT/Tech
Data Engineering
About The Position
The Lead AI Data Engineer serves as a scientific, technical, and managerial lead for data-heavy research projects. The Lead AI Data Engineer will overlap with a multidisciplinary team of government and contract researchers, academic experts and consumers. The Lead AI Data Engineer will provide technical and management support, oversee project execution, and provide key guidance on data architecture and infrastructure. The team that the Lead AI Data Engineer oversees is responsible for the curation and maintenance of large data pipelines, developing ETL pipelines, defining schemas, identifying bottlenecks, and deriving variables /features directly used in machine learning / AI model development.
The Lead AI Data Engineer serves as a versatile position who designs, builds, and maintains the entire data ecosystem. As project aims evolve, collaborators join, and team processes change, an ideal candidate is flexible and can adapt quickly. The ideal candidate will also be mindful of data privacy / security and will adhere to data governance and data management best practices.
This is a full time hybrid remote position working at the Walter Reed National Military Medical Center in Bethesda, MD that will require working in office/on site at least 2 days per week. Background checks will be administered.
Sleep Physiology Modeling Project:
Sleep & Wearables Operational Readiness for Research & Defense (SWORD) Lab
Salary Range $155,000 - $193,000.
Qualifications- PhD in a relevant field (e.g., Computer Science, Engineering, Data Science, Biomedical Engineering) required
- 8+ years experience with multimodal data analysis and data pipeline engineering required
- Proven experience with multivariate signal processing (e.g., time-series biosensor data)
- Hands-on experience with relational (SQL) and non-relational (No
SQL) databases - Hands-on experience with version control systems (eg, Git) and demonstrated ability to work in (and lead) a collaborative coding environment
- Solid problem-solving and analytical skills to address complex technical challenges
- Hands-on experience using Google Cloud Platform (GCP) cloud infrastructure (or equivalent), including setting up and managing cloud-native data warehouses (eg, Big Query), storage, and compute resources
- Ability to translate high-level scientific hypotheses into scalable engineering solutions and data products
- Ability to work in a fast-paced, multidisciplinary, multi-site (sometimes asynchronous) team environment
- Strong knowledge of sleep science and hands-on experience with handling data from consumer wearable devices (eg, actigraphy, PPG, EEG); familiarity with machine learning workflows, including model development, model tuning, and deploying models at scale;
- Leadership and/or project management experience with the ability to oversee a team of people ingesting data
Foster a collaborative coding and research environment, driving skill development for junior and mid-level data engineers and analysts across the data pipeline;
Serve as the primary technical liaison to senior management, translating high-level research aims into actionable objectives;
Communicate team progress, bottlenecks, and milestones.
- Produce clean, well-documented, efficient code across the entire stack
- Design, develop, and deploy robust, scalable applications (both front-end interfaces and back-end data pipelines) to support large-scale research
- Lead the engineering workflows to acquire, ingest, and clean multimodal datasets, ensuring efficient storage, retrieval, and processing of massive datasets (+1million records)
- Architect and maintain scalable infrastructure to support advanced machine learning…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).