Data Scientist Pulmonary Research
Listed on 2026-09-05
-
IT/Tech
Data Scientist, Data Analyst, Data Engineering
Site:
The Brigham and Women's Hospital, Inc.
Mass General Brigham relies on a wide range of professionals, including doctors, nurses, business people, tech experts, researchers, and systems analysts to advance our mission. As a not-for-profit, we support patient care, research, teaching, and community service, striving to provide exceptional care. We believe that high-performing teams drive groundbreaking medical discoveries and invite all applicants to join us and experience what it means to be part of Mass General Brigham.
Job SummaryThe position will work directly in collaboration with Dr. Shappell, a critical care physician researcher with expertise using data from Electronic Health Records (EHRs) to study the epidemiology and outcomes of critical illness.
The position will primarily be responsible for the creation, maintenance, and ongoing operations of the new MGB Critical Care Data Repository, which aims to provide highly granular research-ready ICU data to MGB investigators for clinical operations, quality improvement, and research projects. This will include working with existing MGB analysts to complete set up of EHR data pipelines and databases using Clarity and MGB's cloud-computing environment, Snowflake;
collaborating with clinical experts on data mapping and cleaning; performing in-depth data quality assurance including the production of visualizations and reports; and providing requested data to PCCM Division leadership as needed.
In addition, the Data Scientist will support Dr. Shappell's health services research on critical illness including sepsis, shock, severe viral infections, and respiratory failure; they will also support the full integration of the MGB data repository into an established critical care research consortium via application of a novel common data model. This work will involve completing and ensuring compliance with latest version of common data model, developing and coding new research projects that utilize consortium data, running code locally from other consortium PI-led projects, and attending weekly consortium meetings.
There is opportunity for expansion into work with other EHR data modalities, including unstructured data from clinical notes, waveform data, and/or imaging data, if desired.
Summary:
Responsible for analyzing data, uncovering the underlying data patterns and logic, and developing data-driven applications. They will work towards developing solutions for the entire problem-solving cycle: find/prioritize the problems, research the best algorithms to solve the problems, design robust, practical solvers, and implement them.
Does this position require Patient Care? No
Essential Functions- Analyze complex and high-dimensional clinical and operational data for dependencies, patterns, outliers, inaccuracies, and validity.
- Apply knowledge of statistics, machine learning, programming, data modeling, simulation, and/or advanced mathematics to recognize patterns, identify opportunities, pose business questions, and make valuable discoveries.
- Contribute to the design and evaluation of optimal data-driven metrics (such as physician/facility performance criteria, bottleneck metrics, productivity limits, processing delay/error reports, etc.).
- Use a flexible, analytical approach to design, develop, and evaluate predictive models and advanced algorithms that lead to optimal value extraction from the data.
- Generate and test hypotheses and analyze and interpret the results.
- Leads creation and maintenance of databases in Clarity's cloud-computing environment, Snowflake, using SQL.
- Data cleaning and management. Prepares analytic datasets using data analysis software (E.g. R tidyverse packages, Pandas for Python) from electronic health care record data sets.
- Follows the standards of transparent and reproducible data science including use version control software such as git and code with a literate programming style.
- Anticipates data management needs, including identifying when data pulls will be necessary and obtaining newer versions of datasets.
- Data visualization with appropriate statistical software, such as the ggplot2 package in R.
- Reports results of data analysis, prepares figures and tables.
- Supports PI and trainees in the drafting and revision of manuscripts and grant proposals.
- Collaborates effectively with local and external PIs and data analysts/scientists.
- Performs other related work as needed.
Education
Bachelor's Degree Related Field of Study required or Master's Degree Related Field of Study preferred
Can this role accept experience in lieu of a degree?
Yes
Experience working in data science-type positions and with large data sets 2-3 years preferred
Knowledge,Skills and Abilities
- Technical expertise with SQL, Stata, R, Python. Experience with database creation and engineering.
- Ability to create reports and dashboards with Tableau.
- Skilled in data analysis with Python and SQL.
- Knowledge of statistics.
- Knowledge of machine learning.
- Preferred SQL…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).