Data Engineer II - OIT
Listed on 2026-07-24
-
IT/Tech
Data Engineering, Data Warehousing, Data Analyst
Discover Your Career at Emory University
Emory University is a leading research university that fosters excellence and attracts world-class talent to innovate today and prepare leaders for the future. We welcome candidates who can contribute to the excellence of our academic community.
Data Engineer IIThe Data Engineer II is a key contributor to the development and evolution of the Unified Data Platform (UDP), built on Azure Fabric. This role works closely with researchers, architects, and cross-functional teams to design and deliver scalable, high-quality data solutions that support research, analytics, and clinical insights. The position involves working with complex datasets, modern cloud data platforms, and emerging technologies.
While experience with Azure Fabric is preferred, candidates with strong hands-on experience in Azure Synapse or Databricks with PySpark will be equally considered. This role requires a strong foundation in data engineering, biomedical informatics principles, data governance, and compliance standards such as HIPAA. The ideal candidate is a collaborative team player who can translate business and research requirements into efficient and scalable technical solutions.
The following are key responsibilities:
- Collaborate as a core member of cross-functional teams including Business Analysts, Project Managers, Data Analysts, and Architects to deliver high-quality solutions within scope and timeline.
- Work directly with researchers and stakeholders to gather, analyze, and translate requirements into technical solutions.
- Design, develop, and maintain scalable data pipelines, ETL/ELT processes, and data integration workflows across multiple systems.
- Build and optimize data solutions that integrate data from disparate sources, ensuring data quality, integrity, and consistency.
- Evaluate emerging technologies and develop proof-of-concepts to support innovation and continuous improvement.
- Apply biomedical informatics standards, methodologies, and principles to research data solutions.
- Ensure adherence to HIPAA and institutional data governance policies and standards.
- Develop and maintain metadata, data standards, and data governance processes for complex datasets.
- Create and maintain clear technical documentation, including data pipelines, workflows, and system designs.
- Communicate technical concepts effectively to both technical and non-technical stakeholders.
- Partner with data and system architects and demonstrate an understanding of data modeling concepts and best practices.
- Manage workload effectively and provide timely updates on task progress and deliverables.
- Works as a positive team member of a project that may consist of Business Analysts, Project Managers, Information Architects, Data Analysts, and/or Database Administrators to deliver quality applications and components within scope, on time, and within budget.
- Manages workload effectively and report status of tasks in a timely manner.
- Works directly with researchers to document, analyze, and translate their needs into technical designs and informatics solutions.
- Participates in the evaluation of emerging technologies and develops proof-of-concepts.
- Contributes to technical teams.
- Follows standard operational procedures and HIPAA regulations.
- Develops strategies for managing complex data sets through maintaining data standards and metadata.
- Applies biomedical informatics technical standards, methodologies, and principles to research-specific program needs, objectives, and outcomes.
- Develops complex reports, data pipelines, and ETL processes from disparate systems and ensures their accuracy.
- Gathers user requirements and creates technical documentation
- Performs other related duties as required.
Minimum qualifications include a bachelor's degree in a related field and three years of related experience, or an equivalent combination of education, training, and experience.
Preferred qualifications include:
- Strong experience designing and maintaining data pipelines and ETL/ELT processes
- Proficiency in Python, SQL, and Apache Spark (PySpark)
- Experience with modern data platforms such as Azure Fabric (preferred), Azure Synapse Serverless, or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).