Data Engineer
Listed on 2026-10-08
-
IT/Tech
Data Engineering
The Data Engineer is responsible for designing, building, and maintaining the data pipelines and integrations that power GEI’s AI solutions and digital initiatives. This role focuses on ensuring enterprise data is accessible, reliable, and governed so that AI capabilities can be deployed and scaled with confidence.
The Data Engineer plays a hands‑on role by preparing and integrating the data foundations that AI solutions depend on. This includes building ingestion pipelines, managing data stores, implementing quality and governance controls, and supporting retrieval patterns such as RAG. This role works closely with AI Engineers, solution architects, and platform teams to ensure data infrastructure is production-ready, secure, and aligned with GEI standards.
- Design, build, and maintain data pipelines that ingest, transform, and deliver enterprise data to AI solutions and business applications.
- Develop and manage integrations across enterprise data sources using APIs, Graph connectors, event‑driven architectures, and batch/streaming patterns.
- Build and maintain data stores and indexing infrastructure that support retrieval‑augmented generation (RAG) and other AI consumption patterns.
- Implement data quality, validation, and lineage controls to ensure accuracy and trustworthiness of data feeding AI workflows.
- Support and optimise data models underpinning Power BI dashboards and AI‑enabled analytics.
- Collaborate with AI Engineers to define data contracts and ensure pipeline outputs meet solution requirements for schema, latency, and freshness.
- Instrument data pipelines for monitoring, alerting, cost control, and performance optimisation.
- Implement data governance and security controls including access management, encryption, and compliance with organisational data policies.
- Identify data‑related risks and support mitigation strategies in collaboration with architecture and platform teams.
- 4+ years of data engineering experience, with demonstrated ability to build and operate production data pipelines.
- Proficiency in Python and SQL; experience with PySpark or Spark is strongly preferred.
- Experience with Azure data services, including Azure Data Factory, Azure Synapse Analytics, Azure Data Lake Storage, and Azure SQL.
- Familiarity with Azure AI Search, Cosmos DB, or similar services used to support AI retrieval and storage patterns.
- Experience building and managing ETL/ELT pipelines with structured, semi‑structured, and unstructured data sources.
- Knowledge of data modeling, schema design, and indexing strategies for both analytical and AI workloads.
- Familiarity with infrastructure‑as‑code and CI/CD practices for data pipeline deployment (e.g., Terraform, Azure Dev Ops).
- Knowledge of data governance principles, including data cataloging, lineage, access control, and privacy requirements.
- Knowledge of security best practices for data solutions, including encryption at rest and in transit, role‑based access control, and private networking.
- Experience with Databricks, including Delta Lake and Unity Catalog, is a plus.
- Prior experience in professional services, engineering, or construction environments is a plus.
- Azure or Databricks certifications (e.g., Azure Data Engineer Associate, Azure Solutions Architect Expert, Databricks Data Engineer Professional) are a plus.
Some of the world’s most pressing problems – from climate change to sustainable development, to critical infrastructure and the future of our energy supply – need our brightest and diverse minds working together to create safer, more resilient communities for tomorrow.
We are technical experts, collaborators, and entrepreneurs who draw from diverse backgrounds to solve our clients’ most complex challenges.
With several…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).