Data Engineer
Listed on 2026-09-14
-
Software Development
Data Engineering
Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired bya collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world.
Job DescriptionWe are seeking a highly skilled GCP Data Engineer with strong Python expertise to design, build, and optimize scalable data solutions on Google Cloud Platform (GCP). The ideal candidate will have hands‑on experience developing batch and real‑time data pipelines, working with large‑scale datasets, and enabling analytics and AI/ML use cases.
Key Responsibilities Data Engineering & Pipeline Development- Develop and optimize ETL/ELT workflows for structured and unstructured data processing using GCP services such as Dataflow, Dataproc, and Pub/Sub
- Implement event-driven data processing using Cloud Functions and Pub/Sub
- Build and manage data ingestion frameworks for streaming and batch data sources
- Design and optimize data lakes and data warehouses using Big Query and Cloud Storage
- Develop efficient data models to support analytics, reporting, and machine learning workloads
- Optimize performance and cost of data pipelines and queries
- Develop solutions using Python
- Automate workflows and orchestration using Cloud Composer (Airflow)
- Implement CI/CD pipelines and deployment automation
- Collaborate with analytics, AI/ML, and business teams for data consumption needs
- Troubleshoot data issues and perform root cause analysis
- Continuously improve pipeline reliability, scalability, and performance
- 5+ years of overall data engineering or software engineering experience
- 2+ years of hands‑on Google Cloud Platform experience
- 2+ years of Python development
- 2+ years of experience building data pipelines (batch and streaming)
- Experience with Dataproc (Spark/PySpark) for large-scale processing
- Familiarity with event-driven architectures
- Knowledge of Terraform or Infrastructure as Code
- Understanding of cost optimization (Fin Ops)
- Google Cloud Professional Data Engineer Certification
- Experience supporting AI/ML data pipelines
The base compensation range for this role in the posted location is: 75000 to 96000
Capgemini provides compensation range information in accordance with applicable national, state, provincial, and local pay transparency laws. The base compensation range listed for this position reflects the minimum and maximum target compensation Capgemini, in good faith, believes it may pay for the role at the time of this posting. This range may be subject to change as permitted by law.
The actual compensation offered to any candidate may fall outside of the posted range and will be determined based on multiple factors legally permitted in the applicable jurisdiction.
These may include, but are not limited to:
Geographic location, Education and qualifications, Certifications and licenses, Relevant experience and skills, Seniority and performance, Market and business consideration, Internal pay equity.
It is not typical for candidates to be hired at or near the top of the posted compensation range.
In addition to base salary, this role may be eligible for additional compensation such as variable incentives, bonuses, or commissions, depending on the position and applicable laws.
Capgemini offers a comprehensive, non‑negotiable benefits package to all regular, full‑time employees.
- Paid time off based on employee grade (A-F), defined by policy:
Vacation: 12-25 days, depending on grade, Company paid…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).