More jobs:
Job Description & How to Apply Below
Build, and optimize scalable data pipelines and ETL/ELT workflows for large, complex datasets
Design and implement foundational data architecture supporting identity resolution and systems
Develop and enhance systems supporting identity resolution and construction (data ingestion, normalization, matching, and deduplication)
Process and unify multi-source datasets (cookies, device IDs, behavioral data, third-party and proprietary data)
Write efficient, testable, and maintainable code using Python and SQL for large-scale data processing
Optimize data models, queries, and storage strategies for performance, scalability, and cost efficiency
Build and maintain data validation, monitoring, and alerting systems to ensure data quality and reliability
Troubleshoot, debug, and improve existing data pipelines and infrastructure
Own and drive complex data problems end-to-end, from initial design through production deployment
Make and influence key technical decisions related to data architecture, scalability, and system design
Collaborate with data, platform, Dev Ops, and product teams to deliver scalable, production ready solutions
Translate business and product requirements into practical, performant data solutions
Document data pipelines, systems, and workflows clearly
Continuously improve system performance, data quality, and pipeline resilience.
Contribute to building new capabilities that improve how customers understand and leverage data insights.
Key Requirements:
8-12+ years of hands-on experience in data engineering or large-scale data processing
Proven experience building and maintaining production-grade data pipelines and distributed systems
Demonstrated experience architecting and delivering large-scale data platforms or mission critical data systems
Strong expertise in: o SQL and relational databases (Postgres, Big Query, Redshift, etc.) Python for data processing and analysis
Experience with Google Cloud Platform (Big Query, Dataflow, Pub/Sub, Cloud Storage, Cloud Functions) and/or AWS (S3, Redshift, EMR, RDS)
Experience working with large-scale datasets (hundreds of millions to billions of records)
Strong understanding of data modeling, partitioning, indexing, and query optimization
Experience with distributed data processing and parallelization techniques
Experience moving large volumes of data across systems and architectures
Familiarity with CI/CD, containerization, and orchestration tools (Docker, Kubernetes, Git Hub Actions, etc.)
Strong debugging and troubleshooting skills in complex data environments
Experience with version control (Git) and Agile tools (Jira, Confluence, etc.)
Highly analytical with strong attention to detail and a data-driven mindset
Ability to hit the ground running, quickly understand systems, and deliver independently
Comfortable working in a remote, fast-paced, and collaborative environment
Proven ability to drive system design and implementation.
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×