More jobs:
Data Quality Engineer
Job in
Dallas, Dallas County, Texas, 75215, USA
Listed on 2026-05-03
Listing for:
Select Minds LLC
Full Time
position Listed on 2026-05-03
Job specializations:
-
IT/Tech
Data Engineer, Big Data
Job Description & How to Apply Below
Benefits
- W2 OPPORTUNITY
- Competitive salary
- Opportunity for advancement
Job Title: Data Quality Engineer (Databricks, Kafka, AWS)
Location: Dallas, TX (Hybrid – 3 days onsite)
Job Type: Long-term Contract
Work Authorization: Open - W2 opportunity
We are looking for a Data Quality Engineer to own validation across batch and streaming data pipelines. This role focuses on ensuring data correctness, reliability, and performance across platforms built on Databricks, Kafka, AWS, SQL, and Python.
This is a hands‑on role focused on building scalable data validation frameworks and ensuring production‑grade data systems.
Key Responsibilities End-to-End Data Validation- Validate data pipelines for accuracy, completeness, consistency, and timeliness
- Build SQL-based validations for business rules and transformations
- Implement reconciliation between source and downstream systems
- Ensure data lineage and traceability
- Test pipelines built on AWS (Glue, Lambda, EMR, Step Functions)
- Validate transformations using SQL and Python
- Test ingestion, transformation, aggregation, and serving layers
- Handle backfills, reprocessing, and historical data loads
- Validate Spark pipelines (PySpark/Scala) on Databricks
- Validate data integrity, ordering, and delivery guarantees
- Test producer and consumer logic and serialization formats (Avro, JSON, Protobuf)
- Validate topics, partitions, offsets, retention, and schema evolution
- Simulate late events, duplicates, and failure scenarios
- Build Python-based data testing frameworks
- Develop reusable validation utilities and synthetic datasets
- Integrate data tests into CI/CD pipelines
- Enable automated alerts for data quality issues
- Validate throughput, latency, and concurrency at scale
- Test retry logic, idempotency, and recovery mechanisms
- Perform regression, soak, and failover testing
- Validate logs, metrics, and alerts using tools such as Cloud Watch, Prometheus, and Grafana
- Define and monitor data SLAs and SLOs
- Support incident response, root cause analysis, and postmortems
- 7+ years of total experience in QA, SDET, or Data Quality Engineering
- Minimum 4–6 years of hands‑on experience working with data platforms, data pipelines, or data engineering ecosystems
- 3+ years of hands‑on experience with Databricks and Apache Spark
- Strong SQL skills for data validation, reconciliation, and complex analysis
- Proficiency in Python for automation and data validation
- Experience testing ETL/ELT pipelines (batch and streaming)
- Hands‑on experience with Kafka or similar streaming platforms
- Strong understanding of AWS data services (S3, Glue, Lambda, Redshift, Athena)
- Experience working with large-scale distributed data systems
- Strong debugging, analytical, and problem‑solving skills
- Experience with data quality or observability tools such as Great Expectations or Monte Carlo
- Knowledge of schema registry and data contracts
- Experience with CI/CD tools such as Git Hub Actions or Jenkins
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×