Big Data Engineer
Listed on 2026-09-12
-
Software Development
Data Engineering, Python
ABOUT US:
Headquartered in the United States,
TP-Link Systems Inc. is a global provider of reliable networking devices and smart home products, consistently ranked as the world’s top provider of Wi‑Fi devices. The company is committed to delivering innovative products that enhance people’s lives through faster, more reliable connectivity. With a commitment to excellence, TP-Link serves customers in over 170 countries and continues to grow its global footprint.
We believe technology changes the world for the better! At TP-Link Systems Inc, we are committed to crafting dependable, high-performance products to connect users worldwide with the wonders of technology.
Embracing professionalism, innovation, excellence, and simplicity, we aim to assist our clients in achieving remarkable global performance and enable consumers to enjoy a seamless, effortless lifestyle.
KEY RESPONSIBILITIES- Build, run, and own ETL pipelines on EMR (Spark) orchestrated in Airflow — including their monitoring, alerting, and recovery.
- Write SQL and Python for data ingestion, transformation, and the datasets that analysts and dashboards depend on.
- Own data quality for what you build — run the checks before you ship and add new ones where they're missing; confirm the results are correct, not only that the job completed.
- Build and extend the data models for your area: fact and dimension tables following the team's layering conventions.
- Keep your jobs efficient — watch runtime and cost, and raise slow or expensive jobs rather than living with them.
- Use and extend the team's shared patterns and templates, and turn work you find yourself repeating into something automated or reusable.
- Work with analysts and business stakeholders to turn requests into datasets that actually get used.
- 2–4 years of hands‑on data development in a production environment — pipelines that run on a schedule with real downstream consumers, that you were responsible for including when they broke.
- Strong SQL: window functions, complex multi‑table joins, incremental loads, and the ability to work out why a query is slow.
- Python for production ETL and tooling (PySpark, pandas, boto3) — code that runs on a schedule, not only notebooks.
- Hands‑on experience with Spark on a cloud platform: AWS EMR, Databricks, Glue, or equivalent.
- Production experience with a scheduler, Airflow preferred: DAG design, dependencies, retries, and reruns that are safe to repeat.
- Working understanding of dimensional modeling: fact and dimension tables, star schema, and warehouse layering.
- Comfortable working on AWS (S3 with Parquet / ORC), Linux, and Git — and comfortable picking up new tools as the platform evolves.
- Effective use of AI to solve data problems — using it to move faster on SQL, debugging, and unfamiliar schemas, with the judgment to catch output that looks right but isn't.
- Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.
- Data ingestion or CDC tooling:
DataX, Sqoop, Debezium, Fivetran, or similar. - An OLAP / MPP engine:
Star Rocks, Doris, Click House, Redshift, or similar. - Spark performance work: partitioning, shuffle, skew, and memory tuning.
- AWS cost optimization: EMR instance sizing and Spot strategy, S3 lifecycle policies.
- Lakehouse formats (Iceberg, Hudi, Delta), streaming (Kafka, Flink), dbt, or a data quality framework.
- Quick Sight or another BI tool: dataset and permission design.
- Self‑directed learning: a project, an open‑source contribution, or a tool you picked up on your own and put to real use
Base Salary Range: $100 - 120K
- Free snacks and drinks
- Fully paid medical, dental, and vision insurance (partial coverage for dependents)
- Contributions to 401K funds
- Bi‑annual reviews, and annual pay increases
- Health and wellness benefits, including free gym membership
- Quarterly team‑building events
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).