More jobs:
Job Description & How to Apply Below
Location:
Bangalore (Yamlur)
Experience:
6-10 Years
Role: Principal Data Engineer
About Styli Marketplace
Launched in 2019 by Landmark Group, Styli Marketplace is the first e-commerce venture of the group, quickly becoming a leading online destination for fashion and lifestyle across the GCC, including Saudi Arabia, the UAE, Kuwait, Bahrain, and beyond. Styli connects global sellers and creators with millions of fashion-forward customers, offering the latest trends, exceptional value, and convenient services like same-day to 48-hour delivery and flexible payment options.
Our mission is to make style accessible, aspirational, and exciting for all, backed by a passionate team fostering a culture of creativity and innovation.
Role Overview
We are looking for a Principal Data Engineer to take technical ownership of Styli's Data Platform and to lead its evolution from a Big Query-centric warehouse to a modern Lakehouse on Databricks . This is a hands-on architecture and leadership role: you will set the technical direction for the platform, drive the migration end-to-end (design, staffing, partner management, cutover), and mentor a team of senior and staff data engineers.
You will work closely with Platform/Dev Ops, Security, ML, and business stakeholders across the GCC and India to ensure the platform scales reliably, cost-efficiently, and securely as Styli grows.
What You'll Do
Data Platform Architecture & Migration Leadership
Own the end-to-end target-state architecture for Styli's data platform, leading the migration from Big Query to a Databricks Lakehouse built on Delta Lake / Apache Iceberg, Unity Catalog, and GCS.
Lead build-vs-buy and platform evaluations (e.g., Databricks vs. Snowflake), and translate the recommendation into an executable, phased migration roadmap with clear cutover and rollback criteria.
Define and defend architecture decision records (ADRs) covering table format strategy, catalog ownership, and cross-engine interoperability between Databricks and existing warehouses via the Iceberg REST catalog.
Own the staffing and delivery model for the migration — partnering with Databricks implementation/SI partners, structuring hybrid onsite/offshore teams, and tracking cost, timeline, and risk.
Drive Unity Catalog adoption for centralized governance, fine-grained access control, and data lineage across work spaces and markets.
Data Pipeline Development
Design, build, and maintain batch and real-time pipelines migrating from Airflow/Big Query patterns to Databricks Workflows and Delta Live Tables, alongside Apache Airflow and dbt where appropriate.
Handle ingestion from diverse sources:
MySQL/PostgreSQL, Kafka event streams, S3/GCS data lakes, REST APIs, and SaaS platforms.
Ensure pipelines are fault-tolerant, recoverable, and instrumented with data quality checks at every stage, with clear migration parity/reconciliation testing against legacy Big Query pipelines.
Data Modelling & Lakehouse Engineering
Design and own the medallion (bronze/silver/gold) data architecture on Delta Lake, including dimensional and analytical models (star schema, OBT, wide tables) for BI and ML consumption.
Own warehouse/lakehouse performance engineering — partitioning, Z-ordering/clustering, Photon acceleration, and cost-efficient query and compute (cluster policy) patterns on Databricks, alongside the legacy Big Query estate during transition.
Set and enforce standards for dbt models, tests, documentation, and lineage across a well-governed transformation layer.
Streaming & Real-Time Data
Build and maintain real-time pipelines for live inventory updates, order event processing, and customer behavior streams, using Databricks Structured Streaming, Apache Flink, Spark Streaming, or ksqlDB.
Design Kafka topic schemas, partition strategies, and consumer group management for high-throughput e-commerce event flows.
Data Lake & Cloud Infrastructure
Build and manage a well-structured lakehouse on GCS, with Databricks running natively on GCP alongside existing cloud-native data services.
Partner with the Platform/Dev Ops team on infrastructure-as-code (Terraform) for data infrastructure provisioning and containerized data…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×