Sr. Data Platform Engineer
About Tripstack
Founded in Toronto, Canada in 2016, Tripstack has been part of Etraveli Group since 2019. It is a B2B Flights as a Service provider and a world leader in virtual interlining. Operating from offices in Canada, India, and Poland, Tripstack is the gateway into Etraveli Group’s world leading tech platform – giving partners access to global flight content, virtual interlining, and a full suite of services including payments, fraud prevention, pricing, and customer support.
Its technology ingest over 30B price points and handles over 240 million searches daily.
Through partnerships with airlines, OTAs, and other distribution channels across the globe, Tripstack expands networks, drives new revenue streams, and offers more choice at competitive prices, all backed by robust technology and traveler protection.
The RoleTripstack is moving its entire data stack from bare‑metal VMs to Kubernetes on Open Stack in a new data centre. We are looking for a senior infrastructure engineer who has done stateful migrations before, who can plan, codify, and execute this one safely, and who will own the day‑to‑day operational health of the data platform once we are there. This is a hands‑on platform and SRE role with a clear, time‑bounded mission.
You will partner closely with our SRE team on networking, hardware, and Kubernetes fundamentals, and own the data applications — Druid, Spark, Redpanda, Airflow, PostgreSQL, Elasticsearch — end‑to‑end. It is not an ML role. We have a separate plan for evolving our MLFlow platform, and the right hire here may grow into more of that work over time, but day‑one impact is the migration and the operational health of the platform.
Lead the Data‑Stack Migration
- Plan and execute the migration of Druid, Spark, Redpanda, and our orchestration layer from bare‑metal VMs to Kubernetes on Open Stack, with no downtime on stateful workloads.
- Design Stateful Set, PVC, pod‑disruption‑budget, and rolling‑upgrade patterns that are safe for production data systems.
- Codify the migration with Infrastructure as Code — Terraform for Open Stack, Helm or Kustomize for Kubernetes, Git Ops via ArgoCD or Flux — so the result is reproducible and supportable by the whole team.
- Own the operational health of Druid, Spark, Redpanda, Airflow, PostgreSQL, and Elasticsearch as production systems — segment lifecycle, JVM tuning, ingestion specs, broker/coordinator/overlord internals, partition design, consumer lag, replication tuning.
- Build the KPIs, alerting, dashboards, and runbooks that let us see cluster exhaustion before it becomes an incident, and diagnose it quickly when it does.
- Own the query, report, segment, and tiering optimisations that keep our analytics cost‑effective and responsive under load.
- Build the Prometheus, Grafana, and distributed‑tracing coverage our data systems need. Treat SLOs, error budgets, and post‑incident discipline as table stakes.
- Partner with SRE on hardware, networking, and Kubernetes fundamentals — but own the data applications themselves end‑to‑end.
- Strong Kubernetes experience with stateful workloads — Stateful Sets, PVCs, pod disruption budgets, and rolling upgrades for data systems. You have done a real stateful migration before and can talk through what went wrong.
- Infrastructure as Code at a senior level — Terraform, Helm or Kustomize, Git Ops with ArgoCD or Flux. You have shipped production infrastructure this way, not just experimented with it.
- Observability and RCA discipline — Prometheus, Grafana, distributed tracing, SLOs, error budgets, and the habit of writing the runbook that stops the next incident.
- Production operations experience with at least one of Apache Druid, Apache Kafka or Redpanda, Apache Spark, or Elasticsearch — deep enough to be credible on internals and willing to learn the others.
- 7+ years building and operating production data or platform systems, at least 2 of them on self‑hosted or bare‑metal infrastructure. You have been on‑call for what you built.
- Clear written and verbal English; comfortable working across Kraków, Toronto, Pune, and…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: