Senior Software Engineer; Data Platform
Listed on 2026-08-30
-
Software Development
Software Engineer, Backend Developer, Cloud Engineer - Software, DevOps
Location: Greater London
The role
We’re looking for a Software Engineer with strong C++ expertise to join the team building and operating Nebius Data Platform — a distributed storage and a processing platform that acts as the company’s “source of truth” and the backbone of many internal (and some external) products. Nebius Data Platform is a single multi-tenant ecosystem based on YTsaurus — instead of running separate HDFS/Kafka/HBase-style systems, we provide storage, compute, and analytics capabilities inside one platform.
Built on top of the open-source YTsaurus ecosystem, we run and extend our own Nebius distribution and develop significant in-house functionality (core and platform-level). We can design, implement, and roll out features end-to-end on our clusters without waiting for upstream approvals and contribute upstream when it makes sense. At scale today, this includes ~500 servers, ~20k CPU cores and ~10 PB of compressed data in our largest production cluster, supporting workloads ranging from business-critical pipelines and financial transactions to large-scale ML/LLM training datasets and compute.
inside the platform
You’ll work on a system that includes (and ties together):
Distributed Storage (Cypress): transactional semantics, tiered storage, erasure coding, replication, and strong reliability expectations. Compute & ETL: a cluster-wide job scheduler (tens of thousands of cores), Map Reduce, YQL for SQL-like data processing, and SPYT (Spark over YTsaurus) for modern data engineering. Interactive analytics (CHYT):
Click House® instances spun up directly on compute nodes for fast SQL over data in-place. Dynamic Tables: low-latency No
SQL KV with distributed ACID transactions for OLTP-style workloads and feature stores. Orchestracto: workflow orchestration deeply integrated with the platform (Airflow-like, but platform-native).
We’re looking for engineers who combine strong systems skills with product sense: understanding who uses the platform, why certain capabilities matter, and making pragmatic trade-offs to maximize impact. On our team, engineering work is expected to be connected to real users and outcomes — you’ll regularly align with internal stakeholders, clarify requirements, and help drive prioritization.
In this role, you will:- Design and implement new functionality in YTsaurus core (C++) with production reliability in mind.
- Build and evolve platform-level capabilities: platform architecture and operating model—multi-cluster growth, shared primitives, and a consistent experience that scales with new teams and use cases.
- Improve end-to-end platform experience for internal (and external-facing) users: APIs, guardrails, debugging workflows, and automation.
- Own production quality: incident response / on-call rotation, root cause analysis, and turning learnings into durable fixes.
- Roll out sharded YTsaurus masters (incl. Kubernetes operator support) and build automatic balancing of metadata across master cells (consensus groups) to remove control‑plane bottlenecks and unlock 10–100x cluster growth.
- Make CHYT interactive SQL faster and more predictable at high load via performance work like data‑skipping / min‑max‑style indexes and improved execution introspection.
- Turn Orchestracto into a platform product by defining the building blocks, developer experience, and governance for how teams create and share workflows.
- Scale and harden Parquet-on‑S3 for native YTsaurus workloads by tackling replication/movement, consistent lifecycle semantics, and master‑server metadata optimizations for performance and reliability.
- Design and ship complete, trustworthy audit trails for data changes (who/what/when) across heterogeneous storage and compute paths.
- Core: modern C++ (C++20, async + multithreaded primitives)
- Services & tooling:
Go and Python (microservices, utilities, integration tests)
5+ years of software engineering experience. Strong C++ skills (you’ll write core code). Working knowledge of Python and/or Go (you don’t have to be expert, but should be comfortable navigating them). Experience developing and/or operating high‑load, distributed services.…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: