Member of Technical Staff - Data Engineering
Verfasst am 2026-10-05
-
Software Entwicklung
Dateningenieur, Datenwissenschaftler, Maschinelles Lernen
About us
Albs is an AI research lab building real-time multimodal intelligence for machines, enabling them to see, hear, reason, and interact. We treat model architecture, inference, and runtime as one system. Our purpose-built models and optimized runtimes unlock the full potential of each device within defined limits for hardware cost, power consumption, and response time. Companies can adapt our technology to their own machines without building the underlying AI from scratch.
We founded Albs at the intersection of LLM architecture research and on-device engineering, and we are currently in stealth, but well funded, with dedicated compute for large-scale training and experimentation. Publishing is a core part of our research culture, and we contribute our work to top venues. We share more details about the company, the team and our backing in the first conversation.
Who we're looking forThis is a Staff / Senior IC role. We are looking for experienced researchers and engineers, typically with several years of research or industry experience or an equivalent track record. The exact scope of each role depends on your background: some people go deep on one part of the stack, others shape the technical direction of a whole area. We agree on scope together with you during the interview process.
We welcome applications from all qualified candidates, regardless of gender, age, ethnic origin, religion, disability or sexual orientation.
Build the data platform behind our models: ingestion, storage and processing for pre-training, post‑training and evaluation. For small on‑device models, data quality is one of the biggest levers on model quality.
Build and scale pipelines for petabyte‑scale text corpora: crawling, extraction, deduplication, model‑based quality filtering and decontamination against our benchmarks.
Build synthetic data and rephrasing pipelines together with the research team, and measure whether they actually help.
Make data reproducible and compliant: versioning, lineage, dataset cards, licensing and PII handling, so every training run can be traced back to the exact data it saw.
Keep the GPUs fed: high‑throughput data loading and streaming for multi‑node training, together with the pre‑training team.
Build dataset visualization and QA tooling so researchers can inspect, slice and compare data before they train on it.
Work directly with the research team on data mixtures and ablations, and turn data hypotheses into experiments that run within days.
You have several years of experience building large‑scale data systems, in industry or research.
You write strong Python, plus e.g. Rust or C++.
You have run distributed data processing (Spark, Ray, Dask, Beam or similar) on terabyte to petabyte datasets, with object storage and columnar formats (Parquet, Arrow), and you know how to make it fast and cheap.
You have built data pipelines for ML training, ideally for LLM pre‑training, including tokenization and GPU‑based filtering or scoring models.
You care about data quality as much as throughput, and you write tests and monitoring for your pipelines.
Bonus: experience with multimodal data, synthetic data generation, or data governance under GDPR and the EU AI Act.
We are a small, focused team of experts. You would join early, work directly with the founders, and help shape how we build. Fast iteration, short lines of communication, and in‑person discussion matter a lot to us. Our culture is built around the office in Freiburg, Germany, with a default of three days a week on site. Alternatively, you can work remotely and join us on a regular cadence.
We will discuss what works best for you during the interview process.
We hold ourselves…
(Wenn dieser Job tatsächlich in Ihrem Zuständigkeitsbereich liegt, verwenden Sie möglicherweise einen Proxy oder VPN, um auf diese Seite zuzugreifen. Um weiterzukommen, sollten Sie Ihre Verbindung zu einem anderen Mobilgerät oder PC wechseln).