×
Hier anmelden um sich kostenlos auf Stellen zu bewerben oder Stellenanzeigen aufzugeben. X

Senior Machine Learning Ops Engineer

in 10115, Berlin, Berlin, Deutschland
Unternehmen: United States Digital Space LLC
Vollzeit position
Verfasst am 2026-08-13
Berufliche Spezialisierung:
  • IT/Informationstechnik
    Maschinelles Lernen, Site Reliability Ingenieur/in, Künstliche Intelligenz Ingenieur, Cloud Computing: IT-Infrastruktur & Betrieb
Gehalts-/Lohnspanne oder Branchenbenchmark: 120000 - 150000 EUR pro Jahr EUR 120000.00 150000.00 YEAR
Stellenbeschreibung

the company, part of Booking Holdings (NASDAQ: BKNG), is a leading travel search engine. With billions of queries across our platforms, we help people find their perfect flight, stay, rental car and vacation package. We're also transforming business travel with a new corporate travel solution, the company for Business.

As an employee of the company, you will be part of a travel company that operates a portfolio of global metasearch brands including momondo, Cheap flights and Hotels Combined, among others. From start-up to industry leader, innovation is in our DNA and every employee has an opportunity to make their mark. Our focus is on building the best travel search engine to make it easier for everyone to experience the world.

Every machine learning model the company ships depends on reliable, scalable infrastructure to move from experiment to production — and that's exactly what this role makes possible. the company is seeking a Senior MLOps Engineer who will focus on the design and implementation of our machine learning infrastructure and production lifecycle. This is a senior, hands-on role where you will bridge the gap between data science and production engineering.

You will join the Machine Learning Platform team and be responsible for building and maintaining scalable infrastructure & automated pipelines for model training, deployment, and monitoring, ensuring our ML models are reliable, reproducible, and performant. You will work closely with Data Scientists, ML Engineering and Operations teams to transform experimental code into robust, production-ready services at scale.

This role requires commuting to the Berlin office 3 times a week.

In this role, you will:

  • Build and maintain ML infrastructure end-to-end:
    Extend and operate the infrastructure that powers every model we ship — including CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.
  • Own model deployment and serving:
    Help define and evolve the standards and tooling for model serving, ensuring low latency and high availability across our ML services.
  • Develop core MLOps capabilities:
    Establish and maintain essential infrastructure that functions as reliable, self-service systems for the entire machine learning organization — with a focus on feature stores, model registries, and automated monitoring for performance and data drift.
  • Operationalize infrastructure for the ML team:
    Collaborate with Operations to enable Kubernetes (k8s) autoscaling and GPU provisioning, turning these into accessible, self-service tools for ML practitioners — including standing up and operating a Kubernetes-based development cluster and taking models from experimentation to GPU-backed production.
  • Improve platform reliability and performance:
    Partner with Operations to design resilient monitoring using advanced observability tooling. Define service-level objectives and implement automation to reduce manual interventions and improve system reliability.
  • Empower Data Scientists through standardized, optimized workflows:
    Amplify the impact of the ML team by building clear, well-supported "golden paths" — standardized workflows that streamline the model development lifecycle and let Data Scientists focus on modeling while you handle the infrastructure.

Please apply if you have:

  • Experience building and operating ML platforms in production environments.
  • Solid working knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale.
  • Familiarity with ML lifecycle tooling, including orchestration frameworks, feature stores, model registries, and drift or performance monitoring.
  • Experience owning production systems: defining service-level objectives (SLOs), building observability (for example, using tools such as Prometheus, Grafana, or Datadog), participating in incident response, and diagnosing large-scale failures systematically. You look for opportunities to automate repetitive work rather than absorb it.
  • Comfort writing production-quality code in Python or a comparable language.
  • Experience modernizing production infrastructure with attention to…
Stellen-Anforderungen
10+ Jahre Berufserfahrung
Bitte beachten Sie, dass derzeit keine Bewerbungen aus Ihrem Zuständigkeitsbereich für diese Stelle über diese Jobseite akzeptiert werden. Die Präferenzen der Kandidaten liegen im Ermessen des Arbeitgebers oder des Personalvermittlers und werden ausschließlich von diesen bestimmt.
Um nach Stellen zu suchen, sie anzusehen und sich zu bewerben, die Bewerbungen aus Ihrem Standort oder Land akzeptieren, klicken Sie hier, um eine Suche zu starten:
 
 
 
Suchen Sie hier nach weiteren Stellen:
(nach Beruf, Fähigkeit)
Standort
Suchradius erweitern (Meilen)
0
200
Filter
Mindest-Bildungsgrad für die Stelle
Mindest-Berufserfahrung für die Stelle
Veröffentlicht in den letzten:
Gehalt