×
Hier anmelden um sich kostenlos auf Stellen zu bewerben oder Stellenanzeigen aufzugeben. X

AI Inference Platform Engineer

in 10115, Berlin, Berlin, Deutschland
Unternehmen: XpertDirect
Vollzeit position
Verfasst am 2026-09-14
Berufliche Spezialisierung:
  • IT/Informationstechnik
    Künstliche Intelligenz Ingenieur, Maschinelles Lernen, Site Reliability Ingenieur/in, Cloud Computing: IT-Infrastruktur & Betrieb
Gehalts-/Lohnspanne oder Branchenbenchmark: 70000 - 110000 EUR pro Jahr EUR 70000.00 110000.00 YEAR
Stellenbeschreibung

AI Infrastructure | Model Serving | GPU Computing | Inference Engineering | ML Platforms

Our client, a growing AI Infrastructure company based in Berlin, is looking for an AI Inference Platform Engineer to build and optimise the platform used to serve production AI models across GPU-enabled infrastructure.

You'll work at the intersection of AI Infrastructure, Distributed Systems, and Platform Engineering
, focusing on inference performance, GPU utilisation, autoscaling, latency, and reliability.

What You’ll Work On
  • Build and operate Kubernetes infrastructure for production AI inference
  • Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton
  • Improve GPU utilisation, throughput, and inference latency
  • Design autoscaling strategies for dynamic AI workloads
  • Build platform tooling and automation in Python
  • Provision and manage infrastructure using Terraform
  • Develop observability across models, GPUs, Kubernetes, and serving infrastructure
  • Profile and troubleshoot inference performance bottlenecks
  • Improve batching, concurrency, caching, and resource allocation strategies
  • Build reliable deployment workflows for new models and model versions
  • Partner with ML Engineers to move models efficiently into production
Core Skills
  • 4+ years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering, or similar roles
  • Strong understanding of Linux and distributed production systems
Nice to Have
  • CUDA
  • NVIDIA GPU Operator
  • Py Torch
  • KServe
  • Ray Serve
  • LLM inference optimisation
  • Multi-GPU inference
  • AWS / GCP GPU infrastructure
  • Experience operating high-throughput or latency-sensitive inference services
Um Jobs auf dieser Seite anzusehen und sich zu bewerben, die Bewerbungen aus Ihrem Standort oder Land akzeptieren, klicken Sie unten auf den Button, um eine Suche zu starten.
(Wenn dieser Job tatsächlich in Ihrem Zuständigkeitsbereich liegt, verwenden Sie möglicherweise einen Proxy oder VPN, um auf diese Seite zuzugreifen. Um weiterzukommen, sollten Sie Ihre Verbindung zu einem anderen Mobilgerät oder PC wechseln).
 
 
 
Suchen Sie hier nach weiteren Stellen:
(nach Beruf, Fähigkeit)
Standort
Suchradius erweitern (Meilen)
0
200
Filter
Mindest-Bildungsgrad für die Stelle
Mindest-Berufserfahrung für die Stelle
Veröffentlicht in den letzten:
Gehalt