×
Register Here to Apply for Jobs or Post Jobs. X

MLOps & Platform Engineer

Job in 20093, Cologno Monzese, Lombardia, Italy
Listing for: Altro
Full Time position
Listed on 2026-07-27
Job specializations:
  • Software Development
    DevOps, Cloud Engineer - Software, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 35000 - 40000 EUR Yearly EUR 35000.00 40000.00 YEAR
Job Description & How to Apply Below
Location: Cologno Monzese

TXT Group , a company within the  TXT Group , is seeking an  MLOps & Platform Engineer  to join its Industrial Business Unit. The ideal candidate will be responsible to design, build and operate the company’s application and AI platform, ensuring secure, scalable and highly available environments for both enterprise applications and AI/ML workloads. At least three years’ relevant experience in platform and infrastructure engineering is required, with production ownership of:
Kubernetes, CI/CD pipelines and cloud infrastructure.
Main responsibilities   Design and manage cloud-native platforms, Kubernetes clusters and containerised applications;
Build and maintain CI/CD pipelines, and automate infrastructure provisioning and application deployment;
Design, deploy and manage infrastructure for AI systems: LLM and embedding model serving, vector databases, application databases and caching systems;
Define and maintain CI/CD pipelines for AI applications and data pipelines, with reproducible environments and secure release strategies;
Manage versioning of models, prompts and configurations, and support fine-tuning and retraining pipelines where required;
Collaborate with software developers and AI engineers to streamline delivery and operations;
Implement monitoring, logging, tracing and alerting for LLM applications and data pipelines, covering latency, error rate, cost per call, drift and production quality metrics;
Ensure scalability, high availability and operational continuity of AI services, including knowledge base ingestion and update pipelines;
Optimise inference and embedding costs through resource sizing, quantisation, batching and infrastructure-level caching;
Ensure overall platform security, observability, performance and operational reliability;
Manage secrets, API keys, access control and environment isolation for AI services;
Support offline evaluation and A/B testing activities by providing the necessary infrastructure and telemetry data.
Indispensable technical skills   Solid experience with containerisation and orchestration technologies such as Docker, Kubernetes, Helm and Docker Compose;
Proficiency with CI/CD and Dev Ops tooling, including Git, Git Lab and Git Lab CI/CD (or equivalents such as Git Hub Actions), Git Ops workflows, container registries and release management practices, with automation skills in Python and Bash;
Strong background in infrastructure automation on Linux, using Terraform and Ansible to implement Infrastructure as Code and manage virtualised environments;
Solid understanding of networking and security fundamentals (TCP/IP, HTTP/HTTPS, DNS), reverse proxies and ingress controllers (Nginx, Traefik), TLS/SSL, identity and access management (Keycloak, OAuth2/OpenID Connect), secrets management and IAM;

Experience with observability stacks such as Prometheus, Grafana, Loki, Open Telemetry, Open Search and Alert manager (or an equivalent ELK-based stack), applied to production ML and LLM systems;
Hands-on experience with AI/MLOps tooling, including MLflow and experiment tracking platforms such as Weights & Biases, model registries, and model serving frameworks such as vLLM, Ollama, Triton Inference Server, Sage Maker or Vertex AI, applied to GPU-based inference workloads;
Experience deploying and operating vector databases (Qdrant, pgvector), embedding models and RAG pipelines as part of production AI inference services;
Proficiency in backend and data technologies, including Python, FastAPI and REST APIs, together with operational experience running PostgreSQL, Microsoft SQL Server and Redis in production;
Practical experience operating production databases across SQL/No

SQL and vector stores, covering provisioning, backup and scaling;
Familiarity with LLMOps practices, including prompt and experiment tracking, inference cost monitoring and model version management;
Understanding of inference optimisation techniques such as quantisation, batching, caching and GPU-level optimisation (e.g. TensorRT);
Experience managing the ML/LLM model lifecycle, including fine-tuning and retraining pipelines, offline evaluation and A/B testing;
Working knowledge of at least one major cloud…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary