Senior CloudOps Engineer (On-site
Job Description & How to Apply Below
About Translated Translated is a leading provider of AI-powered language solutions. Founded in 1999 by a linguist and a computer scientist, we are on a mission to allow everyone to understand and be understood in their own language.
About Translated Translated is a leading provider of AI-powered language solutions. Founded in 1999 by a linguist and a computer scientist, we are on a mission to allow everyone to understand and be understood in their own language.
We envision a world where people from different cultures can communicate seamlessly, gaining unprecedented access to knowledge, cultural exchange, and opportunity. To make this possible, we combine proprietary translation AI (Lara Translate) with advanced text, audio, and video translation technologies (Matecat, Matesub, and Matedub) and the world’s largest network of vetted, native-speaking language professionals.
We welcome complex technical challenges from our customers and engineer tailored solutions that often become part of our core products, whether integrated into our enterprise localization platform, the tools designed to support translators, or Lara Translate, our online AI translator for teams and individuals.
Everything we build reflects a simple principle: technology should amplify human potential, not replace it.
We believe in humans.
We are looking for a Senior Cloud Ops Engineer to join our Infrastructure Team.
We run a genuinely hybrid platform, combining AWS cloud environments with our own on-premise infrastructure and custom GPU clusters built for AI training and inference. In this role, your core focus will be driving Cloud Governance, scaling our Kubernetes footprint (both cloud and on-prem), and enabling our AI development through modern MLOps tooling.
You will own the infrastructure lifecycle end-to-end, acting as an enabler for our engineering teams by embedding modern Infrastructure-as-Code (IaC) and cost-optimization practices across the entire stack.
Responsibilities Cloud Governance:
Lead cloud architecture governance on AWS. Own cost optimization, capacity planning, access management, and ensure 100% of infrastructure is declared via clean, modular Infrastructure-as-Code (Terraform / Ansible).
Hybrid Kubernetes Architecture:
Design, build, and operate enterprise-grade Kubernetes clusters across both cloud and on-premise environments. Standardize deployments using Git Ops (Argo CD), Helm, and modern networking (Cilium).
MLOps
Infrastructure: Build and maintain high-performance infrastructure pipelines for AI model deployment and inference. Optimize GPU usage and monitoring for AI/ML teams.
Team Enabling & Developer
Experience:
Champion the "you build it, you run it" culture. Provide guidance, templates, and self-service tooling to help dev teams deploy, monitor, and scale their services safely and autonomously.
Infrastructure Operations (Nice-to-Have Support):
Continuously optimize platform reliability, observability (Datadog/ELK/Cloud Watch), CI/CD pipelines (Git Hub Actions), and hybrid connectivity alongside our core infrastructure team.
Requirements 5+ years of hands-on experience in a Cloud Ops / Dev Ops / SRE role managing production infrastructure.
Cloud Governance & IaC:
Strong background in AWS infrastructure management, cost optimization, and deep expertise with Terraform and Ansible.
Kubernetes Expertise:
Proven production experience building and operating Kubernetes clusters in both cloud (AWS EKS) and bare-metal/on-prem environments.
Experience with Argo CD (Git Ops) and Helm.
MLOps Foundations:
Experience serving, scheduling, and scaling machine learning inference/training workloads (NVIDIA GPU drivers, device plugins, or MLOps frameworks).
Engineering Mindset:
Strong scripting skills (Python, Bash, or Go) to build automation, CLI tools, and platform integrations.
Team Enablement:
Excellent communication skills with a passion for mentoring developers, driving best practices, and improving developer experience.
Nice to Have Production experience with bare-metal virtualization (Proxmox VE, Ceph).
Advanced Kubernetes networking using Cilium or eBPF.
Experience with Docker Swarm to Kubernetes migrations.
Relatio…
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×