Senior DevOps Engineer
Listed on 2026-09-29
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, SRE/Site Reliability, Systems Engineer
Senior Dev Ops / Cloud Platform Engineer – ML AI Infrastructure Job Summary We are looking for a Senior Dev Ops / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services. The ideal candidate will have hands‑on experience with AWS EKS, Sage Maker, Bedrock, Docker, Kubernetes, Terraform, Helm, Git Hub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments .
This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost‑efficient platforms for ML services across development and production environments. In this role you will...
- AWS Kubernetes Infrastructure Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
- Manage Amazon EKS clusters , including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
- Work with AWS Sage Maker, AWS Bedrock, EKS, ECS, and related AWS services .
- Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB .
- Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security .
- Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
- Implement best practices for security, reliability, availability, and scalability.
- CI/CD Azure-to-AWS Migration Build and maintain CI/CD pipelines for ML and AI services.
- Develop and manage Git Hub Actions and Azure Dev Ops pipelines using YAML.
- Migrate repositories and CI/CD workflows from Azure Dev Ops to Git Hub/AWS .
- Automate build, test, containerization, deployment, and release processes.
- Establish deployment strategies across development, staging, and production environments.
- ML Service Deployment Deploy and manage ML services across AWS EKS/ECS and Sage Maker .
- Build and maintain Docker containers and Kubernetes deployments.
- Manage environment segregation and configuration across Dev, QA, and Production.
- Develop and maintain Kubernetes manifests and Helm charts .
- Troubleshoot ML service deployment, networking, scaling, and runtime issues.
- App Runner to EKS Migration Lead migration of existing services from AWS App Runner to Amazon EKS .
- Containerize applications and develop Kubernetes manifests/Helm charts.
- Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
- Ensure minimal service disruption during migration and establish operational best practices on EKS.
- Self-Hosted LLM AI Infrastructure Deploy and manage self-hosted Large Language Models and inference services.
- Work with model serving frameworks such as vLLM .
- Design containerized infrastructure for GPU-based model serving and inference.
- Manage model versions, deployments, configurations, and rollback strategies.
- Support migration of ML services from managed APIs/services to self-hosted models .
- Work with engineering teams on API integration and inference infrastructure.
- Databricks Administration Administer Databricks work spaces, clusters, permissions, and access controls .
- Manage cluster configuration, policies, and resource utilization.
- Support LMI Insights and related ML/AI workloads.
- Troubleshoot Databricks infrastructure and connectivity issues.
- Implement appropriate security and access-control practices.
- Elasticsearch Infrastructure Design, deploy, and manage Elasticsearch clusters .
- Perform cluster sizing, scaling, configuration, and performance optimization.
- Manage indices, mappings, retention, and data lifecycle requirements.
- Support Kibana configuration, dashboards, and troubleshooting.
- Monitor Elasticsearch health,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).