DevOps Engineer; AI Inference
Listed on 2026-07-06
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Location: Town of Poland
This position is available only under an employment (labor) agreement.
The world’s digital experiences run on something invisible: the infrastructure and software that keep them fast, reliable, and secure. AtGcore,
you’llhelp design and deliver that foundation for an AI-driven world.
We’rea global provider of infrastructure and software solutions for
AI, cloud, network, and security,powering everything from real-time communication and streaming to enterprise AI and secure web applications. With
210+ edge locations, 50+ cloud regions, and thousands of GPUs
, your work here can reach users and businesses across the globe.
We’ll collaborate with leading technology partners such as
Intel, NVIDIA, Dell, and Equinix
, and work on platforms that power digital products used around the world. Our vision is simple: to connect the world to AI, anywhere, anytime.
Want to work on technology that goes beyonda single productor industry? Join a global team of
550+ professionals building infrastructure and software that supports the entire digital ecosystem.
We are looking for a talented Dev Ops Engineer to join our AI Inference Operations Team.
Job DescriptionAs a Dev Ops Engineer, you will be responsible for designing, deploying, and maintaining infrastructure and services that enable scalable and secure AI inference workloads on-premises.
What You Will Do- Design, develop, and maintain infrastructure for AI inference workloads, including GPU scheduling, model deployment pipelines, and data access patterns in on-prem environments
- Build and manage monitoring and observability tools for AI inference platforms, including dashboards, alerts, and runbooks for model health and system performance
- Collaborate with ML engineers and platform teams to design system architecture for AI workloads, integrate inference runtimes, and test performance at scale
- Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components.
- Hands-on experience operating and troubleshooting production Kubernetes clusters.
- Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity and performance issues.
- Ability to develop automation and operational tooling using Python, Go, or Bash.
- Experience with Terraform, Ansible, or similar IaC/configuration management tools.
- Experience with Victoria Metrics/Grafana or similar monitoring, alerting, and troubleshooting tools.
- Strong experience with Git-based workflows and CI/CD pipelines.
- Familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies.
- Hands-on operation or administration of Slurm clusters.
- Knowledge of Argo CD, Git Ops workflows, Helm, or Helmfile.
- Background working with managed platforms, PaaS, or cloud services.
- Exposure to bare metal, GPU, HPC, or other high-performance computing environments.
- Familiarity with the NVIDIA GPU stack, RDMA/Infini Band, or high-performance networking.
- Knowledge of Open Stack or similar cloud infrastructure platforms.
- Hands-on experience developing Kubernetes operators or controllers.
- Competitive compensation
- Flexible working hours and hybrid or remote options, depending on your role
- Work from anywhere in the world for up to
45 days per year - Private medical insurance for you and your family*
- Extra paid vacation and sick leave days*
- Support for life’s important moments and celebrations
- Language courses to help you connect and grow
- Modern, welcoming offices with snacks, drinks, and entertainment*
- Team sports and social activities*
* Benefits may vary depending on your location.
Equal Opportunity EmployerWe provide equal opportunity to all applicants without regard to race, color, religion, sex, sexual orientation, age, gender identity, gender expression, national origin, disability, or any other legally protected characteristics.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).