Infrastructure Production Engineer
Northern, Floyd County, Kentucky, USA
Listed on 2026-09-12
-
IT/Tech
Systems Engineer, IT Infrastructure, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description Who We Are
Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation.
Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company.
- 100% company-paid insurance premiums for employee medical, dental and vision plans.
- 401(k) plan that matches 100% up to 4%, with immediate vesting
- Professional Development Reimbursement of $2,500 each year
- 11 Holidays + Paid Time Off Accrual + Rollover Plan
- Commitment matters to Vultr! Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
- $500 stipend for remote office setup in first year + $400 each following year
- Internet reimbursement up to $75 per month
- Gym membership reimbursement up to $50 per month
- Company paid Wellable subscription
Vultr is seeking a highly skilled and experienced Associate Infrastructure Production Engineer to ensure GPU hardware validation and readiness across our global production environment. The ideal candidate is an analytical problem-solver with strong Linux skills, Python scripting ability, and a detail-oriented approach to infrastructure validation and automation. This is a highly visible role in a high-growth technology company, which will require collaboration with engineering teams to refine testing processes, close gaps in automation coverage, and build technical solutions that reduce deployment risk and improve production standards.
This is your opportunity to join our fast growing team and leave your mark on Vultr and the future of Cloud Infrastructure.
- Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware (NVIDIA and AMD) using vendor tooling, Python, and infrastructure automation technologies.
- Engineer and support Python-based agents, services, APIs, and Ansible automation (playbooks, roles, pipelines) that orchestrate hardware provisioning, telemetry collection, health monitoring, and production onboarding workflows.
- Analyze workload performance, utilization, thermals, and diagnostic output to identify hardware and system issues, enhance validation methodologies, and improve infrastructure readiness standards.
- Execute and evolve testing and verification processes for production onboarding and Return Material Authorization (RMA), developing automation enhancements to improve reliability, scalability, and coverage.
- Contribute to production stability by building tooling that gates hardware deployment, enforces quality standards, and reduces systemic infrastructure risk.
- Document system designs, automation logic, validation methodologies, and operational guidance; maintain accurate Jira records reflecting engineering activities, findings, and outcomes.
- Identify gaps or inefficiencies within testing, validation, or automation processes and design technical solutions to address them in collaboration with engineering teams.
$60,000 - $80,000
Final compensation will vary depending on years of experience, background/skill set, location, and applicable laws
Inclusion & PrivacyWe are an equal opportunity employer and are committed to creating an inclusive environment for all employees. We welcome applications from individuals of all backgrounds…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).