Data Center Operations Engineer - Chicago ORD
Listed on 2025-11-28
-
IT/Tech
Systems Engineer, Cloud Computing
Lambda, The Superintelligence Cloud, builds Gigawatt-scale AI Factories for Training and Inference. Lambda’s mission is to make compute as ubiquitous as electricity and give every person access to artificial intelligence. One person, one GPU.
If you’d like to build the world’s best deep learning cloud, join us.
* Note:
This position requires presence in our Chicago/Elk Grove Village Data Center location 5 days per week.
The Operations team is at the heart of keeping our AI-IaaS infrastructure running smoothly from start to finish. They handle everything from sourcing the right hardware and components to keeping our data centers performing at their best day in and day out. The team also works closely across the company, making sure our operational capabilities stay in sync with product goals and overall strategy.
By managing the entire lifecycle — from procurement through deployment and ongoing efficiency — the Operations team ensures our AI infrastructure stays reliable, scalable, and ready to support the business as it grows.
Make sure new servers, storage, and networking gear are racked, labeled, cabled, and configured the right way.
Keep data center layouts and network topologies up to date in our DCIM software.
Coordinate with supply chain and manufacturing teams so systems are deployed on time, especially for large-scale projects.
Evaluate current and future data center needs based on growth and technology trends.
Manage parts depot inventory and track equipment as it moves from delivery → storage → staging → deployment → handoff.
Work closely with hardware support teams to get tickets resolved quickly.
Create and manage RMA tickets when needed, making sure faulty parts are replaced and reinstalled without delay.
Develop and maintain installation standards (placement, labeling, cabling) to ensure consistency across all data centers.
Act as a subject matter expert on data center deployments, supporting sales engagements for major deployments in our facilities or at customer sites.
Researching, evaluating, and securing the right hardware and infrastructure components.
Building relationships with peers and supply chain to ensure cost-effective and timely supply.
Monitoring day-to-day performance of data centers to maintain uptime and efficiency.
Troubleshooting and resolving hardware or infrastructure issues quickly.
Performing regular maintenance and upgrades to keep systems running at peak performance.
Overseeing the full lifecycle of infrastructure, from initial setup to ongoing optimization.
Coordinating deployments of new hardware and ensuring seamless integration with existing systems.
Managing capacity planning to make sure infrastructure can scale with business growth.
Working with product management, support, and other teams to align operational capabilities with company goals.
Translating business priorities into technical and operational requirements.
Supporting cross-functional projects where infrastructure plays a critical role.
Ensuring infrastructure remains stable, secure, and scalable as demand increases.
Continuously improving processes to boost efficiency and reduce downtime risks.
Certifications:
Any Linux or project management.Military background.
Experience in the machine learning or computer hardware industry.
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About LambdaFounded in 2012, ~400 employees (2025) and growing fast
We offer generous cash & equity compensation
Our investors include Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, US Innovative Technology, Gradient Ventures, Mercato Partners, SVB, 1517, Crescent Cove.
We are experiencing extremely high demand for our systems, with quarter over…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).