×
Register Here to Apply for Jobs or Post Jobs. X

Senior Network Engineer - InfiniBand​/UFM

Job in Seattle, King County, Washington, 98127, USA
Listing for: Lightning AI
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Infrastructure, AI Engineer (Applied/Software), Systems Engineer
Salary/Wage Range or Industry Benchmark: 170000 - 210000 USD Yearly USD 170000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Senior Network Engineer - InfiniBand / UFM

Senior Network Engineer - Infini Band / UFM

New York, New York, United States;
San Francisco, California, United States;
Seattle, Washington, United States

Who We Are

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.

Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer‑first software with cost‑efficient, large‑scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.

We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and First minute.

Who We Look For

The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here’s what that looks like in practice:

  • Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.
  • Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.
  • Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.
  • Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.
  • Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.
  • Think Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.
The Role

We are seeking an experienced Senior Network Engineer (Infini Band / UFM) to design, deploy, automate, and operate next‑generation AI Factory networking infrastructure supporting large‑scale GPU clusters. This role is responsible for building and maintaining high‑performance NVIDIA Quantum Infini Band fabrics that power AI training and inference environments utilizing NVIDIA UFM, NCCL, RoCEv2, and modern data center technologies.

The ideal candidate has deep expertise in Infini Band networking, NVIDIA Unified Fabric Manager (UFM), large‑scale GPU deployments, Linux networking, automation, and troubleshooting distributed AI workloads.

What You’ll Do
  • Design, deploy, and maintain large‑scale NVIDIA Infini Band fabrics supporting AI/ML GPU clusters.
  • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
  • Configure and optimize NVIDIA Quantum and Quantum‑2 Infini Band switches.
  • Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
  • Implement and validate fat‑tree, Dragonfly+, Clos, and spine‑leaf network architectures.
  • Perform firmware lifecycle management for Infini Band switches, adapters (HCAs), and UFM infrastructure.
  • Optimize congestion control, adaptive routing, QoS, and traffic engineering for high‑performance GPU communication.
  • Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments.
  • Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure‑as‑Code methodologies.
  • Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms.
  • Support high availability, maintenance windows, incident response, root‑cause analysis, and capacity planning.
  • Participate in architecture reviews and define networking standards for AI infrastructure.
Required Qualifications
  • 7+ years of data center networking experience.
  • 3+ years supporting NVIDIA Infini Band environments.
  • Hands‑on experience with NVIDIA UFM Enterprise.
  • Experience deploying and operating Quantum and Quantum‑2 Infini Band switches.
  • Strong understanding of:
    • Infini Band Architecture
    • Subnet…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary