AI Systems Performance Engineer
Listed on 2026-07-18
-
Software Development
Unix/Linux, DevOps
We are seeking a highly talented and experienced Senior AI Fabric Performance Engineer to take on a critical role within our Performance Lab.
Responsibilities- Benchmarking & Execution:
Install, configure, and run industry‑standard AI performance benchmarks with an emphasis on MLPerf (Training and Inference) and NCCL tests. - Fabric Optimization:
Tune and optimize Ethernet network parameters to ensure seamless data flow for distributed AI workloads running on server clusters. - Deep Debugging:
Identify, isolate, and troubleshoot complex system performance bottlenecks spanning the Linux OS, server hardware, and Ethernet switches. - Automation Development:
Design, develop, and implement robust performance‑testing frameworks and automation tools to streamline continuous benchmarking. - Cross‑Functional
Collaboration:
Document test methodologies, communicate performance findings, and provide actionable improvement recommendations to hardware, software, and networking stakeholders.
- Education:
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, plus 10–12 years of related industry experience. - OS Expertise:
Deep familiarity with Linux operating systems, including system‑level performance tuning and troubleshooting. - Programming
Skills:
Strong proficiency in Python and C++ for scripting and development. - AI/ML Knowledge:
Familiarity with modern machine learning frameworks, particularly PyTorch, and an understanding of how AI models consume compute and network resources. - Networking & Fabric:
Proven experience in performance testing and validating Ethernet switch systems. - Analytical Capabilities:
Extensive experience with performance metrics, profiling, and benchmarking tools. - Problem Solving:
Strong ability to diagnose root causes in complex, distributed systems. - Preferred Qualifications (optional but recommended for a critical role):
Experience with RDMA (Remote Direct Memory Access) and RoCEv2; prior experience building CI/CD pipelines for automated hardware or software performance regression testing; familiarity with containerization and orchestration tools such as Docker and Kubernetes.
The annual base salary range is $141,300 – $226,000.
Eligible for a discretionary annual bonus, a competitive new‑hire equity grant, and annual equity awards.
Broadcom offers a comprehensive benefits package, including medical, dental, and vision plans; 401(k) participation with company matching;
Employee Stock Purchase Program (ESPP);
Employee Assistance Program (EAP); company‑paid holidays; paid sick leave and vacation time. The company follows all applicable laws for Paid Family Leave and other leaves of absence.
Broadcom is proud to be an equal‑opportunity employer. We will consider qualified applicants without regard to race, color, creed, religion, sex, sexual orientation, national origin, citizenship, disability status, medical condition, pregnancy, protected veteran status, or any other characteristic protected by federal, state, or local law. We will also consider qualified applicants with arrest and conviction records consistent with local law.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).