More jobs:
AI Senior Systems Engineer
Job in
San Jose, Santa Clara County, California, 95199, USA
Listed on 2026-07-23
Listing for:
Cadence Design Systems
Full Time
position Listed on 2026-07-23
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software), Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology. We are seeking a highly skilled and experienced AI Systems Engineer to join our team. This is a hands‑on, senior individual contributor role that will be pivotal in leading the development, operations, and support of our entire AI infrastructure.
Responsibilities- AI Infrastructure Architecture & Strategy
- Lead the design and implementation of next‑generation AI infrastructure to support Agentic AI initiatives.
- Define technical strategy for on‑premise GPU clusters, storage solutions, and networking to ensure optimal performance, scalability, and reliability for all AI workloads.
- Cloud AI Service Integration
- Support and secure the use of public cloud AI services, including Azure OpenAI and Google Cloud Platform (GCP) services like Gemini.
- Manage secure access, monitor usage, track billing for cost‑effectiveness, and provide hands‑on experience with compute, GPUs, and AI services on both GCP and Azure.
- GPU Cluster Management
- Configure, install, and optimize GPU server clusters, troubleshoot hardware and software, tune performance, and implement best practices for cluster utilization and resource management.
- Administer job schedulers such as LSF in a production environment and integrate with Docker for containerized job submission.
- Full‑Stack AI Tech Stack Development & Operations
- Architect and deploy scalable AI tech stacks, including deep learning frameworks (PyTorch, Tensor Flow), Docker, Kubernetes, and CI/CD pipelines for AI model development.
- Advanced LLM Deployment & Optimization
- Lead the deployment, serving, and optimization of Large Language Models using model quantization, distillation, and high‑performance serving frameworks (vLLM, TGI, TensorRT‑LLM).
- Agentic AI Workflow & Service Engineering
- Architect and build production‑grade Agentic AI workflows and services, integrating LLMs with external tools, APIs, and databases.
- Mentor other engineers on building robust and scalable AI agent applications.
- Automation & Monitoring
- Develop and maintain automation scripts using Python, Bash, or Perl to streamline system maintenance, deployment, and reporting.
- Implement and manage monitoring solutions for system health, GPU utilization, container performance, and job statuses.
- AI Systems Support & Mentorship
- Act as the final escalation point for complex technical issues related to AI infrastructure.
- Serve as a technical leader and mentor to other engineers, providing guidance on best practices, performance tuning, and operational excellence.
- Security and Compliance
- Develop and implement security best practices for AI systems and data, ensuring compliance with relevant regulations and protecting intellectual property.
Skills and Qualifications
- 10+ years of experience in a senior technical role, with at least 5 years focused on building and operating high‑performance computing or AI infrastructure.
- Proven track record as a Principal or Senior Staff Engineer.
- Expert‑level knowledge of NVIDIA GPU architecture, CUDA, and cuDNN.
- Extensive experience with multi‑GPU and multi‑node training and inference.
- Proven experience managing access, usage, and billing for Azure OpenAI and GCP services.
- Extensive hands‑on experience with Docker (image management, container orchestration, troubleshooting).
- Proficiency in scripting languages such as Python, Bash, or Perl.
- Deep expertise in Linux system administration (RHEL preferred), networking, storage, and performance tuning.
- Familiarity with user authentication and integration via LDAP or Active Directory.
- Strong problem‑solving and communication skills with the ability to work in a multi‑platform, cross‑functional, geographically distributed team.
- Understanding of AI job profiling and tuning (memory, GPU, I/O).
- Experience administering LSF clusters in production or research environments.
- Familiarity with other job schedulers such as Slurm.
- Experience with LSF Docker integration and container image job submission.
- Experience with macOS/Apple Silicon system admin tasks and troubleshooting.
The annual salary range for California is $136,500 to…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×