Cloud Support Engineer
Listed on 2026-10-05
-
IT/Tech
Unix/Linux, Systems Administrator, Cloud Computing: Infrastructure & Operations
At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.
Job SummaryWe are seeking a highly motivated Cloud Support Engineer to support our secure, cloud‑based silicon chip design environments used by our external customers for mission‑critical EDA and HPC workloads. This role is customer‑facing and service‑oriented, requiring not only strong technical expertise, but also a commitment to responsiveness, professionalism, and delivering an exceptional customer experience. The Cloud Support Engineer will be responsible for maintaining system reliability, resolving issues quickly and accurately, and proactively improving the platform through automation and agentic AI.
Success in this role is measured not just by technical outcomes, but by customer satisfaction, trust, and confidence in the service. This position involves working with export‑restricted data (ITAR/CUI) and supporting secure environments with high operational and compliance standards.
- Customer Support & Service Excellence
Serve as a primary technical support contact for external customers using secure cloud‑based silicon design platforms
Deliver timely, responsive, and high‑quality support, ensuring customer issues are acknowledged, communicated, and resolved effectively
Actively focus on delighting customers by minimizing downtime, anticipating needs, and exceeding service expectations
Clearly communicate technical issues, status updates, and resolutions to customers with varying levels of technical expertise
Build trust with customers through professionalism, accountability, and consistent follow‑through - Platform & Environment Support
Support and troubleshoot Linux‑based cloud environments used for silicon design and verification workflows
Maintain availability, performance, and reliability of secure multi‑tenant environments
Diagnose and resolve system‑level issues across compute, storage, networking, and identity services - HPC, Licensing & Performance Management
Monitor HPC cluster performance, job throughput, and queue health
Identify and remediate HPC job performance issues, including scheduler configuration, resource contention, I/O bottlenecks, and memory constraints
Troubleshoot and resolve license availability, utilization, and checkout issues impacting customer workloads
Support distributed resource managers (e.g., Slurm, LSF, SGE, or similar technologies) - Automation & Agentic AI
Develop and maintain automation to streamline recurring operational tasks, including:
System health, performance, and capacity monitoring User provisioning and de‑provisioning License usage tracking and alerting Detection of abnormal system or job behavior
Use agentic AI, Python, shell scripting, Perl, or similar technologies to reduce manual effort and improve mean time to resolution (MTTR)
Continuously improve operational efficiency while enhancing the customer experience - Security, Compliance & Operations
Operate and support systems containing ITAR‑controlled and CUI data in compliance with regulatory and company requirements
Follow documented security, access control, auditing, and change management processes
Participate in incident response, root cause analysis, and corrective actions
Create and maintain runbooks, knowledge base articles, and customer‑facing documentation
- Technical Skills
Strong hands‑on experience with Linux system administration and troubleshooting
Experience supporting HPC or large‑scale compute environments
Proficiency in Python, shell scripting, Perl, or other automation‑focused programming languages
Experience with monitoring, logging, and alerting systems
Familiarity with license management systems (e.g., Flex Net/FLEXlm or equivalent) - Custome…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).