Senior Cloud Software Engineer, DGXC Data Services
Listed on 2026-07-15
-
Software Development
Cloud Engineer - Software, Backend Developer, DevOps
Overview
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload.
Whatyou will be doing
- Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access.
- Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows.
- Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems.
- Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale.
- Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification.
- BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience.
- Strong foundation in algorithms, data structures, distributed systems, and practical software design.
- Experience building, shipping, and operating backend or cloud-native services using Kubernetes, cloud providers such as AWS, GCP, or Azure, and languages such as Go, Python, Rust, C/C++, or Java.
- Ability to design APIs, document systems, reason through tradeoffs, communicate clearly, and break ambiguous problems into practical execution plans.
- Experience working across engineering, product, platform, and operations teams to deliver reliable production software.
- Curiosity and practical judgment around AI-assisted or agentic engineering workflows, including using clear intent, specifications, acceptance criteria, tests, and verification to guide development.
- Hands‑on experience building, scaling, or operating large‑scale data, storage, or ML infrastructure services.
- Experience solving enterprise‑grade data management, governance, analytics, or AI workflow problems with modern data and ML infrastructure technologies.
- Strong background in distributed systems, storage systems, cloud infrastructure, performance engineering, observability, or agentic engineering practices.
Base salary ranges: $152,000 - $241,500 USD (Level
3) and $184,000 - $287,500 USD (Level
4). The base salary is determined based on location, experience, and peer pay. Eligible for equity and benefits.
Applications for this job will be accepted at least until July 14, 2026.
Equal Opportunity EmployerNVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).