Product Manager - GPUaaS and OE Telemetry
Listed on 2026-09-04
-
IT/Tech
AI Business & Operations, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and Open Router.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
The Role:We are seeking an experienced Product Manager to lead the strategy, roadmap, and execution of our GPU-as-a-Service (GPUaaS) platform and Observability (OE) platform. This role is responsible for defining products that enable customers to seamlessly consume GPU infrastructure while empowering engineering and operations teams with comprehensive observability across infrastructure, platform services, and AI workloads through telemetry, monitoring, analytics, and automation.
Preferred
Location:
Remote, USA.
- Drive the product strategy and roadmap for the GPUaaS and OE Telemetry platform in collaboration with the Infrastructure Engineering team.
- Translate customer and engineering requirements into prioritized product roadmaps.
- Define product requirements, user stories, acceptance criteria, and success metrics.
- Lead product planning, roadmap reviews, and release planning.
- Measure product adoption, operational impact, and business outcomes using data-driven KPIs.
- Define capabilities requirements for GPU provisioning, lifecycle management, scheduling, orchestration, self-service portal, APIs, multi-tenancy, billing, metering, quotas, and access management.
- Define telemetry and observability capabilities requirements across GPU, compute, storage, networking, Kubernetes, Slurm, AI workloads, and supporting infrastructure.
- Drive capabilities adoption including (but not limited to):
- Metrics, logs, traces, and events collection
- Open Telemetry adoption and instrumentation
- Real-time dashboards and visualization
- Intelligent alerting and incident detection
- Service health and dependency mapping
- Distributed tracing
- Root cause analysis
- AI-driven anomaly detection and predictive insights
- SLO/SLI measurement and reliability reporting
- Partner with Infra Engineering teams, SRE, Platform Developers, Data Center Operations to improve observability, reliability, scalability and operation experience for GPUaaS platform.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- 5+ years of Product Management experience in cloud infrastructure, AI infrastructure, or enterprise platforms.
- Strong understanding of GPU computing (NVIDIA H100/H200/B200/B300 or equivalent), Kubernetes, Slurm, Linux, Cloud infrastructure, APIs, Telemetry and observability platforms
- Experience with monitoring technologies such as Prometheus, Grafana, Open Telemetry, Elasticsearch/Open Search, or similar.
- Experience translating customer requirements into technical product specifications.
- Strong analytical and communication skills.
- Experience building AI cloud or GPU cloud platforms.
- Knowledge of NVIDIA AI Enterprise and related software ecosystems NVSentinel, Fleet Intelligence, etc.
- Experience with multi-tenant IaaS platforms.
- Experience with Agile/Scrum product development.
- Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).