×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - AI Cloud Infrastructure

Job in Oakland, Alameda County, California, 94616, USA
Listing for: Emerald AI
Full Time position
Listed on 2026-08-11
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 190000 - 240000 USD Yearly USD 190000.00 240000.00 YEAR
Job Description & How to Apply Below

About Emerald AI

We’re at a pivotal moment for AI and energy. Demand for compute is skyrocketing, but power constraints are becoming a critical bottleneck. Emerald AI sits at the intersection of these two worlds, enabling AI data centers to scale without overwhelming the grid.

Our Emerald Conductor software platform makes data centers flexible and responsive, allowing them to adjust power usage dynamically. This unlocks massive AI growth without major new infrastructure, while also strengthening the grid and supporting the expansion of renewable energy.

We’re a team of experts across AI, cloud, software, and energy—on a mission to scale AI sustainably. We’re backed by leading investors and partners including Radical Ventures and NVIDIA.

Learn more about our vision, team, and backers at .

About

The Role

Emerald AI is building the world's first power flexible managed cloud infrastructure. We are hiring a senior infrastructure engineer to architect and stand up our managed cloud services from end to end. The work covers the platform, the control plane, and the customer experience that together make up a managed AI cloud.

The right person has done this before. They have built or served as a core early engineer on a managed cloud or AI platform, whether at a GPU cloud, an internal machine learning platform run at scale, a hyperscaler AI service, or a HPC research computing centre operated as a service. This is a role for an architect who still builds.

You will make the major design decisions and then implement them yourself.

Key Responsibilities
  • Architect our managed services from 0→1. Define the productization of GPU capacity, encompassing isolation boundaries, tenant models, provisioning flows, and service catalogs that scale across diverse providers.
  • Engineer the platform core. Build robust control‑plane services, self‑service customer interfaces, and automated lifecycle systems, including usage metering integrated with billing infrastructure.
  • Onboard and vet infrastructure partners. Conduct deep technical assessments of bare‑metal GPU vendors, evaluating fabric quality, network isolation, and economics to automate the path from handoff to active tenant.
  • Design end‑to‑end multi‑tenancy. Implement rigorous isolation across compute, storage, and networking (Infini Band/VLANs), ensuring secure boundaries, QoS, and encryption even when customers possess root access.
  • Drive workload orchestration. Manage Kubernetes and Slurm environments for large‑scale training and inference, overseeing node health, driver fleets, and kernel management across heterogeneous clouds.
  • Lead high‑performance storage strategy. Deploy and integrate parallel storage solutions like Lustre, VAST, or Weka, leveraging your deep experience with these systems to ensure they fold cleanly into our provisioning model.
  • Ensure operational excellence. Define SLOs, observability standards, and incident response protocols that bridge our internal standards with underlying provider SLAs to deliver a reliable, sellable product.
Minimum Requirements
  • At least 7+ years of experience in infrastructure or platform engineering, including the architecture and launch of a managed cloud or AI platform that reached production users.
  • Strong experience with Kubernetes and Slurm and offering them as managed service
  • Production experience deploying or operating Lustre or a comparable parallel file system such as GPFS, Weka, VAST, or BeeGFS, with a solid understanding of parallel file system architecture, tuning, and failure modes.
  • A strong grasp of cloud service fundamentals, including control planes, tenancy and isolation models, APIs, quota and metering systems, and the operational discipline of running a service that customers pay for.
  • Deep Linux systems knowledge, mature infrastructure as code practice with tools such as Terraform and Ansible, and solid programming ability in Python or Go.
  • Familiarity with GPU infrastructure, including high performance networking with Infini Band, RoCE, and RDMA, and the GPU software stack.
Preferred Requirements
  • Prior time at a GPU cloud, a hyperscaler AI service, or an HPC center that delivers compute and storage as a service,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary