×
Register Here to Apply for Jobs or Post Jobs. X

GPU Inference SDET

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Engg
Full Time position
Listed on 2026-10-09
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.

OpenAI recently announced a multi-year partnership  with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About the Role As a Staff GPU Inference SDET, you will be the founding quality, reliability, and validation lead for a new GPU Inference Development team. Working closely with engineering leads and cross-functional systems infrastructure teams, you will design, build, and scale the end-to-end release qualification and automated test ecosystem for our GPU inference stack and rack-scale accelerated compute fleets. In this high-impact role, you will be responsible for building automated test suites to validate multi-node GPU cluster bring-up, verifying prefill worker optimizations, testing open-source and custom serving engines, and ensuring numerical correctness and performance stability under real-world streaming workloads.

You will be the primary technical anchor ensuring production-grade reliability, fault isolation, and peak inference performance across accelerated GPU infrastructure.

WHAT YOU’LL DO
  • Build GPU Release Qualification Systems:
    Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack—spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware.
  • Inference Serving & Workload Validation:
    Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism.
  • Performance & Performance Modeling Verification:
    Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including Time-to-First-Token (TTFT), Inter-Token Latency (ITL), request throughput, tail latency (P99), and capacity efficiency.
  • Numerical Correctness & Quality Gates:
    Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions.
  • Fault Injection & Fleet Resilience:
    Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters.
  • Observability & CI/CD Integration:
    Integrate automated test pipelines with telemetry tools (e.g., Prometheus, Grafana) to turn one-off investigations into repeatable engineering gates and continuous performance monitoring.
REQUIREMENTS
  • 8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
  • GPU & Cluster Infrastructure Expertise:
    Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD ecosystem) across public cloud infrastructure or enterprise data center environments.
  • Inference Stack Knowledge:
    Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching.
  • Automation & Scripting:
    Expert-level…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary