AI Compute Systems Intern: Optimize, Benchmark & Deploy
Job in
Mountain View, Santa Clara County, California, 94039, USA
Listed on 2026-08-03
Listing for:
Socket.dev
Full Time, Apprenticeship/Internship
position Listed on 2026-08-03
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below
About Naïve
Naïve is an autonomous company builder — we let businesses deploy AI employees to create/run entire companies. We're a $120M company backed by Y Combinator, Liquid2, DEEPCORE (Softbank), and more.
We're hiring interns who ship and hope to convert to full-timers.
What You'll Do- Research and ship the systems that make running thousands of AI agents dramatically cheaper, faster, and more reliable
- Optimize local / self-hosted model inference — quantization, batching, speculative decoding, KV-cache strategy, tensor & pipeline parallelism
- Build model routing that sends every request to the cheapest model that can actually do the job — frontier API when it matters, local when it doesn't
- Benchmark and deploy across hardware — GPUs, edge, on-prem, alternative accelerators — and turn the numbers into real deployment decisions
- Push on agent infrastructure: orchestration, caching, context management, and parallelization for fleets of concurrent agents
- Prototype recursive self-improvement loops — agents that improve their own tooling, prompts, and evals
- Own a research question end-to-end — frame it, run the experiments, ship the result into production
You're not here to write papers nobody reads. You're here to find the cost/performance frontier and ship past it.
Every dollar and millisecond you save compounds across an entire fleet of AI employees. Our best interns take a benchmark on Monday and land a production cost win by Friday — and write the changelog entry themselves. This is research with a deploy button.
Must-Haves- High agency
- Strong systems + ML engineering — comfortable in Python and PyTorch, can profile, optimize, and ship without hand-holding
- Real understanding of how transformers actually run — attention, KV cache, memory bandwidth, throughput vs. latency tradeoffs
- Built and shipped something real — side project, OSS, hackathon win, research artifact, prior internship
- Comfortable with LLMs both as tooling and as objects of study — API calls, prompts, tool use, and what's happening under the hood
- Move fast, take feedback, push back when you're right
- Hands-on with inference engines (vLLM, TensorRT-LLM, SGLang, llama.cpp)
- GPU kernel or low-level perf work (CUDA, Triton)
- Hardware benchmarking or deployment experience (cloud GPUs, on-prem, edge, alt accelerators)
- Built with agent frameworks
- Published, open-sourced, or blogged research/tooling
- Currently enrolled in a CS / EE / math program — or dropped out of one to build
* P.S. If you're serious about this role, send Sean a connection request with a note over Linked In.*
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×