Member of Technical Staff — Inference
Listed on 2026-10-08
-
Software Development
AI Reliability/ Performance Engineer, Software Engineer, Machine Learning/ ML Engineer, DevOps
Building the autonomous company
Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.
We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.
The roleMake model execution dependable enough for business operations and efficient enough to improve continuously. You will build the serving and routing systems behind agents, evaluations, and training rollouts, and measure performance in terms that matter to complete tasks.
What you’ll doBuild inference services and model routing with explicit reliability, latency, and cost targets.
Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.
Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.
Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.
Implement versioned rollouts, observability, capacity planning, and practical failure recovery.
Partner with research and infrastructure engineers to support evaluation and post-training workloads.
Experience building and operating machine learning services or performance-sensitive distributed systems.
Strong Python and an understanding of model serving, accelerator memory, and concurrency.
Hands-on experience with an inference engine or substantial production model-serving workloads.
The ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.
Ownership of reliability, debugging, and clear operational documentation.
GPU profiling, quantization, KV cache management, distributed serving, capacity scheduling, or contributions to inference software.
What success looks likeAgents and researchers have predictable access to models, and the team can explain and improve the cost and latency of completing a task.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).