×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff — Inference

Job in New York, New York County, New York, 10261, USA
Listing for: Human Intuition Inc.
Full Time position
Listed on 2026-10-08
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer, Software Engineer, Machine Learning/ ML Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 140000 - 195000 USD Yearly USD 140000.00 195000.00 YEAR
Job Description & How to Apply Below
Location: New York

Building the autonomous company

Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.

We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.

The role

Make model execution dependable enough for business operations and efficient enough to improve continuously. You will build the serving and routing systems behind agents, evaluations, and training rollouts, and measure performance in terms that matter to complete tasks.

What you’ll do
  • Build inference services and model routing with explicit reliability, latency, and cost targets.

  • Profile representative agent workloads, including long contexts, tool calls, streaming, and concurrent requests.

  • Improve throughput and resource use through scheduling, batching, caching, and informed deployment choices.

  • Develop reproducible benchmarks that connect serving changes to task quality as well as speed and cost.

  • Implement versioned rollouts, observability, capacity planning, and practical failure recovery.

  • Partner with research and infrastructure engineers to support evaluation and post-training workloads.

What you’ll bring
  • Experience building and operating machine learning services or performance-sensitive distributed systems.

  • Strong Python and an understanding of model serving, accelerator memory, and concurrency.

  • Hands-on experience with an inference engine or substantial production model-serving workloads.

  • The ability to investigate performance bottlenecks and distinguish measured improvements from assumptions.

  • Ownership of reliability, debugging, and clear operational documentation.

Useful experience

GPU profiling, quantization, KV cache management, distributed serving, capacity scheduling, or contributions to inference software.

What success looks like

Agents and researchers have predictable access to models, and the team can explain and improve the cost and latency of completing a task.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary