×
Register Here to Apply for Jobs or Post Jobs. X

ML Research Engineer

Job in New York, New York County, New York, 10261, USA
Listing for: Nodi
Full Time position
Listed on 2026-09-25
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Data Scientist
Salary/Wage Range or Industry Benchmark: 235000 - 295000 USD Yearly USD 235000.00 295000.00 YEAR
Job Description & How to Apply Below
Location: New York

Compensation $235,000 – $295,000
· per year
· USD

Interested in this role?

The Role

Our scientific agent decides what experiment to run next. This role builds the machinery that lets us understand, evaluate, and continuously improve the agent's decisions.

You will work closely with the Principal Scientist to turn hypotheses about agent behavior into reproducible experiments and production improvements. Your initial focus will be evaluation datasets, replay, trajectory analysis, and the policies governing how the agent plans, responds to results, and recovers from failure.

You will own the evaluation and experimentation systems, contribute research ideas, and carry promising changes from prototype through production in partnership with platform engineers.

The profile is engineering-first, ML-strong, and research-capable. You should be as comfortable designing a data schema and CI pipeline as turning an ML or scientific idea into a controlled experiment and shipping the resulting improvement.

What You'll Work On
  • Research and experimentation: work with the Principal Scientist to formulate hypotheses about agent behavior, design controlled experiments, and implement and evaluate improvements to planning, tool use, and decision-making.
  • Reproducible replay: reconstruct historical agent runs from recorded context so models, prompts, and policies can be compared under consistent conditions, with clear limits on what those comparisons establish.
  • Versioned evaluation datasets: curated task sets, splits, and ground truth that evolve without invalidating past results
  • Evaluation harness and metrics: decision quality, task success, recovery from error, cost, and latency, with honest treatment of noise and small sample sizes
  • Trajectory analysis: instrument agent runs, characterize failure modes, and convert recurring failures into regression coverage
  • Control policies: prototype, test, and ship the logic that decides when the agent revises an approach, pivots, escalates to a scientist, or terminates
  • Experiment infrastructure: data pipelines, artifact and dataset versioning, experiment tracking, and the tooling that lets the team run a clean comparison in minutes rather than days
  • Close partnership with our scientists and platform engineers to connect agent behavior to what actually happens in the Workcell
What We're Looking For

Required:

  • 4-8 years building production ML systems, research infrastructure, or data-intensive backends; strong enough to design, build, and ship end-to-end
  • Excellent Python and strong software engineering discipline: testing, typing, packaging, code review, CI
  • Direct experience with ML evaluation and experimentation: offline evaluation harnesses, dataset and split design, metric design, and sound reasoning about noisy or underpowered results
  • Data pipeline skills: schema design, versioning, orchestration (Airflow, Prefect, Dagster, Ray, or equivalent), object storage
  • A reproducibility instinct: pinned environments, seeded runs, versioned artifacts, results that hold up when someone else reruns them six months later
  • The ability to take ambiguous system behavior, reduce it to a controlled experiment, and write up what the result does and does not support
  • Ability to turn ideas into working prototypes quickly, then harden successful ones into reliable production systems
  • Initiative and ownership in an environment where the roadmap is still being written

Nice to have:

  • Experience with LLM agents: tool calling, planning loops, context management, trajectory logging, and agent evaluation
  • Background in sequential decision making: bandits, Bayesian optimization, reinforcement learning, or off-policy evaluation
  • Experiment tracking, agent observability, and artifact versioning tools (Langfuse, MLflow, Weights and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary