×
Register Here to Apply for Jobs or Post Jobs. X

Applied Research Engineer

Job in New York, New York County, New York, 10261, USA
Listing for: Verona
Full Time position
Listed on 2026-08-30
Job specializations:
  • Research/Development
    AI Evaluation, Research Scientist
Salary/Wage Range or Industry Benchmark: 90000 - 130000 USD Yearly USD 90000.00 130000.00 YEAR
Job Description & How to Apply Below
Location: New York

Applied research at Verona begins with a demanding real-world question: how do we know an AI system is useful, reliable, and improving on the workflow that matters? You will develop the evaluation methods, datasets, experiments, and learning systems that make those answers rigorous.

This role sits between research and production engineering. You will study failures from live enterprise deployments, form precise hypotheses about model and system behavior, build the infrastructure needed to test them, and turn the results into systems that improve continuously in the field.

What you’ll do
  • Develop evaluations for task success, answer quality, robustness, safety, cost, and the business outcome a system is intended to improve.
  • Build high-signal datasets, simulations, adversarial tests, and human-review protocols that capture what actually matters in a customer workflow.
  • Establish evaluation integrity by validating automated judges, monitoring drift, and making sure reported improvements represent real progress.
  • Investigate production traces and unexpected failures deeply enough to explain why they happened and which intervention is most likely to work.
  • Partner with forward deployed engineers to move research ideas into customer systems and measure their impact under real operating conditions.
  • Turn field results into reusable methods, internal standards, research infrastructure, and product capabilities across Verona.
What we’re looking for
  • A record of strong applied machine-learning research or engineering work, demonstrated through shipped systems, publications, open-source contributions, or equivalent projects.
  • Excellent experimental judgment: you can turn surprising system behavior into testable hypotheses and distinguish a meaningful result from a misleading metric.
  • Fluency in Python and the engineering ability to build durable research and evaluation infrastructure, not only one-off notebooks.
  • Strong foundations in statistics, machine learning, and the practical limitations of automated evaluation.
  • Clear technical communication and the ability to collaborate with researchers, production engineers, customer teams, and domain experts.
  • A bias toward research whose value can be observed in deployed systems and real user outcomes.
You might excel here if
  • Experience with LLM or agent evaluation, post-training, red teaming, synthetic data, reward modeling, or human-feedback systems.
  • Experience designing evaluations for open-ended workflows where quality is subjective, multidimensional, or difficult to observe directly.
  • Systems-level understanding of how models, retrieval, tools, prompts, data, and application code interact in production.
  • Experience working with messy domain data or subject-matter experts to turn tacit judgment into reliable measurement.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary