×
Register Here to Apply for Jobs or Post Jobs. X

Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering

Job in Toronto, Ontario, C6A, Canada
Listing for: vectorinstitute
Full Time position
Listed on 2026-09-17
Job specializations:
  • Research/Development
    AI Evaluation, Research Scientist
Salary/Wage Range or Industry Benchmark: 126000 - 157000 CAD Yearly CAD 126000.00 157000.00 YEAR
Job Description & How to Apply Below

POSITION SUMMARY

As an Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering, you will lead applied research on evaluation, observability, stress-testing, and systematic improvement of AI agents. The role focuses on assessing agent performance and safety across long-horizon, multi-step tasks, and on building methods and tools to help organizations understand whether those systems are working, why they fail, and how to make them measurably better.

A core objective is developing adaptive evaluation approaches tailored to Canadian organizations, moving beyond static public benchmarks towards rigorous, organization-specific test environments of end-to-end agentic systems. Working alongside Vector researchers, research professionals, and external partners, the role balances high-quality applied research with the creation of practical technical systems that improve the reliability, safety, security, and effectiveness of deployed agents.

KEY RESPONSIBILITIES
  • Research and implement state-of-the-art methods for evaluating agents operating over long horizons, multiple tools, changing environments, and partially observable states;
  • Develop evaluations that assess complete agent trajectories, including planning quality, tool selection, intermediate decisions, state transitions, recovery behaviour, verification, termination decisions, resource consumption, and final outcomes;
  • Develop methods for creating organization-specific evaluations from production traces, human feedback, incidents, near misses, support interactions, domain-expert knowledge, and synthetic scenario generation;
  • Create techniques for converting discovered failures into durable regression evaluations that can be rerun across model, prompt, policy, tool, and harness changes;
  • Partner with Vector researchers, Applied ML Specialists, research professionals, and external collaborators – including Vector industry partners and members of the Canadian AI Safety Institute – to identify consequential agent use cases and create tools, reference agents, and evaluations required for trustworthy deployment;
  • Develop schemas and infrastructure for capturing structured traces of active agents;
  • Research representations of agent trajectories, such as event streams, causal graphs, tool-call graphs, state-transition graphs, and compact trajectory embeddings;
  • Develop approaches for identifying recurrent failure patterns and attributing outcomes to specific components or decisions within an agent system;
  • Build privacy-preserving and security-conscious methods for collecting and analyzing traces in sensitive organizational environments;
  • Research and build agent harnesses incorporating tools, memory, retrieval, sandboxes, permissions, validators, execution loops, recovery strategies, state management, and human approval mechanisms;
  • Develop automated or semi-automated methods for optimizing agent harnesses based on evaluation results and execution traces;
  • Develop safe mechanisms for agents to propose modifications to their own prompts, tools, policies, memory structures, workflow logic, or evaluation criteria while preserving auditability and human control;
  • Lead or contribute to peer-reviewed publications, technical reports, open-source software, benchmark releases, and reference implementations;
  • Contribute to training programs and technical workshops that help Vector partners and external stakeholders design, evaluate, debug, and govern agent systems;
  • Serve as a Vector expert on emerging methods in agent evaluation and harness engineering and connect external stakeholders with relevant members of the Vector research community; and,
  • Other related duties as assigned from time to time.
KEY SUCCESS MEASURES
  • Development of novel, scientifically rigorous evaluation methods for tool-using and long-horizon agents;
  • Release of novel evaluation tooling in collaboration with the Canadian AI Safety Institute;
  • Creation of adaptive evaluation systems that discover materially important failures beyond those captured by static benchmarks;
  • Adoption of evaluation artefacts by Vector partners and stakeholders across Canada;
  • High-quality technical outputs, including peer-reviewed publications, technical reports, open-source tools, benchmarks, datasets, and reference architectures;
  • Effective collaboration with internal research and engineering teams and with external partners across sectors; and,
  • Meaningful contribution to Vector’s technical leadership in agent evaluation, agent observability, harness…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary