×
Register Here to Apply for Jobs or Post Jobs. X

Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering

Job in Brampton, Ontario, C6S, Canada
Listing for: Vector Institute
Full Time position
Listed on 2026-08-01
Job specializations:
  • Research/Development
    AI Evaluation, Data Scientist, Research Scientist
Job Description & How to Apply Below

Full Time
- Professional Staff Professional Staff Toronto, ON, CA

4 days ago Requisition

Salary Range: $ To $ Annually

Applied Machine Learning Scientist
- Agent Evaluation and Harness Engineering

POSITION SUMMARY

As an Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering, you will lead applied research on evaluation, observability, stress-testing, and systematic improvement of AI agents. The role focuses on assessing agent performance and safety across long-horizon, multi-step tasks, and on building methods and tools to help organizations understand whether those systems are working, why they fail, and how to make them measurably better.

A core objective is developing adaptive evaluation approaches tailored to Canadian organizations, moving beyond static public benchmarks towards rigorous, organization-specific test environments of end-to-end agentic systems. Working alongside Vector researchers, research professionals, and external partners, the role balances high-quality applied research with the creation of practical technical systems that improve the reliability, safety, security, and effectiveness of deployed agents.

KEY RESPONSIBILITIES
  • Research and implement state-of-the-art methods for evaluating agents operating over long horizons, multiple tools, changing environments, and partially observable states;
  • Develop evaluations that assess complete agent trajectories, including planning quality, tool selection, intermediate decisions, state transitions, recovery behaviour, verification, termination decisions, resource consumption, and final outcomes;
  • Develop methods for creating organization-specific evaluations from production traces, human feedback, incidents, near misses, support interactions, domain-expert knowledge, and synthetic scenario generation;
  • Create techniques for converting discovered failures into durable regression evaluations that can be rerun across model, prompt, policy, tool, and harness changes;
  • Partner with Vector researchers, Applied ML Specialists, research professionals, and external collaborators – including Vector industry partners and members of the Canadian AI Safety Institute – to identify consequential agent use cases and create tools, reference agents, and evaluations required for trustworthy deployment;
  • Develop schemas and infrastructure for capturing structured traces of active agents;
  • Research representations of agent trajectories, such as event streams, causal graphs, tool-call graphs, state-transition graphs, and compact trajectory embeddings;
  • Develop approaches for identifying recurrent failure patterns and attributing outcomes to specific components or decisions within an agent system;
  • Build privacy-preserving and security-conscious methods for collecting and analyzing traces in sensitive organizational environments;
  • Research and build agent harnesses incorporating tools, memory, retrieval, sandboxes, permissions, validators, execution loops, recovery strategies, state management, and human approval mechanisms;
  • Develop automated or semi-automated methods for optimizing agent harnesses based on evaluation results and execution traces;
  • Develop safe mechanisms for agents to propose modifications to their own prompts, tools, policies, memory structures, workflow logic, or evaluation criteria while preserving auditability and human control;
  • Lead or contribute to peer-reviewed publications, technical reports, open-source software, benchmark releases, and reference implementations;
  • Contribute to training programs and technical workshops that help Vector partners and external stakeholders design, evaluate, debug, and govern agent systems;
  • Serve as a Vector expert on emerging methods in agent evaluation and harness engineering and connect external stakeholders with relevant members of the Vector research community; and,
  • Other related duties as assigned from time to time.
KEY SUCCESS MEASURES
  • Development of novel, scientifically rigorous evaluation methods for tool-using and long-horizon agents;
  • Release of novel evaluation tooling in collaboration with the Canadian AI Safety Institute;
  • Creation of adaptive evaluation systems that discover materially important…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary