×
Register Here to Apply for Jobs or Post Jobs. X

AI Engineer Sr

Job in Toronto, Ontario, C6A, Canada
Listing for: Dayforce
Full Time position
Listed on 2026-10-08
Job specializations:
  • Software Development
    AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 120000 - 180000 CAD Yearly CAD 120000.00 180000.00 YEAR
Job Description & How to Apply Below
About the opportunity

We are looking for an AI Engineer - Agentic Systems Evaluation to help define how we measure, test, and improve the quality of enterprise AI systems.

As AI evolves from conversational assistants and Retrieval-Augmented Generation (RAG) into tool using agents, multi-agent systems, and autonomous business workflows, evaluating only the final response is no longer enough.

An agent may reach the right answer while choosing the wrong tool, taking unnecessary steps, retrieving incorrect context, failing to escape to a human or violating business rules.

This role will build the evaluation frameworks needed to understand not only whether an AI system succeeded, but how it succeeded, how reliably it can repeat that outcome, and whether the architecture is appropriate for the problem.

What you’ll get to do

Build Agentic Evaluation Frameworks

Design evaluation methodologies covering the complete AI execution lifecycle:

Intent Planning Retrieval Tool Use Reasoning Action Business Outcome

Evaluate systems across dimensions including:

  • task and business outcome accuracy
  • planning and decision quality
  • retrieval quality and groundedness
  • tool selection and execution
  • agent routing and delegation
  • human escalation decisions
  • reliability and failure recovery
  • safety and policy compliance
  • latency, token consumption, and cost

Evaluate Agentic Design Patterns

Design experiments and benchmarks that help engineering teams determine which architecture works best for a given problem.

Evaluate patterns such as:

  • Single-agent vs. multi-agent systems
  • Supervisor/router architectures
  • Planner–executor patterns
  • Sequential and parallel workflows
  • Tool-using agents
  • Human-in-the-loop workflows
  • Long-running agents with state and memory

Measure whether additional agent complexity actually improves task success, reliability, and business outcomes enough to justify increased latency, cost, and operational complexity.

Advance RAG & Knowledge Evaluation

Build rigorous evaluation approaches for enterprise retrieval and knowledge systems, including:

  • retrieval precision and relevance
  • groundedness and citation accuracy
  • source authority and freshness
  • chunking, metadata, and indexing strategies
  • semantic vs. hybrid search and reranking
  • permission-aware retrieval

Evaluate how retrieval decisions ultimately impact downstream agent performance, rather than treating RAG evaluation as an isolated problem.

Build Automated Evaluation & Regression Testing

Develop scalable evaluation infrastructure including:

  • golden and synthetic datasets
  • scenario and adversarial test suites
  • deterministic graders
  • LLM-as-a-Judge evaluation
  • human evaluation workflows
  • trace-based evaluation
  • automated regression testing

Integrate evaluations into AI development and release pipelines so changes to models, prompts, retrieval, tools, or agent architectures can be measured before reaching production.

Build Agent Trace & Failure Analysis

Analyze complete agent execution traces including planning, retrieved context, tool calls, handoffs, retries, exceptions, latency, and cost.

Develop failure taxonomies that distinguish between:

Model | Retrieval | Planning | Tool | Routing | Memory | Integration | Policy | Orchestration failures

Turn production failures and user feedback into measurable regression tests and engineering improvements.

Define Production AI Quality

Establish measurable quality standards and release criteria for AI systems.

Metrics may include Task Success Rate, First-Pass Success Rate, Tool Selection Accuracy, Agent Routing Accuracy, Plan Execution Fidelity, Failure Recovery Rate, Human Escalation Accuracy, Business Outcome Accuracy, Cost / Latency per Successful Task

Help teams answer a fundamental question:

Is this AI system reliable, safe, efficient, and valuable enough to operate in…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary