More jobs:
Agentic AI Optimization Developer
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-08-16
Listing for:
United States Digital Space LLC
Full Time
position Listed on 2026-08-16
Job specializations:
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer, AI QA / Validation Engineer, Backend Developer
Job Description & How to Apply Below
Canada
Technology
Full time
8/12/2026
J
To play this video, please accept Functional & Personalisation cookies in your cookie preferences.
Manage Preferences
the company is where you can power your possible. If you want to achieve your true potential, chart new paths, develop new skills, collaborate with bright minds, and make a meaningful impact, we want to hear from you.
Synopsis of the role At the company, we are moving past passive AI chat interfaces to build the future of autonomous workflows. We are creating intelligent, self-correcting multi-agent systems that can navigate complex software environments, utilize external tools, and solve open-ended business problems with minimal human intervention. To ensure these systems are safe, reliable, and enterprise-grade, we are seeking an analytical Agentic AI Evaluation & Tuning Engineer.
In this role, you will be the guardian of our production AI reliability. You will bridge the gap between raw Large Language Model (LLM) capabilities and flawless autonomous execution. Unlike traditional software testers or prompt engineers, you will focus on the behavior, decision‑making logic, tool‑use efficiency, and long‑term stability of multi‑agent architectures. Your mission is to build the automated evaluation frameworks that keep our agents accurate, cost‑effective, and hallucination‑free.
What you will do Golden Dataset Curation & Automated Evaluation
Build the Golden Set :
Curate, maintain, and augment high‑quality reference datasets (Golden Sets) of documents, user queries, and expected agent trajectories to serve as the ultimate source of truth for testing.
Automate Eval Cycles:
Design and implement automated, continuous evaluation pipelines to measure agent accuracy, latency, token spend, and fallback reliability before code hits production.
Trajectory & Reasoning Auditing:
Trace and dissect complex, multi‑step agent thought processes (e.g., ReAct, Reflection loops) to pinpoint exactly where an agent deviates from its intended logic path.
Agent Tuning & Developer Collaboration
Behavioral Optimization:
Refine system prompts, context windows, and few‑shot examples to optimize how agents execute complex, multi‑step workflows.
Tool & Function-Calling Optimization:
Fine‑tune how agents interact with external APIs, databases, and UiPath RPA workflows—minimizing execution errors, redundant calls, and token overhead.
Augment Development:
Partner closely with AI Solution Leads and AI Agent Developers to feed evaluation insights back into the development lifecycle, helping them build robust, reusable, and self‑correcting agent components.
Production Guardrails & Lifecycle Management (LLMOps)
Defeat Drift & Hallucinations:
Actively monitor deployed agents to identify, troubleshoot, and mitigate semantic drift, prompt injections, infinite execution loops, and hallucinations.
Maintain Autonomous Integrity:
Implement robust guardrail frameworks to ensure agents maintain reliable, fact‑based autonomous decision‑making post‑deployment in production.
RAG & Knowledge Integration:
Optimize Domain‑Specific Knowledge Bases and Retrieval-Augmented Generation (RAG) pipelines to ensure agents pull from accurate data rather than assumptions.
What Experience You Need
Experience:
3+ years of professional experience in software quality engineering, test automation, or data/ML engineering, with a dedicated focus on LLM testing, prompt tuning, or orchestration patterns over the last 1–2 years.
Agentic & LLM Frameworks:
Proven hands‑on experience working with LLM orchestration frameworks (e.g., Lang Graph, ADKs or specialized internal SDKs).
Function Calling Mastery:
Deep understanding of JSON schema design for LLM tool‑calling, function‑calling, and structured outputs.
Advanced Debugging & Automation:
Strong background in writing automated test scripts (Python‑heavy) and using tracing/observability concepts to debug cascading errors in asynchronous, non‑deterministic systems.
What Could Set You Apart
Experience with AI evaluation and observability platforms
Live production experience testing Agentic workflows and GenAI solutions
Familiarity with Google Cloud AI suite…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×