×
Register Here to Apply for Jobs or Post Jobs. X

Applied AI Developer; Agent Evaluation

Job in Vancouver, BC, Canada
Listing for: Autodesk
Full Time position
Listed on 2026-07-29
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 99000 - 145200 CAD Yearly CAD 99000.00 145200.00 YEAR
Job Description & How to Apply Below
Position: Applied AI Developer (Agent Evaluation)

Job Requisition  #

26WD96920

Position Overview

As an Applied AI Developer – Agent Evaluation on the Fusion Platform Services team within Product Development and Manufacturing Solutions (PDMS), you'll be part of a team of technologists dedicated to creating cutting-edge AI and generative AI solutions that enhance developer productivity and experience. You'll work closely with AI engineers, software architects, and product engineering teams to build, deploy, and rigorously evaluate intelligent agentic systems while developing MCP (Model Context Protocol)-based tooling that integrates seamlessly with IDEs such as VS Code and Cursor.

Responsibilities
  • Develop and orchestrate multi-agent AI systems for automated test generation, test execution, and end-to-end development workflow optimization using frameworks such as Lang Graph, Auto Gen, or the Anthropic Agent SDK (Claude Code).
  • Design and implement agentic workflows that coordinate multiple AI agents to autonomously drive test automation across UI, API, integration, and system levels, from test case synthesis to result evaluation, while integrating seamlessly with existing developer tools and MCP-compatible services.
  • Build evaluation frameworks and benchmark suites for agentic systems, including comparisons of AI agents against commercial or domain-specific solutions, using tools such as Agent Bench and Langfuse.
  • Evaluate and optimize MCP servers and tool integrations across agentic pipelines by measuring latency, accuracy, context fidelity, and end-to-end task completion.
Minimum Qualifications
  • BS/MS in Computer Science, Machine Learning, or a related applied AI field.
  • Expertise in Python and ML frameworks (PyTorch, Transformers, scikit-learn).
  • Experience with Large Language Models applied to software understanding, test generation, or agentic AI systems.
  • Knowledge of AI evaluation methodologies and metrics for agentic task completion and test quality.
  • Strong foundation in statistical analysis and experimental design.
  • Experience with developer workflow and productivity measurement frameworks.
Preferred Qualifications
  • Background in software engineering or QA with close collaboration with development teams.
  • Familiarity with test automation frameworks (e.g., Playwright, Selenium, Pytest, Appium) and CI/CD pipelines.
  • Experience designing evaluation frameworks or benchmark suites for AI agents against commercial or domain-specific solutions.
  • Hands-on experience with MCP (Model Context Protocol), including building, evaluating, and optimizing MCP servers and tool integrations.
  • Experience using AI evaluation and observability tools such as Lang Smith, Langfuse, Agent Bench, RAGAS, Deep Eval, or similar platforms.
  • Experience with Azure AI Foundry/ML or AWS cloud ML platforms.

____________________________________________________

Développeur IA Appliquée – Évaluation d'Agents IA Présentation du poste

En tant que Développeur IA Appliquée – Évaluation d'Agents IA au sein de l'équipe Fusion Platform Services de Product Development and Manufacturing Solutions (PDMS), vous rejoindrez une équipe de technologues dédiée à la création de solutions d'intelligence artificielle et d'IA générative de pointe visant à améliorer la productivité et l'expérience des développeurs. Vous collaborerez étroitement avec des ingénieurs IA, des architect es logiciels et des équipes de développement produit afin de concevoir, déployer et évaluer rigoureusement des systèmes d'IA agentiques, tout en développant des outils basés sur le Model Context Protocol (MCP) intégrés à des environnements tels que VS Code et Cursor.

Responsabilités
  • Développer et orchestrer des systèmes d'IA multi-agents pour la génération automatique de tests, leur exécution et l'optimisation des workflows de développement à l'aide de frameworks tels que Lang Graph, Auto Gen ou l'Anthropic Agent SDK (Claude Code).
  • Concevoir et mettre en œuvre des workflows agentiques permettant à plusieurs agents IA de piloter de manière autonome l'automatisation des tests (UI, API, intégration et système), depuis la génération des cas de test jusqu'à l'évaluation des résultats, tout en assurant une intégration fluide avec les outils de développement…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary