Applied AI Developer; Agent Evaluation
Listed on 2026-07-29
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Job Requisition #
26WD96920
Position OverviewAs an Applied AI Developer – Agent Evaluation on the Fusion Platform Services team within Product Development and Manufacturing Solutions (PDMS), you'll be part of a team of technologists dedicated to creating cutting-edge AI and generative AI solutions that enhance developer productivity and experience. You'll work closely with AI engineers, software architects, and product engineering teams to build, deploy, and rigorously evaluate intelligent agentic systems while developing MCP (Model Context Protocol)-based tooling that integrates seamlessly with IDEs such as VS Code and Cursor.
Responsibilities- Develop and orchestrate multi-agent AI systems for automated test generation, test execution, and end-to-end development workflow optimization using frameworks such as Lang Graph, Auto Gen, or the Anthropic Agent SDK (Claude Code).
- Design and implement agentic workflows that coordinate multiple AI agents to autonomously drive test automation across UI, API, integration, and system levels, from test case synthesis to result evaluation, while integrating seamlessly with existing developer tools and MCP-compatible services.
- Build evaluation frameworks and benchmark suites for agentic systems, including comparisons of AI agents against commercial or domain-specific solutions, using tools such as Agent Bench and Langfuse.
- Evaluate and optimize MCP servers and tool integrations across agentic pipelines by measuring latency, accuracy, context fidelity, and end-to-end task completion.
- BS/MS in Computer Science, Machine Learning, or a related applied AI field.
- Expertise in Python and ML frameworks (PyTorch, Transformers, scikit-learn).
- Experience with Large Language Models applied to software understanding, test generation, or agentic AI systems.
- Knowledge of AI evaluation methodologies and metrics for agentic task completion and test quality.
- Strong foundation in statistical analysis and experimental design.
- Experience with developer workflow and productivity measurement frameworks.
- Background in software engineering or QA with close collaboration with development teams.
- Familiarity with test automation frameworks (e.g., Playwright, Selenium, Pytest, Appium) and CI/CD pipelines.
- Experience designing evaluation frameworks or benchmark suites for AI agents against commercial or domain-specific solutions.
- Hands-on experience with MCP (Model Context Protocol), including building, evaluating, and optimizing MCP servers and tool integrations.
- Experience using AI evaluation and observability tools such as Lang Smith, Langfuse, Agent Bench, RAGAS, Deep Eval, or similar platforms.
- Experience with Azure AI Foundry/ML or AWS cloud ML platforms.
____________________________________________________
Développeur IA Appliquée – Évaluation d'Agents IA Présentation du posteEn tant que Développeur IA Appliquée – Évaluation d'Agents IA au sein de l'équipe Fusion Platform Services de Product Development and Manufacturing Solutions (PDMS), vous rejoindrez une équipe de technologues dédiée à la création de solutions d'intelligence artificielle et d'IA générative de pointe visant à améliorer la productivité et l'expérience des développeurs. Vous collaborerez étroitement avec des ingénieurs IA, des architect es logiciels et des équipes de développement produit afin de concevoir, déployer et évaluer rigoureusement des systèmes d'IA agentiques, tout en développant des outils basés sur le Model Context Protocol (MCP) intégrés à des environnements tels que VS Code et Cursor.
Responsabilités- Développer et orchestrer des systèmes d'IA multi-agents pour la génération automatique de tests, leur exécution et l'optimisation des workflows de développement à l'aide de frameworks tels que Lang Graph, Auto Gen ou l'Anthropic Agent SDK (Claude Code).
- Concevoir et mettre en œuvre des workflows agentiques permettant à plusieurs agents IA de piloter de manière autonome l'automatisation des tests (UI, API, intégration et système), depuis la génération des cas de test jusqu'à l'évaluation des résultats, tout en assurant une intégration fluide avec les outils de développement…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: