×
Register Here to Apply for Jobs or Post Jobs. X

Principal AI Quality Engineer

Job in Winnipeg, Manitoba, Canada
Listing for: Worky
Full Time position
Listed on 2026-08-16
Job specializations:
  • Software Development
    AI QA / Validation Engineer, AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Job Description & How to Apply Below
At Nexaminds, we’re on a mission to redefine industries with AI. We’re passionate about the limitless potential of artificial intelligence to transform businesses, streamline processes, and drive growth.

Join us on our visionary journey. We’re leading the way in AI solutions, and we’re committed to innovation, collaboration, and ethical practices. Become a part of our team and shape the future powered by intelligent machines. If you’re driven by ambition, success, fun, and learning, Nexaminds is where you belong.

Nexaminds is looking for an AI Quality Engineer to lead the validation and quality strategy for AI-powered systems, including Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) applications, and agent-based workflows. The ideal candidate has deep experience in Quality Engineering, AI evaluation, and automated validation frameworks, with a strong understanding of how to measure, monitor, and improve the reliability of non-deterministic AI systems.

This role focuses on building enterprise-grade AI validation capabilities, designing automated evaluation pipelines, implementing AI guardrails, and partnering closely with engineering, AI/ML, and data governance teams to ensure AI-powered solutions are accurate, safe, and production-ready.

Location:

Canada (Remote)

Qualifications weare looking for:

7+ years of experience  in Quality Engineering, Software Development, or a related technical discipline, including recent hands-on experience with AI/ML or LLM-based systems.

Experience  designing validation or evaluation frameworks fo r non-deterministic or probabilistic systems such as  Large Language Models (LLMs) or Machine Learning applications.

Strong understanding of  Retrieval-Augmented Generation (RAG), agentic workflows, and LLM evaluation  concepts, including hallucination detection, groundedness, relevance, and consistency.

Hands-on experience with  Large Language Models (LLMs) and prompt engineering,  including prompt design, optimization, and regression validation.

Practical experience using Claude Code, Claude SDK, or similar AI-assisted development platforms and agentic coding tools.

Proficiency with  Node.js  for building automation tools, validation pipelines, and engineering utilities.

Experience working with  MongoDB  to store, manage, and analyze evaluation results, validation data, or AI output logs.

Experience integrating automated validation processes into  CI/CD pipelines .

Experience designing semantic validation approaches that evaluate meaning, relevance, and contextual accuracy beyond traditional software testing.

Strong understanding of AI quality metrics, confidence scoring, output consistency, and production monitoring.

Solid knowledge of data validation principles, including schema validation, business rule compliance, and data consistency.

Excellent cross-functional collaboration skills, with experience partnering across Quality Engineering, AI/ML Engineering, Software Engineering, and Data Governance teams.

Nice to have:

Experience with prompt regression testing tools or frameworks.

Familiarity with AI governance, fairness, or bias-detection practices.

Experience with tools such as Playwright, Rest Sharp, or similar UI/API test automation.

Prior experience standing up a validation or evaluation function from scratch.

Exposure to confidence scoring or guardrail systems for production AI.

Experience designing or orchestrating multi-agent systems in production environments

Job duties:

Define and lead the AI validation strategy for LLMs, RAG systems, and AI agents.

Build automated validation pipelines integrated into CI/CD.

Establish evaluation frameworks covering correctness, relevance, groundedness, consistency, and hallucination rate.

Design and maintain prompt regression testing to catch silent quality degradation after model or prompt changes.

Introduce semantic validation techniques that go beyond traditional QA (evaluating meaning, not just structure).

Lead production monitoring of AI output quality and define acceptable output ranges for non-deterministic systems.

Partner with engineering teams to validate AI-generated code, test cases, and AI-assisted decision…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary