×
Register Here to Apply for Jobs or Post Jobs. X

Manager, AI Engineering; Tester

Job in O'Fallon, St. Charles County, Missouri, 63366, USA
Listing for: Dynamic Yield
Full Time position
Listed on 2026-07-27
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Position: Manager, AI Engineering (Tester )

Manager, AI Engineering (Tester)

Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready.

This role will lead AI quality engineering efforts — defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates — while fostering a culture of rigorous, first-class AI testing across the program.

Roles and Responsibilities:

  • Design and own end-to-end LLM evaluation frameworks — including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
  • Build comprehensive test suites for agentic AI systems — validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
  • Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, Tru Lens, and Deep Eval.
  • Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation — building and maintaining an evolving adversarial test library.
  • Execute fairness, bias, and Responsible AI audits — testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
  • Design and run inference performance benchmarks — measuring latency, throughput, token efficiency, and degradation under peak load — and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
  • Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, Cloud Watch).
  • Define the AI testing roadmap and quality standards for the program — establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI work streams.
  • Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one — reviewing prompt architectures, agent designs, and system workflows for testability and risk.
  • Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, Truthful QA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.

All

About You:

  • Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
  • Demonstrated expertise testing LLM and Gen AI systems — including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
  • Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, Deep Eval, Tru Lens, Lang Smith, Prompt Flow, Weights & Biases Evals, or equivalent platforms.
  • Strong understanding of Gen AI failure modes — hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures — and proven methods to surface and document them systematically.
  • Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.
  • Working knowledge of LLM ecosystems — OpenAI, Anthropic, Hugging Face, Lang Chain/Lang Graph — sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
  • Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, Sage Maker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
  • Experience with…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary