×
Register Here to Apply for Jobs or Post Jobs. X

Manager, AI Engineering; Tester

Job in O'Fallon, St. Charles County, Missouri, 63366, USA
Listing for: Mastercard
Full Time position
Listed on 2026-07-30
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 231000 USD Yearly USD 140000.00 231000.00 YEAR
Job Description & How to Apply Below
Position: Manager, AI Engineering (Tester )

Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title

And Summary

Manager, AI Engineering (Tester)

Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready.

This role will lead AI quality engineering efforts — defining evaluation frameworks, red team strategies, and LLMOps quality gates — while fostering a culture of rigorous, first-class AI testing across the program.

Roles And Responsibilities
  • Design and own end-to-end LLM evaluation frameworks — including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
  • Build comprehensive test suites for agentic AI systems — validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
  • Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, Tru Lens, and Deep Eval.
  • Lead structured red team and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation — building and maintaining an evolving adversarial test library.
  • Execute fairness, bias, and Responsible AI audits — testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
  • Design and run inference performance benchmarks — measuring latency, throughput, token efficiency, and degradation under peak load — and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
  • Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, Cloud Watch).
  • Define the AI testing roadmap and quality standards for the program — establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI work streams.
  • Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one — reviewing prompt architectures, agent designs, and system workflows for testability and risk.
  • Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, Truthful QA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.
All About You
  • Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
  • Demonstrated expertise testing LLM and Gen AI systems — including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
  • Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, Deep Eval, Tru Lens, Lang Smith, Prompt Flow, Weights & Biases Evals, or equivalent platforms.
  • Strong understanding of Gen AI failure modes — hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures — and proven methods to surface and document them systematically.
  • Strong Python programming skills with the ability to…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary