×
Register Here to Apply for Jobs or Post Jobs. X

Manager, AI Engineering; Tester

Job in O'Fallon, St. Charles County, Missouri, 63366, USA
Listing for: MasterCard
Full Time position
Listed on 2026-07-26
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Position: Manager, AI Engineering (Tester )
Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title and Summary

Manager, AI Engineering (Tester )

Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready.

This role will lead AI quality engineering efforts - defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates - while fostering a culture of rigorous, first-class AI testing across the program.

Roles and Responsibilities:

* Design and own end-to-end LLM evaluation frameworks - including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.

* Build comprehensive test suites for agentic AI systems - validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.

* Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, Tru Lens, and Deep Eval.

* Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation - building and maintaining an evolving adversarial test library.

* Execute fairness, bias, and Responsible AI audits - testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.

* Design and run inference performance benchmarks - measuring latency, throughput, token efficiency, and degradation under peak load - and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).

* Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, Cloud Watch).

* Define the AI testing roadmap and quality standards for the program - establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI work streams.

* Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one - reviewing prompt architectures, agent designs, and system workflows for testability and risk.

* Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.

All

About You:

* Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.

* Demonstrated expertise testing LLM and Gen AI systems - including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.

* Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, Deep Eval, Tru Lens, Lang Smith, Prompt Flow, Weights & Biases Evals, or equivalent platforms.

* Strong understanding of Gen AI failure modes - hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures - and proven methods to surface and document them systematically.

* Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.

* Working knowledge of LLM ecosystems - OpenAI, Anthropic, Hugging Face, Lang Chain/Lang Graph - sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.

* Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, Sage Maker) and experience integrating automated quality gates into CI/CD workflows for AI systems.

* Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production.

* Strong analytical, communication, and stakeholder management skills - with the ability to translate complex AI failure patterns…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary