Manager, AI Engineering; Tester
Listed on 2026-07-27
-
Software Development
AI Engineer (Applied/Software), AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Manager, AI Engineering (Tester)
Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready.
This role will lead AI quality engineering efforts — defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates — while fostering a culture of rigorous, first-class AI testing across the program.
Roles and Responsibilities:
- Design and own end-to-end LLM evaluation frameworks — including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
- Build comprehensive test suites for agentic AI systems — validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
- Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, Tru Lens, and Deep Eval.
- Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation — building and maintaining an evolving adversarial test library.
- Execute fairness, bias, and Responsible AI audits — testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
- Design and run inference performance benchmarks — measuring latency, throughput, token efficiency, and degradation under peak load — and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
- Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, Cloud Watch).
- Define the AI testing roadmap and quality standards for the program — establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI work streams.
- Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one — reviewing prompt architectures, agent designs, and system workflows for testability and risk.
- Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, Truthful QA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.
All
About You:
- Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
- Demonstrated expertise testing LLM and Gen AI systems — including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
- Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, Deep Eval, Tru Lens, Lang Smith, Prompt Flow, Weights & Biases Evals, or equivalent platforms.
- Strong understanding of Gen AI failure modes — hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures — and proven methods to surface and document them systematically.
- Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.
- Working knowledge of LLM ecosystems — OpenAI, Anthropic, Hugging Face, Lang Chain/Lang Graph — sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
- Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, Sage Maker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
- Experience with…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).