AI Eval/Testing; Eval Engineer
Job in
Dallas, Dallas County, Texas, 75215, USA
Listed on 2026-06-19
Listing for:
NTT DATA North America
Full Time
position Listed on 2026-06-19
Job specializations:
-
Software Development
AI Engineer (Applied/Software), AI QA / Validation Engineer
Job Description & How to Apply Below
Req
NTT DATA seeks an AI Eval / Testing (Eval Engineer) in Dallas, Texas, United States.
Experience level: 10+ years.
Job SummaryWe are looking for an AI Evaluation & Test Engineer to ensure generative AI models and applications are safe, accurate, trustworthy, and deliver an elegant user experience.
Responsibilities- Build and maintain AI evaluation pipelines to test, measure, and evaluate the behavior and performance of AI systems.
- Implement traces, spans, and session tracking for observability and identify error propagation in multi-step pipelines.
- Define AI quality metrics and KPIs around factuality, faithfulness, toxicity, grounding precision/recall, latency, cost, etc., with clear acceptance bars.
- Implement evaluation and testing automation to enable end-to-end system and regression testing at scale.
- Define criteria for and implement release gates in the CI/CD pipeline.
- Find creative ways to break products.
- Assist in root cause analysis and troubleshooting of bugs and field issues.
- Collaborate with cross-functional teammates from product, engineering, linguistics, and customer support to shape human-AI interaction paradigms and ensure that our AI models and applications deliver the desired outcome and user experience.
- AI Platform Admin (M365, copilot Studio): manages AI platforms and environments, including access provisioning, governance controls, and policy enforcement (e.g., DLP, security, and compliance).
- AI Reusable Utility: develops reusable components (e.g., prompts, connectors, APIs, templates) to accelerate AI solution delivery and promote standardization across use cases.
- AI Common Infrastructure, Framework & Observability Architect (AWS and Azure): designs and maintains the foundational AI infrastructure, frameworks, and observability capabilities (telemetry, monitoring, metrics) required for scalable, reliable, and governed AI operations.
- Adversarial Testing (Red Teaming): design prompts to manipulate agent behavior, stress‑test edge cases, and expose security vulnerabilities (e.g., prompt injection or PII leakage) before deployment.
- Pipeline Automation: build and maintain automated regression testing, CI/CD release gates, and testing data sets (golden sets) to measure system drift.
- Grader Development: implement "LLM-as-a-judge" frameworks, rule‑based checks, and human‑in‑the‑loop scoring rubrics to objectively evaluate open‑ended AI outputs.
- Root Cause Analysis: trace multi‑turn conversations and agent tool interactions to diagnose when and why the AI chose the wrong path.
- Metric Definition: establish and monitor AI KPIs such as factual accuracy, latency, cost, and grounding precision.
- Programming: 5+ years of strong proficiency in Python and testing frameworks like pytest.
- AI & LLM Frameworks: 5+ years of hands‑on experience with evaluation tools like Lang Smith, Deep Eval, Tru Lens, or Promptfoo.
- Orchestration Tools: 3 to 5 years of familiarity with agentic workflows built on Lang Chain, CrewAI, or Llama Index.
- Observability: understanding of tracing and session tracking to map how errors propagate in RAG systems.
- 5+ years of strong software testing fundamentals and expertise in writing test plans, executing test cases, and generating detailed reports and dashboards.
- Strong analytical and debugging skills, and attention to detail.
- 5+ years of proficiency in Python, scripting, and software testing automation frameworks and tools such as Pytest, Selenium, Robot Framework, etc.
- Working knowledge of generative AI models, AI agents, and related concepts such as retrieval‑augmented generation (RAG), prompt engineering, context engineering, explainability, traceability, observability, guard rails, reasoning, specificity, etc.
- Sound understanding of the fundamental differences in the approach for testing conventional software versus evaluating generative AI systems.
- Team player with excellent interpersonal skills and the ability to collaborate effectively with remote and cross‑functional team members.
- Go‑getter attitude and ability to flourish in a fast‑paced, startup environment.
- Experience in any of the following would be a big…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×