×
Register Here to Apply for Jobs or Post Jobs. X

AI Eval​/Testing; Eval Engineer

Job in Dallas, Dallas County, Texas, 75215, USA
Listing for: NTT DATA North America
Full Time position
Listed on 2026-06-19
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 100000 - 125000 USD Yearly USD 100000.00 125000.00 YEAR
Job Description & How to Apply Below
Position: AI Eval / Testing (Eval Engineer)

Req
NTT DATA seeks an AI Eval / Testing (Eval Engineer) in Dallas, Texas, United States.

Experience level: 10+ years.

Job Summary

We are looking for an AI Evaluation & Test Engineer to ensure generative AI models and applications are safe, accurate, trustworthy, and deliver an elegant user experience.

Responsibilities
  • Build and maintain AI evaluation pipelines to test, measure, and evaluate the behavior and performance of AI systems.
  • Implement traces, spans, and session tracking for observability and identify error propagation in multi-step pipelines.
  • Define AI quality metrics and KPIs around factuality, faithfulness, toxicity, grounding precision/recall, latency, cost, etc., with clear acceptance bars.
  • Implement evaluation and testing automation to enable end-to-end system and regression testing at scale.
  • Define criteria for and implement release gates in the CI/CD pipeline.
  • Find creative ways to break products.
  • Assist in root cause analysis and troubleshooting of bugs and field issues.
  • Collaborate with cross-functional teammates from product, engineering, linguistics, and customer support to shape human-AI interaction paradigms and ensure that our AI models and applications deliver the desired outcome and user experience.
Platform & Enablement Roles
  • AI Platform Admin (M365, copilot Studio): manages AI platforms and environments, including access provisioning, governance controls, and policy enforcement (e.g., DLP, security, and compliance).
  • AI Reusable Utility: develops reusable components (e.g., prompts, connectors, APIs, templates) to accelerate AI solution delivery and promote standardization across use cases.
  • AI Common Infrastructure, Framework & Observability Architect (AWS and Azure): designs and maintains the foundational AI infrastructure, frameworks, and observability capabilities (telemetry, monitoring, metrics) required for scalable, reliable, and governed AI operations.
Core Responsibilities
  • Adversarial Testing (Red Teaming): design prompts to manipulate agent behavior, stress‑test edge cases, and expose security vulnerabilities (e.g., prompt injection or PII leakage) before deployment.
  • Pipeline Automation: build and maintain automated regression testing, CI/CD release gates, and testing data sets (golden sets) to measure system drift.
  • Grader Development: implement "LLM-as-a-judge" frameworks, rule‑based checks, and human‑in‑the‑loop scoring rubrics to objectively evaluate open‑ended AI outputs.
  • Root Cause Analysis: trace multi‑turn conversations and agent tool interactions to diagnose when and why the AI chose the wrong path.
  • Metric Definition: establish and monitor AI KPIs such as factual accuracy, latency, cost, and grounding precision.
Required Skills & Tech Stack
  • Programming: 5+ years of strong proficiency in Python and testing frameworks like pytest.
  • AI & LLM Frameworks: 5+ years of hands‑on experience with evaluation tools like Lang Smith, Deep Eval, Tru Lens, or Promptfoo.
  • Orchestration Tools: 3 to 5 years of familiarity with agentic workflows built on Lang Chain, CrewAI, or Llama Index.
  • Observability: understanding of tracing and session tracking to map how errors propagate in RAG systems.
  • 5+ years of strong software testing fundamentals and expertise in writing test plans, executing test cases, and generating detailed reports and dashboards.
  • Strong analytical and debugging skills, and attention to detail.
  • 5+ years of proficiency in Python, scripting, and software testing automation frameworks and tools such as Pytest, Selenium, Robot Framework, etc.
  • Working knowledge of generative AI models, AI agents, and related concepts such as retrieval‑augmented generation (RAG), prompt engineering, context engineering, explainability, traceability, observability, guard rails, reasoning, specificity, etc.
  • Sound understanding of the fundamental differences in the approach for testing conventional software versus evaluating generative AI systems.
  • Team player with excellent interpersonal skills and the ability to collaborate effectively with remote and cross‑functional team members.
  • Go‑getter attitude and ability to flourish in a fast‑paced, startup environment.
  • Experience in any of the following would be a big…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary