×
Register Here to Apply for Jobs or Post Jobs. X

Senior AI Test Automation Engineer

Job in Northern, Floyd County, Kentucky, USA
Listing for: Motion
Full Time position
Listed on 2026-09-06
Job specializations:
  • Software Development
    AI QA / Validation Engineer, Software Testing
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Location: Northern

Senior AI Test Automation Engineer

Summary:

The Senior AI Test Automation Engineer designs, develops, maintains, and executes automated testing and evaluation solutions for both traditional software and LLM-powered applications. Operating in a forward-deployed capacity, this role works directly with users and delivery teams to capture real-world usage patterns and feedback, translating them into a continuously growing evaluation framework that validates LLM behavior against the consistent flows users actually follow.

This role partners with delivery teams, QA, development, product, AI engineering, and the QA Center of Excellence (CoE) to establish scalable automation and evaluation practices, increase test and eval coverage, and integrate quality controls throughout the software delivery lifecycle, from pre-deployment regression testing through production observability. You must be eligible to work in the US without Visa Sponsorship.

ResponsibilitiesLLM Evaluation & Observability (Core)
  • Act as a forward-deployed quality engineer: engage directly with users and stakeholders to collect feedback, observe real usage patterns, and identify the consistent flows users follow through LLM-powered features.
  • Translate user feedback and production traces into curated evaluation datasets in Lang Smith, and continuously expand the eval framework as new feedback, edge cases, and failure modes are discovered.
  • Design, build, and maintain offline evaluation suites (regression, benchmarking, and backtesting) that gate prompt, model, and Lang Graph workflow changes before deployment.
  • Develop and calibrate evaluators, heuristic/code-based checks, LLM-as-judge evaluators, and pairwise comparisons, and validate judge reliability against human review.
  • Instrument and maintain end-to-end tracing across Lang Graph agents and workflows using Lang Smith, ensuring trace coverage, quality, and useful metadata for debugging and analysis.
  • Manage annotation queues and human-in-the-loop feedback workflows, routing interesting or problematic production runs to reviewers and feeding results back into datasets and evaluator calibration.
  • Analyze agent trajectories and multi-step Lang Graph executions (tool calls, state transitions, retrieval steps) to pinpoint failure points and distinguish nondeterministic LLM variance from genuine product defects.
  • Integrate eval runs into CI/CD pipelines so that dataset versions, experiments, and quality thresholds provide automated feedback on every relevant change.
  • Support the adoption of production auditing and monitoring capabilities — such as online evaluations on live traffic, quality drift detection, and alerting — to help teams detect issues in production (supportive to the role, not its core focus).
Test Automation (Core)
  • Design, develop, and implement automated test scripts for UI, API, integration, and regression testing, including deterministic E2E coverage of LLM-powered application surfaces.
  • Integrate automated tests into CI/CD pipelines to enable timely feedback and continuous quality validation.
  • Collaborate with QA, development, product, and business teams to translate requirements, acceptance criteria, and expected agent behaviors into effective automated test and eval coverage.
  • Analyze and triage automation test failures, differentiating framework or script issues from valid product defects, including the added dimension of expected LLM nondeterminism.
  • Support test data management (including eval dataset versioning, splits, and provenance) and help identify or resolve test-environment stability issues.
  • Report on test execution results, eval experiment outcomes, automation and eval coverage, quality trends, and risks.
  • Participate in code reviews and contribute to automation and evaluation standards, reusable components, and best practices.
  • Engage with the QA CoE to align automation and AI evaluation practices with enterprise standards while contributing domain-specific feedback, lessons learned, and continuous-improvement opportunities.
Required Experience
  • 5+ years of experience in test automation engineering, software quality assurance, or a related role.
  • 1-2+ years of hands-on experience testing…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary