×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

AI TEVV Engineer

Job in Ithaca, Tompkins County, New York, 14850, USA
Listing for: Ursa Space Systems
Full Time position
Listed on 2026-08-09
Job specializations:
  • IT/Tech
    AI Evaluation
Job Description & How to Apply Below

AI TEVV Engineer (Test, Evaluation, Verification & Validation)

Ursa Space Systems is building an AI-native geospatial insights platform that guides the acquisition, analysis, and integration of satellite and geospatial data into customer workflows, giving decision makers an edge. Leveraging hundreds of data sources, AI agents, and proprietary analytics, Ursa Space provides fast, actionable information to a range of industries, including finance, energy, and defense. Our customers receive contextual, comprehensive reporting that goes beyond surface-level observations.

Ursa Space is looking for an AI TEVV Engineer to define how we prove our AI-native geospatial platform works and to whom. This is an evaluation-science role at its core, not a test-automation role. The central skill is measurement under uncertainty: where classical QA asks "does this function return the correct output" (a deterministic pass/fail question), AI evaluation asks "what is the error rate, on what distribution of inputs, under what operating conditions, and is that rate acceptable for this mission" (a measurement question with confidence intervals).

That work is closer to experimental design and psychometrics than to writing test suites. You will own the strategy, methodology, and evidence that let us and our customers' risk officers trust the platform's outputs and demonstrate where, and how well, they hold.

This position reports to the Director of System Requirements, with a functional reporting line and charter that preserve evaluation independence from the teams whose outputs are under test. This position is fully remote and exempt.

Responsibilities
  • Design, develop, and plan the TEVV strategy across the platform's algorithms, AI/ML models, agentic workflows, data pipelines, and analytic products, aligned to the NIST AI RMF Measure function
  • Design statistically defensible evaluations: error metrics and acceptance criteria on representative input distributions, with explicit confidence intervals
  • Define and document the platform's context of use — the validated operating envelope (modalities, geographies, resolutions, conditions, target classes) within which accuracy claims hold
  • Evaluate ground-truth and "golden" datasets, including annotation and adjudication protocols, inter-rater reliability, and quantified uncertainty in the reference data itself
  • Implement a layered evaluation posture: a verifiable core (accuracy, groundedness, format), a rubric-scored middle layer with documented inter-rater reliability, and an honest residual of expert holistic review
  • Evaluate generative and natural-language outputs for claim-level groundedness whether each assertion is traceable to a citable source alongside rubric-based, human-adjudicated assessment
  • Stand up continuous monitoring and re-validation certification gates plus ongoing surveillance watching for model, prompt, retrieval, and agent-behavior drift
  • Author and maintain the TEVV evidence set: test plans, traceability matrices, metrics, acceptance criteria, and credibility-assessment documentation
  • Support DoD AI test-and-evaluation expectations (including DoD Directive 3000.09), contractual milestones, acceptance testing, and demonstrations to government stakeholders.
  • Distinguish internal TEVV from organizationally independent IV&V, and partner with external IV&V agents where required
  • Partner with Engineering teams to embed evaluability, observability, and traceability from design onward.
  • Contribute to emerging standards (NIST AI TEVV consortium, ISO/IEC SC 42 / 42001), aligning our methodology so evidence packages map to customers' compliance frameworks.
  • 30% travel.
  • Perform all other duties as assigned.
Requirements
  • B.S. in Computer Science, Statistics, or Systems Engineering, or a related quantitative discipline (M.S./Ph.D. a plus)
  • 10+ years of relevant experience, centered on evaluation, measurement, or test-and-evaluation of AI/ML or data-driven systems — not solely software QA or test automation
  • Demonstrated ability to design statistically defensible evaluations: input-distribution design, error-rate estimation, confidence intervals, and context-tied acceptance criteria
  • Hands-on…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary