×
Register Here to Apply for Jobs or Post Jobs. X

Data Scientist - Evaluating Foundation Models and AI Agents National Security

Job in Richland, Benton County, Washington, 99352, USA
Listing for: Pacific Northwest National Laboratory
Full Time position
Listed on 2026-09-27
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), AI Evaluation, Data Scientist, Machine Learning/ ML Engineer
  • Research/Development
    AI Evaluation, Data Scientist
Salary/Wage Range or Industry Benchmark: 114000 - 182100 USD Yearly USD 114000.00 182100.00 YEAR
Job Description & How to Apply Below

Overview

At PNNL, our core capabilities are divided among major departments that we refer to as Directorates within the Lab, focused on a specific area of scientific research or other function, with its own leadership team and dedicated budget.

Our Science & Technology directorates include National Security, Integrated Discovery Sciences, and Energy and Environment. In addition, we have an Environmental Molecular Sciences Laboratory, a Department of Energy, Office of Science user facility housed on the PNNL campus.

The National Security Directorate (NSD) drives science-based, mission-focused solutions to take on complex,real-world threats to our nation and the world.

The AI and Data Analytics Division, part of NSD, combines profound domain expertise and creative integration of advanced hardware and software to deliver computational solutions that address complex data and analytic challenges. Working in multidisciplinary teams, we connect foundational research to engineering to operations, providing the tools to innovate quickly and field results faster. Our strengths are integrated across the data analytics lifecycle, from data acquisition and management to analysis and decision support.

Read more about the AI & Data Analytics division: https://www.pnnl.gov/ai-and-data-analytics

Responsibilities

As a Data Scientist in the AI and Data Analytics Division, you will work with experienced technical staff to design and implement evaluation studies, prepare mission-representative test data and scenarios, develop evaluation components and agentic harnesses, execute controlled experiments, and analyze model behavior. The work extends beyond aggregate benchmark scores to examine individual cases, meaningful data slices, component interactions, system performance under realistic operating conditions, and more.

Depending on project needs, research may also address explainable AI, interpretability, or targeted mechanistic-interpretability questions.

Key Responsibilities:

  • Implement defined test and evaluation activities for AI/ML with particular focus on pretrained transformer-based foundation models, including large language models (LLMs), vision-language models (VLMs), and the AI agents and systems that use them.
  • Prepare evaluation datasets, prompts, mission scenarios, meaningful data slices, test cases, scoring rubrics, and deterministic baselines.
  • Develop evaluation environments, test harnesses, and agentic harnesses used to exercise models, tools, retrieval systems, and end-to-end workflows.
  • Execute controlled experiments at the model, component, workflow, and system levels.
  • Analyze model behavior to identify capabilities, systematic weaknesses, failure modes, and sensitivity to data characteristics, retrieval results, tools, and operating conditions.
  • Examine agent trajectories, intermediate decisions, tool calls, retrieved evidence, memory or state, error propagation, recovery behavior, and points of human intervention.
  • Evaluate performance using aggregate metrics together with case-level analysis, behavioral categories, robustness testing, calibration and uncertainty characterization, and adversarial or off-nominal scenarios.
  • Evaluate hybrid systems that combine generative AI with deterministic NLP, computer vision, retrieval, search, rules, or classical machine-learning methods.
  • Help develop and validate automated or model-based graders against human judgments and mission-relevant criteria.
  • Apply functional explainability, model probing, or selected interpretability methods for understanding model and system behavior.
  • Document evaluation methods, assumptions, results, uncertainties, and limitations in reproducible code, visualizations, technical reports, and briefings.
  • Contribute technical material, data,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary