Natural Language Measurement Specialist in Houston, Texas
Listed on 2026-07-28
-
IT/Tech
AI Evaluation, Data Scientist, Machine Learning/ ML Engineer
Natural Language Specialist
This is a remote role. Candidates who live near CB offices have the option of being fully remote or hybrid (Tuesday and Wednesday in office). All CB employees are required to occasionally travel to meet in person for business purposes.
This is a full-time position
About the OpportunityAs a Natural Language Specialist, you will help define and advance how language-based performance is measured in high-stakes educational settings. This role sits at the intersection of natural language processing, large language models, and educational measurement, and is ideal for someone who pairs strong measurement training with a working knowledge of modern AI. You will translate measurement constructs into LLM-based feedback and annotation systems, design and conduct the studies that establish their validity, fairness, and reliability, and ensure that what our models produce holds up to rigorous psychometric standards.
You will serve as the measurement authority for cross-functional partners-psychometricians, engineers, and data scientists-shaping how language-based constructs are defined, evaluated, and applied across the team's portfolio of systems. Your primary lens will be measurement: defining what feedback and annotations mean, evidencing that they mean it, and improving them over time.
- Define and operationalize language-based constructs for automated annotation and feedback generation
- Apply psychometric principles-reliability, validity, dimensionality, and measurement in variance-to LLM-based feedback and annotation systems
- Design and lead validity studies, including human-machine agreement, rater comparison, and fairness analyses across subgroups
- Develop and apply methods for detecting and mitigating bias in language-based scores
- Establish/Recommend annotation guidelines, feedback quality criteria, and standards for acceptable model performance
- Translate measurement requirements into specifications that guide model development and evaluation
- Contribute to prompt design, fine-tuning, and evaluation of LLM-based feedback and annotation systems
- Develop and refine machine learning models for measuring language-based constructs
- Build evaluation frameworks that connect model behavior to measurement outcomes
- Collaborate with senior team members to translate measurement findings into production systems
- Implement high-quality, maintainable code for model development and evaluation
- Lead and contribute to research studies that evaluate model performance and support assessment validity
- Apply statistical and psychometric methods to analyze results and inform model improvements
- Document methodologies and findings in a clear and rigorous manner
- Stay current with advances in educational measurement, NLP, and learning science
- Prepare and curate datasets for measurement studies and model evaluation
- Support reproducible data processing workflows for training, evaluation, and monitoring
- Partner with engineers to integrate feedback and annotation models into scalable systems
- Collaborate closely with psychometricians, data scientists, and engineers
- Contribute to documentation, methodological standards, and team best practices
- Participate in peer reviews and knowledge sharing
- Actively raise the measurement literacy of the broader team-mentoring junior ICs, providing technical feedback on colleagues' work, and building shared standards
You bring strong measurement training and a genuine interest in how modern AI can be used to measure language-based performance. You are excited about applying psychometric rigor to large language models in high-stakes educational settings. You have:
- A Master's or PhD (or near completion) in a quantitative field such as Psychometrics, Educational Measurement, Quantitative Psychology, Statistics, Data Science, or a related discipline (measurement-focused training strongly valued)
- A solid foundation in measurement theory, including reliability, validity, and fairness; familiarity with IRT, generalizability theory, or related frameworks is highly desirable
- Experience analyzing language or text data, with exposure to NLP or large language models
- Programming skills in Python and familiarity with data science libraries (e.g., pandas, Num Py, PyTorch)
- Demonstrated ability to conduct rigorous research (e.g., thesis, publications, or applied research projects)
- Strong skills in statistical analysis and experimental design
- Familiarity with working with structured and unstructured data
- Strong attention to detail and a commitment to producing high-quality, reproducible work
- The ability to travel 5-10 times a year to College Board offices or on behalf of College Board business.
You are:
- Curious and eager to learn new tools, methods, and domains
- Thoughtful about the implications of AI systems, including fairness and validity
- Able to communicate…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).