More jobs:
Postdoctoral Scholar, AI Evaluation & Standards
Job in
Lakewood, Ocean County, New Jersey, 08701, USA
Listed on 2026-08-15
Listing for:
Jobtailor
Full Time
position Listed on 2026-08-15
Job specializations:
-
Research/Development
AI Evaluation, Data Scientist -
IT/Tech
AI Evaluation, AI Engineer (Applied/Software), Data Scientist, Machine Learning/ ML Engineer
Job Description & How to Apply Below
- Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
- Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
- Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
- Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
- Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
- Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
- Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
- Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
- Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
- Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
- Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.
- PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
- Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
- Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
- Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
- Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
- Proficiency in Python and common data science or machine learning tools.
- Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
- Clear written and verbal communication skills.
Demonstrates expertise in designing evaluation frameworks and criteria for GenAI tools in pharmaceutical R&D, with a strong foundation in biomedical science and data analysis. Proficient in translating expert judgment into measurable outcomes and improving evaluation methods through collaboration with cross-functional teams.
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×