Postdoctoral Scholar, AI Evaluation & Standards
Listed on 2026-08-03
-
Research/Development
Data Scientist, AI Evaluation -
IT/Tech
AI Engineer (Applied/Software), Data Scientist, AI Evaluation, Machine Learning/ ML Engineer
Post Doc – Data Analytics & Computational Sciences
At J&J we are developing Generative AI solutions to support pharmaceutical R&D, including literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support.
These systems need to be tested before teams use them in scientific workflows. In pharmaceutical R&D, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error.
We are looking for a postdoctoral researcher to help design methods that test whether GenAI tools produce answers that are accurate, evidence-grounded, traceable, usable, and appropriate for the intended task.
The role reports to the Associate Director, Generative AI Evaluation & Quality Standards. The team defines how J&J Innovative Medicine evaluates GenAI systems before use and helps determine when they are ready for release, expansion, or improvement.
Key Responsibilities- Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
- Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
- Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
- Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
- Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
- Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
- Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
- Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
- Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
- Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
- Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.
- PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
- Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
- Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
- Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
- Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
- Proficiency in Python and common data science or machine learning tools.
- Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
- Clear written and verbal communication skills.
- Experience with LLM APIs, embeddings, vector databases, prompt engineering, agent frameworks, or AI evaluation tools.
- Experience evaluating retrieval quality, generated answers, multi-step workflows, tool use, scientific reasoning, citation quality, or evidence-grounded outputs.
- Experience designing expert review workflows, annotation instructions, adjudication processes, or inter-rater reliability analyses.
- Domain knowledge in one or more biomedical or therapeutic areas.
- Familiarity with biomedical data standards, structured scientific or clinical data, ontologies, knowledge graphs, CDISC, FHIR, or related frameworks.
- Publications or applied research in AI evaluation, NLP, biomedical informatics, machine learning, data science, computational biology, bioinformatics, or a related field.
Required…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).