×
Register Here to Apply for Jobs or Post Jobs. X

Principal Data Scientist

Job in Washington, District of Columbia, 20022, USA
Listing for: Wiley
Full Time position
Listed on 2026-09-30
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Data Scientist
Salary/Wage Range or Industry Benchmark: 140000 - 201000 USD Yearly USD 140000.00 201000.00 YEAR
Job Description & How to Apply Below

Job Description:

We believe in bold ideas, diverse perspectives, and the drive to transform knowledge into impact. Here, your curiosity fuels progress, your voice shapes innovation, and your ambition helps redefine what's possible within science and learning. We are a culture that obsesses over impact, challenges, and drives what's next to power infinite possibilities for our customers, colleagues and society at large.

About the Role:

We'rebuilding the systems that turn one of the world's largest scientific corporation intoresearchintelligence. That meansproductionNLP pipelines running over millions of journal articles, extracting entities, classifications, claim tuples, and summariesoptimizedfor use by downstream agentic applications.

We'relooking for a principal data scientist to owndomain-specificcontentmodeling work end to end, from the eval set through the pipeline stage that ships it.

You'lljoin a small, senior team where data scientists own their models in production.

You'llwrite the code, own the evaluations, ship the changes, and stay accountable for the outcomes. This is a hands-on role for someone who wants to see their models through to real usersin a rapidly evolving market.

Job Responsibilities:

  • Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientificfull-textat scale.

  • Compare NLP approaches to extraction and enrichment against LLM-based approaches, andpick the right tool for each task. That means putting traditional NLP (NER, sequence labeling, classification), embedding-based retrieval, LLM prompting, and fine-tuned smaller models on the same table, and defending each choice with evaluation, cost, and operational tradeoffs. This is a core part of the job, not an occasional exercise.

  • Own evaluation. Build the golden setsin consultation with SMEs and vendors, choose the metrics, and make productive tradeoffs between speed, quality, and cost.

  • Write production-quality Python. Manage concurrency and cost for high-volume LLM workloads. Structure code that engineers canshipand other data scientists can extend.

  • Collaborate with a team of data engineers to orchestrate work in datapipelineand data build tools like Airflow and Dagster. Design idempotent,retryable, evaluable pipeline stages that stay reliable when a run fails at scale.

  • Contribute to agentic AI application work: tool-using systems that reason over the enriched corpus, where your NLP and evaluation background will shape how the agent grounds and defends its answers.

  • Work directly with editors, product managers, and engineers. Bring the modeling perspective into product decisions, and translate stakeholderpushbackinto concrete modeling work.

Job Requirements:

  • Deep Python.

    You'vewritten it in production, at scale, for years. You know when to reach forasyncioversus threads versus a queue, and you can explain the tradeoff clearly.

  • Strong NLP background across modern (LLMs, transformers, embeddings, retrieval) and classical (NER, classification, sequence labeling) approaches.

    You'vebuilt evaluation sand learned from the results.

  • A habit of comparing approaches and choosing the right one for the task. You can defend "prompt a large LLM" and "train a small classifier on 2,000 labels" with equal seriousness, back the choice with an eval and a cost estimate, andknow what to do when performance drifts.

  • A track recordof shipping - not just prototypes and papers, but systems that deliver value to real users.

Preferred:

  • Experience working with scientific or scholarly text.

  • Familiarity with AWS (S3, Batch, Lambda, Sage Maker) and Parquet or Iceberg data lake patterns.

  • Experience running LLMs underreal costand latency budgets in production.

  • Some exposure to agentic AI applications: tool use, multi-step reasoning, guardrails, and evaluation of trajectories rather than single-turn outputs.

We power infinite possibilities.

For more than 200 years, we've transformed knowledge into discoveries that shape the world. Today, our global team of innovators, creators, and experts is driving what's next in science, education, and publishingâcreating impact that reaches everywhere.

We're not just observers of progress. We're the ones accelerating scientific breakthroughs, advancing learning, and sparking innovation that redefines entire fields and improves lives.

Here, your talent matters. Your ideas have room to grow. And your work creates breakthroughs that can change everything.

Wiley is an equal opportunity/affirmative action employer. We…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary