×
Register Here to Apply for Jobs or Post Jobs. X

Director, AI Experimentation & Measurement

Job in North Bothell Area, Snohomish County, Washington, 98021, USA
Listing for: Pfizer
Full Time position
Listed on 2026-07-27
Job specializations:
  • Business
    AI Evaluation, Data Scientist
Salary/Wage Range or Industry Benchmark: 162900 - 271500 USD Yearly USD 162900.00 271500.00 YEAR
Job Description & How to Apply Below

Pfizer US Commercial is scaling a portfolio of AI initiatives, each launched as an experiment measured against a real-world baseline. The value of that portfolio depends on one thing: a shared, credible definition of what good looks like plus the evidence to prove we are reaching it.

We are hiring a Director of Experimentation & Measurement to own that definition end to end. This is not a scorekeeping role. You will decide what is worth proving next, design the experiment that proves it, set the bar for success before the work begins, and tell the story of what the evidence means to the people making the biggest decisions.

You are equal parts systems thinker, experiment designer, and storyteller, with the rigor to make sure the story is true.

This is a build-from-scratch role for someone who can see across a whole portfolio, frame the questions that matter most, and hold the independence to say when a result does not clear the bar.

What You Will Own
  • Success criteria & proof standards. Own the portfolio-wide standard for what counts as proof—defined before work begins, not after results land. Draft and socialize pre-registered success criteria for each initiative (primary KPIs, guardrails, minimum detectable effect, read windows). Maintain a shared measurement playbook (matched controls, power calculations, blinded review, backtesting where appropriate).
  • Portfolio learning agenda. Prioritize what the portfolio must prove next—and in what sequence evidence compounds. Maintain a ranked learning agenda across initiatives and recommend the next proof point per initiative.
  • Experiment design & statistical rigor. Review experiment designs: control arms, audience matching, read windows, confounders, data leakage risks. Run power analysis and flag under powered or ungradeable designs before launch.
  • Evidence narrative for senior leaders. Produce evidence read memos on a fixed cadence: what was tested, what was observed, what it means, what is recommended. Synthesize across initiatives into one portfolio story—not a patchwork of disconnected scorecards.
  • Measurement integrity & bias controls. Guard against failure modes that quietly invalidate results: moving goalposts; letting the people a system is meant to outperform grade it; acting before pre-registered proof exists. Escalate when a read is not gradeable.
  • Evidence governance & readout cadence. Run evidence read sessions at pre-registered dates; own continue / adjust / stop recommendations with rationale; track whether recommendations are acted on.
Key stakeholders & accountability

Partners with initiative facilitators on criteria, design, and readouts. Influences BU Presidents and proxies by helping define their success criteria and reporting back whether evidence meets the bar. Reports results and implications to USLT and Chief Commercial Office as portfolio executive sponsors.

What we require

Experience:

8+ yrs (Bachelor's)

  • 7+ yrs (Master's)
  • 5+ yrs (PhD), spanning experimentation, measurement science, causal inference, or a quantitatively rigorous discipline — paired with strategy or product-facing work.
Required Qualifications
  • Bachelor's degree required with 8+ years of relevant experience spanning experimentation, measurement science, causal inference, or a quantitatively rigorous discipline — paired with strategy or product-facing work.
  • Systems thinking: sees across an entire portfolio, frames the questions that matter most, and designs how proof compounds over time rather than optimizing one test at a time.
  • Storytelling and executive presence: turns a rigorous, technical result into a clear, persuasive narrative that shapes senior decisions.
  • Experiment design and causal credibility: has designed and defended controlled experiments at enterprise scale: power analysis, matched controls, quasi-experimental methods, pre-registration.
  • Independence and spine: able to tell a senior stakeholder that a result does not clear the bar, and make it stick.
  • AI fluency: evaluates AI/ML and generative systems credibly: what a real performance signal looks like versus a plausible artifact.
Preferred Qualifications
  • Regulated-industry measurement experience.
  • Provenance, auditability,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary