Data Scientist – PU Framework and Feature Store
Listed on 2026-07-22
-
IT/Tech
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Data Scientist
CoAd helps businesses navigate the complexities of workforce management through a combination of technology, expertise, and human support. Serving more than 17,000 clients nationwide, CoAd delivers payroll, HR, benefits administration, compliance support, workforce technology, and PEO services that help organizations simplify operations, support their employees, and drive growth.
Built on the belief that workforce solutions should be integrated, intuitive, and people-centered, CoAd empowers employers to manage their workforce with greater confidence while adapting to the evolving needs of today’s workplace.
Position SummaryCoAd is building two interlocking analytic substrates: the productivity-unit (PU) framework that measures realized τ across operational functions, and the analytics feature store that serves production ML models across the company. Both Data Scientist roles work across both substrates. This role reports to the AI Experimentation Lead and works in close coordination with the data team, the Staff MLOps Engineer, and the Principal AI Architect.
This is a hands‑on engineering‑grade data science role. The Data Scientist writes production‑grade Python, owns model code from notebook to production, and uses AI‑assisted coding tools as a daily driver. The role is not deck‑only.
The two Data Scientists collectively own the PU measurement layer and the model portfolio served from the feature store. Both roles are accountable for PU baselines and causal estimators on the measurement side, and for feature definitions and production models on the feature store side. Workload is allocated by the AI Experimentation Lead based on backlog priority and candidate strengths; neither role is scoped to a single domain.
Active work streams span MLR pricing redistribution, propensity to renew, churn, contact volume forecasting, staffing demand, payroll exception rates, and PU baselines for in‑scope function.
- Model development:
Build production ML models that serve from the analytics feature store. This includes problem framing with stakeholders, feature engineering, model selection, training, validation, calibration, and packaging for production. Models are owned end‑to‑end; there is no separate "ML engineer" handoff. - Feature definitions in the feature store:
Author feature definitions, including source contracts, transformation logic, freshness requirements, and lineage metadata. The Data Scientist is accountable for the quality and correctness of every feature they introduce into the store. - Data team collaboration on feasibility:
Work directly with the data team to validate feasibility of proposed features before they are committed to the backlog: source system availability, refresh cadence, data quality, governance and tenant‑isolation constraints. The Data Scientist is expected to be in the data team's review channels and to push back on infeasible features early rather than late. - PU baselines and causal estimators:
Both Data Scientists own portions of the PU measurement substrate: defining the productivity unit for in‑scope functions, building the pipelines that compute c_r (c‑sub‑r) from operational data, and authoring the causal estimators that underwrite τ measurement — difference‑in‑differences, synthetic control, propensity matching, interrupted time‑series. The Data Scientist is the analyst behind several of these readouts and is expected to be the methodological author of record on at least one PU baseline. - Model monitoring and reconciliation:
Define and instrument the production monitoring for each model: drift, calibration, business‑metric reconciliation. Carry the on‑call rotation for model issues alongside MLOps. - Documentation and methodological transparency:
Every model the Data Scientist ships carries a model card: assumptions, training data window, identification strategy, known failure modes, and the reconciliation plan. The bar is reproducibility from underlying data.
- AI‑first coding
- Claude Code, Copilot, or successor tools are the default development surface. Feature definitions, model code, evaluation scripts, and monitoring instrumentation are…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).