×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - Data Flywheel Infra, Frontier Models

Job in Redmond, King County, Washington, 98073, USA
Listing for: Microsoft Corporation
Full Time position
Listed on 2026-08-30
Job specializations:
  • IT/Tech
    Data Engineering, Data Scientist
Job Description & How to Apply Below
** Overview*
* We are looking for a  
** Data Flywheel Infrastructure Engineer
** to build the infrastructure that continuously turns  
** 1P data, 3P data, model signals, evaluation results, and synthetic data
** into high-quality training data for frontier LLM and multimodal models.

This role owns the systems connecting:

** Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement*
* A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is  
** secure, policy-compliant, rights-aware, traceable, and auditable** .

Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.

This role is part of Microsoft AI's Superintelligence Team. The MAIST is a  
** startup-like team inside Microsoft AI** , created to push the boundaries of AI toward  
** Humanist Superintelligence-ultra-capable systems that remain controllable, safety-aligned, and anchored to human values.
** Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society-advancing science, education, and global well-being.

We're also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you're a brilliant, highly-ambitious and low ego individual, you'll fit right in-come and join us as we work on our next generation of models!

** Responsibilities*
* 1.  
** Build 1P & 3P Data Flywheel Infrastructure
** Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows.

2.  
** Own Data Governance, Security & Compliance Infrastructure
** Build governance and policy enforcement directly into the data platform, including:

+ Data provenance and lineage

+ Usage rights, licensing, and consent metadata

+ PII / sensitive-data detection and protection

+ Access control and data isolation

+ Retention and deletion enforcement

+ Geographic and regulatory restrictions

+ Dataset approval and audit workflows

+ Training eligibility and purpose-based usage controls

1.  
** Build Policy-Aware Data Acquisition & Curation Systems
** Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.

Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use.

2.  
** Build Evaluation-to-Data Feedback Loops
** Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation.

Enable rapid iteration from:
** Model Failure → Data Gap → Data Intervention → Training → Evaluation*
* 3.  
** Build Synthetic & AI-Native Data Pipelines
** Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.

Maintain clear provenance between  
** human-created, first-party, third-party, model-generated, and derived data** , and enforce appropriate policies across each category.

4.  
** Build Data Quality, Attribution & Observability
** Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.

Enable researchers to understand  
** which data improves which capabilities and under what governance constraints** .

** Qualifications*
* Required

-  Master's Degree in Computer Science, Math, Software Engineering, Computer…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary