×
Register Here to Apply for Jobs or Post Jobs. X

Senior Data Scientist, AI Training Data

Job in Redwood City, San Mateo County, California, 94061, USA
Listing for: Cognichip
Full Time, Apprenticeship/Internship position
Listed on 2026-09-14
Job specializations:
  • IT/Tech
    Machine Learning/ ML Engineer, Data Scientist, AI Engineer (Applied/Software), Data Engineering
Salary/Wage Range or Industry Benchmark: 150000 - 230000 USD Yearly USD 150000.00 230000.00 YEAR
Job Description & How to Apply Below

About Cognichip

We build AI-native tools for semiconductor engineering, combining large proprietary models, agentic workflows, and domain-specific intelligence to help engineers design, verify, and optimize chips faster.

Job Title

Senior Data Scientist, AI Training Data

About

The Role

We're looking for a Senior Data Scientist to own the data our models learn from. Semiconductor engineering data is specialized, scarce, and often tightly licensed — very different from the general text and code used to train most large models. Turning it into training-ready, evaluation-ready datasets is one of the highest-leverage inputs to our model quality. In this role, you'll design the curation, synthetic data generation, and quality-modeling work that turns raw technical material into usable datasets, running at scale on our internal data infrastructure (managed by a dedicated platform team, so you can focus on the data itself).

You'll work closely with domain engineers to figure out what the models actually need, and with our AI team to connect dataset improvements to measurable model performance gains.

Key Responsibilities
  • Curate licensed and open-source technical datasets — collection, cleaning, annotation, and integration across the engineering lifecycle
  • Build automated pipelines for sourcing, license classification, and normalization of public data
  • Design synthetic and augmented data generation workflows to keep pace with model training demand
  • Develop quality-modeling approaches: deduplication, contamination/leakage detection, license and PII screening, difficulty/diversity scoring, and dataset-to-eval attribution
  • Write large-scale distributed data processing jobs, partnering with a platform team on infrastructure needs (throughput, versioning, lineage, reproducibility)
  • Translate observed model weaknesses and feedback into targeted, well-sourced datasets
  • Build and maintain retrieval/embedding datasets that support product features
  • Run exploratory analysis and produce insights that guide modeling, product, and go-to-market decisions
  • Collaborate across engineering, AI research, product, and business teams to turn ambiguous needs into concrete datasets
  • Establish data governance practices: license provenance, documentation, retention, and compliance
Required Qualifications
  • MS or PhD in Computer Science, Data Science, Statistics, or related field
  • 5–10 years of hands-on experience in data science or ML data work, with ownership of production datasets used by other teams
  • Expert Python and strong SQL, with experience processing large datasets using distributed computing frameworks (e.g., Spark)
  • Practical experience preparing text or code corpora for LLM training, fine-tuning, or evaluation
  • Solid applied statistics and ML foundations, with experience in a major ML framework (PyTorch, Tensor Flow, or scikit-learn)
  • Familiarity with modern data orchestration and versioned storage systems
  • Working knowledge of data governance practices — licensing, provenance tracking, handling of confidential/contractual data
  • Strong ability to work with domain experts and convert ambiguous requests into delivered datasets
Preferred Qualifications

The following items are not required but are great bonuses:

  • Exposure to hardware or engineering domain data (e.g., specialized design/verification formats and workflows). You don't need deep prior expertise — just genuine interest in learning a technical domain deeply
  • Experience building retrieval systems: chunking strategies for technical documents, embedding models, vector databases, and evaluation of retrieval-augmented systems
  • Experience with synthetic data generation using LLMs, including agentic pipelines built with modern orchestration frameworks
  • Familiarity with annotation tooling and workflows…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary