×
Register Here to Apply for Jobs or Post Jobs. X

Human Data Architect, Quality

Job in New York, New York County, New York, 10261, USA
Listing for: Kindredventures
Full Time position
Listed on 2026-07-06
Job specializations:
  • IT/Tech
    AI Evaluation, Data Annotation/ AI Labeling, AI Business & Operations, Data Engineering
Salary/Wage Range or Industry Benchmark: 100000 - 130000 USD Yearly USD 100000.00 130000.00 YEAR
Job Description & How to Apply Below
Location: New York

About Mecka AI

Mecka AI is building the data and deployment infrastructure for embodied intelligence. We collect, curate, and license the world's most useful robotics training data to leading AI labs, and we deploy real robotic systems with enterprise customers across hospitality, retail, QSR, pharmacy, logistics, and healthcare. We work with the foundation model teams shaping the next decade of robotics, and with the operators running real businesses today.

Quality, trust, and execution are core to our partnerships.

The Role

We're hiring a Human Data Architect, Quality to be the person with taste for what robotics training data should look like  will define what good data is — the labeling rubrics, ontologies, schemas, sampling philosophy, and acceptance criteria that every dataset we ship is measured against. You decide what goes in or out of a dataset and why.

This is a standards-and-methodology architecture role, not a QA-management role
. You set the quality bar; data operations and QA teams enforce it. Your output is the spec the entire data org and our customers run on.

You will work shoulder-to-shoulder with foundation-model researchers at our customers to translate model behavior into data structure — what to label, how to label it, how to organize it, how to compose a training set, what the edge cases are, and what makes a dataset trainable versus merely large.

What You'll Own Labeling Rubrics & Quality Criteria (per customer)
  • Define the labeling rubrics, severity levels, rejection taxonomies, and acceptance criteria for each customer program across video, sensor streams, trajectories, action labels, task outcomes, language grounding, and metadata.
  • Translate ambiguous customer requirements ("we want a model that can do X") into precise, measurable, executable data specifications.
  • Maintain customer-specific quality criteria and the canonical data dictionary every program references.
  • Build golden datasets, reference examples, and calibration tasks that define "correct" by demonstration, not just description.
Ontology & Data Organization
  • Own the taxonomy, schema, and class hierarchies for robotics datasets — how attributes are structured, how temporal segmentation works, how event boundaries are defined, how ambiguity is handled, how edge cases are categorized.
  • Decide how data is organized end-to-end so it is trainable, queryable, and composable across customers and modalities.
  • Set dataset versioning conventions, schema evolution rules, and the data-organization philosophy the org runs on.
Dataset Composition — What's In, What's Out
  • Own the philosophy for what goes into a dataset and what gets cut: distribution, diversity, edge‑case representation, redundancy, license/provenance constraints.
  • Decide sampling strategies, balancing rules, and curation principles for each program.
  • Make taste‑driven calls on what data is worth collecting at all — and push back when collection plans won't produce trainable data.
  • Define the acceptance bar that says "this dataset is ready to ship" — and hold it under deadline pressure.
Methodology Iteration from Model Signal
  • Iterate rubrics and ontology based on model‑failure signal from customers — your standards evolve with what models actually struggle to learn.
  • Run cross‑customer reviews of recurring quality misses and translate them into standards improvements.
  • Partner with engineering on automated validation (schema completeness, duplicates, time sync, metadata coverage, model‑assisted review) so the standard is enforceable at scale.
Who You Are Required Background
  • 5+ years working at the intersection of ML and data — annotation methodology, dataset curation, data‑centric ML, ground truth design, or labeling‑specifications work for autonomy, vision, or multimodal teams.
  • Hands‑on experience designing taxonomies, ontologies, or labeling schemas that fed production model training (not just internal analytics).
  • Strong data instincts: you can open a dataset in SQL, a notebook, or Python and tell us what's wrong with it within an hour.
  • Comfortable reading ML papers and translating model‑architecture needs into data‑structure choices.
Strong Signals
  • Built a labeling rubric, ontology, or…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary