×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Designer, AI Evaluation Platform

Job in New York, New York County, New York, 10261, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-09-18
Job specializations:
  • IT/Tech
    AI Evaluation
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

AI systems are only as trustworthy as the methods used to evaluate them. At Apple, where AI powers experiences for billions of people, getting evaluation right is not a support function, it is a foundational science. Our team, part of Apple Services Engineering, builds the platform that teams across Apple use to evaluate the AI and agentic systems they ship. It's where they define what  good  means, prove it, and act on what they find.

Evaluating non-deterministic systems is one of the hardest unsolved problems in production ML, and one Apple has to get right 're looking for a Staff Experience Designer to own that experience end to end. You'll be the first designer on this platform, and evaluation as a design practice doesn't have settled patterns yet, so your work will help define them.

Description

Teams building AI features need to know if what they shipped works. Today that means navigating unfamiliar territory: scoring something non-deterministic, telling a real regression from noise, trusting a judgment a model made instead of a person. Most existing tools here were built for the researchers who invented these methods, not for the people who now need to use them daily.

You'll set the direction here and build the design system everyone else builds on. We'd rather get a rough version in front of real users than a polished one later. And because teams use this platform to decide what to ship, trust matters more than usual: provenance, uncertainty, and edge cases determine whether someone acts on a result or quietly stops believing it.

What makes this team unusual is its interdisciplinary core. Alongside the platform, we run a research group working on evaluation methodology itself, including how to tell whether an evaluator is calibrated, biased, or measuring what it claims to. Their methods ship into this platform, and you decide how anyone first encounters them. You'll be close to that work while it's still forming, instead of picking it up once it's finished.

Minimum Qualifications
  • 8+ years of experience designing digital products or experiences, including end-to-end ownership of complex, data-dense products for technical users.
  • You’ve set direction in spaces with no existing precedent, and you get to clarity by running research yourself with the people who use what you build.
  • A portfolio you can share, including at least one case study that walks through your process from problem framing to shipped outcome.
  • Proven ability to design complex information: dashboards, comparison views, large result sets, and results that carry statistical uncertainty.
  • Strong interaction and visual design skills, with a high bar for craft in dense, information-heavy interfaces.
  • Enough fluency in AI/ML concepts (benchmarks, metrics, model-based judging, agentic systems) to work directly with a research scientist or engineer.
  • Fluency with AI tools in your own practice. You use Claude Code or equivalents to build working prototypes and extend what you can make on your own.
  • Excellent communication skills, with the ability to build buy-in across engineering and research without formal authority.
Preferred Qualifications
  • Experience as the first or only designer on a platform, and/or experience translating research output into shipped product.
  • Experience building and maintaining a design system, including its components, patterns, and naming conventions.
  • An AI-first instinct for interface design: shipped interfaces organized around stated intent, or surfaces meant to be operated by agents as well as people.
  • Experience designing coherent workflows across multiple surfaces (UI, CLI, SDK) that need to stay consistent with each other.
  • Familiarity with modern evaluation, observability, or visualization tooling (e.g. Lang Smith, Braintrust, D3, Vega-Lite).
  • Comfort with SQL and notebooks to explore your own data.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary