×
Register Here to Apply for Jobs or Post Jobs. X

Evaluations Engineering - Member of Technical Staff

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Simile
Full Time position
Listed on 2026-09-08
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 200000 - 400000 USD Yearly USD 200000.00 400000.00 YEAR
Job Description & How to Apply Below

About the Company

Pilots don't train with real passengers. Actors don't rehearse with real audiences. Yet, the most consequential decisions in society are often pushed straight to production.

Simile is changing that. We have built the first AI simulation of society, populated by generative agents based on real humans. Our research pioneered the field of AI-based simulation, proving it is possible to model human behavior with high accuracy. Today, we are developing a Foundation Model to predict human behavior in any situation, at any scale.

We are backed by $100M in funding led by Index Ventures, with participation from Hanabi, A*, Bain Capital Ventures, and AI visionaries including Andrej Karpathy, Fei-Fei Li, Adam D'Angelo, and Guillermo Rauch.

About the Role

As a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.

You will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling. Your initial focus will include streamlining how evaluations are run across models; strengthening evaluation versioning, data models, and access controls; and automating customer validations, survey operations, and human data workflows.

Evaluation at Simile presents unusual engineering challenges. Our models predict distributions of human behavior, and the ground truth used to evaluate them can be noisy and heterogeneous. You will partner closely with Evals, Modeling, Product Engineering, and Data Operations to turn complex methods and inputs into systems that are reproducible, scalable, and useful for model development and business decisions.

In this role, you will:
  • Build evaluation execution infrastructure:
    Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.

  • Strengthen evaluation data systems:
    Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.

  • Automate validation and data collection:
    Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.

  • Build human data workflows:
    Create labeling and review tools that enable external experts and operators to contribute high-quality judgments to evaluation campaigns.

  • Develop evaluation tooling:
    Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.

Requirements Must Haves
  • Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.

  • Data and Systems

    Experience:

    Experience building backend services, data pipelines, automation workflows, and relational data models.

  • End-to-End Execution: Ability to work across data, backend, and interface layers and take ambiguous projects from technical design through deployment and adoption.

  • Evaluation Judgment: Strong intuition for what makes evaluation infrastructure reliable, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and meaningful model comparisons.

  • ML and LLM Fluency: Familiarity with modern model-development and evaluation workflows sufficient to partner effectively with modeling and evaluation researchers.

  • Product and User Judgment: Ability to build clear, efficient tools for researchers, engineers, data operators, and other expert users.

  • Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.

Nice to Haves

We do not expect one person to have all of these. We are hiring a team with complementary strengths.

  • Model-Evaluation Infrastructure: Experience building LLM or ML evaluation systems, benchmark platforms, regression suites, experiment-tracking tools, or model-quality dashboards.

  • Research and Internal Tools: E…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary