×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer- Benchmarking

Job in New York, New York County, New York, 10261, USA
Listing for: Office-Hours
Full Time position
Listed on 2026-08-30
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 160000 - 210000 USD Yearly USD 160000.00 210000.00 YEAR
Job Description & How to Apply Below
Location: New York

Software Engineer, Benchmarking (SF, NYC, or Remote) About Us

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. Experts earn income by sharing their knowledge through advisory work, projects, and AI model training. Our platform handles the complexities behind the scenes— screening, compliance, scheduling, and payments—so knowledge sharing stays focused on meaningful insights and real impact.

We're a hyper-growth and profitable company, quickly expanding our expert network, launching new offices, and new products. We are headquartered in San Francisco, with offices in Brooklyn and Bangalore. Our customers include the fastest-growing digital health companies, technology companies, institutional investment firms, consulting firms and AI Labs. We are backed by top marketplace investors and operators of companies like Door Dash, Airbnb, affirm.

What

we believe

Human knowledge is the world's most valuable asset. And yet, despite being more interconnected than ever, most knowledge still remains stuck in our heads, inaccessible and underutilized. Our vision is to make human knowledge easily accessible and infinitely scalable by building tools for the new age knowledge economy.

About the role

We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.

Our researchers design the methodology. You'll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.

What you'll do
  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.

  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.

  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.

  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.

  • Develop the scoreboard and leader board: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.

  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.

  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.

  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.

What you bring
  • Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.

  • Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.

  • Comfort with containers and environments: Experience with Docker and building reproducible execution environments.

  • Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.

Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.

Tech Stack
  • Evaluations:
    Python, model APIs, agent/evaluation frameworks, custom evaluation tooling

  • Models: APIs from the…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary