×
Register Here to Apply for Jobs or Post Jobs. X

AI Engineer - Harness & Evals

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Build Technologies
Full Time position
Listed on 2026-06-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below

About Build

Build is creating the agentic AI stack for the built world. We help institutional real estate teams automate complex development and acquisitions workflows so important projects can move from concept to completion faster, with less cost, delay, and operational drag.

Our customers include some of the largest built-world institutions: alternative asset investors, developers, infrastructure owners, energy companies, industrial operators, and public‑sector partners. Their work shapes the physical world, but the workflows behind that work are still slow, fragmented, document‑heavy, and dependent on expert coordination.

We believe the next generation of built‑world software will not just organize work. It will help do the work. Agents will reason across documents, drawings, financial models, market data, approvals, constraints, and expert judgment. Human experts will stay in control, but they will operate with far more leverage.

We are backed by leading investors and operators, including executives from Blackstone and OpenAI, alongside top venture firms. We are building a generational company at the intersection of AI and the physical world.

About the role

We are looking for an AI engineer, core to build the infrastructure, systems, and quality loops behind Build’s agentic platform.

This is a hands‑on engineering role for someone who wants to make agents reliable, observable, scalable, and safe enough for high‑stakes real‑world workflows. You will work close to the agent runtime, evaluation systems, retrieval layer, tool orchestration, tracing, workflow execution, and developer platform that power Build’s product.

You should be excited by the engineering problems that appear after the demo works: how agents plan and call tools, how context is assembled, how workflows resume after failure, how quality is measured, how regressions are caught, how cost and latency are controlled, and how engineers can ship agent improvements with confidence.

This is not a research‑only role. It is infrastructure work for production AI systems. Your work will define the foundation that lets Build ship faster while keeping quality, trust, and reliability high.

What you’ll do
  • Build the core agent platform used by product engineers to create, run, evaluate, debug, and deploy AI workflows.

  • Design infrastructure for long‑running agents, tool orchestration, workflow state, retries, fallbacks, human handoff, and resumability.

  • Build context and retrieval systems that help agents use the right documents, structured data, prior decisions, project state, and tool outputs.

  • Create eval infrastructure for agent behavior, document understanding, groundedness, workflow completion, visual reasoning, latency, cost, and regressions.

  • Build observability systems for traces, prompts, model versions, tool calls, intermediate reasoning artifacts, failure modes, human overrides, and production quality metrics.

  • Improve the reliability of LLM‑powered systems through deterministic checks, structured outputs, validation layers, guardrails, monitoring, and failure recovery.

  • Partner with product engineers to turn repeated workflow patterns into reusable primitives, SDKs, templates, and platform capabilities.

  • Evaluate and integrate models, agent frameworks, retrieval techniques, multimodal capabilities, and AI infrastructure tools.

  • Own performance, scalability, security, and maintainability across the AI platform.

  • Help define the engineering standards for production agent systems at Build.

Example projects
  • Build an agent runtime that supports durable execution, resumable workflows, retries, tool permissions, human approval gates, and production traceability.

  • Build an eval platform where engineers can run offline tests, replay production traces, compare model and prompt changes, detect regressions, and review failure clusters.

  • Build a context assembly layer that combines project documents, extracted entities, workflow state, customer configuration, user intent, and tool outputs into reliable agent inputs.

  • Build a retrieval quality system for leases, zoning documents, drawings, investment memos, financial models, and market data, with ranking, citations, and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary