×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff, Core AI

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Sycamore
Full Time position
Listed on 2026-09-25
Job specializations:
  • Software Development
    Backend Developer, AI Engineer (Applied/Software), DevOps, Software Engineer
Salary/Wage Range or Industry Benchmark: 170000 - 260000 USD Yearly USD 170000.00 260000.00 YEAR
Job Description & How to Apply Below

Build the runtime, evaluation, and learning systems that every agent on the platform depends on.

About Sycamore

Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate agentic apps that take on real operational work, with the security and control large organizations need.

We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.

Where you could focus

Core AI owns the horizontal runtime, intelligence, and improvement capabilities that Product and Infrastructure both depend on, along with the Sycamore Forge experiences that make them usable. Engineers here own end-to-end slices, from cloud service and API design through the React interface. The team covers four related areas, and you do not have to pick one to apply. Tell us what you have built and we will work out the fit together.

The

agent runtime

Multi-turn sessions, model routing, tool execution, memory, durable workflows, and the APIs that expose them. The hard part is correctness across long horizons: surviving retries, provider interruptions, partial failure, and context growth without losing the thread.

Harnesses and environments

An agent that writes and runs code needs somewhere to do it. That environment has to start fast, isolate genuinely untrusted execution, reach only the services it legitimately needs, and carry credentials it can use but never read. The same area owns the verification layer that decides whether what an agent produced actually works rather than only appears to.

Evaluation

Agent quality is genuinely hard to measure. A change that looks better on a handful of examples often is not, and a judge model can be confidently wrong in the same direction as the system it grades. This is where the offline suites, replay corpora, and statistical discipline that gate a release come from.

Self-improving systems

Every agent run produces evidence, and almost all of it is currently thrown away. This area turns it into improvement: structured trajectories at fleet scale, failure clusters surfaced from production rather than guessed at, and proposed changes that are versioned, measured, staged, and reversible.

What you will do
  • Develop the learning data plane around agents: structured trajectories, feedback and outcome signals, offline datasets, lineage, privacy controls, and reliable links between an agent version and its behavior.
  • Create evaluation systems for task completion, tool use, long-horizon behavior, safety, latency, and cost, and own the credibility of the numbers they produce.
  • Design experiment and versioning systems for comparing changes through replay, shadow traffic, canaries, or controlled rollouts, with clear promotion and rollback criteria.
  • Build durable orchestration for long-running tasks, checkpoints, approvals, handoffs, and human-in-the-loop interactions, with typed tool interfaces, protocol-based execution, and memory retrieval that enforces tenant, user, and project visibility boundaries.
  • Own the execution environments agents run in, including isolation, startup performance, resource limits, and cost, and the policy layer that gates self-modification so a proposed change is staged and tested rather than applied silently.
  • Publish reliable APIs, event-driven interfaces, and reusable libraries that work across agent categories and enterprise deployments.
The environment you will work in

Our current Core AI environment includes Python cloud services;
React and Type Script product surfaces in Sycamore Forge; asynchronous and streaming systems; typed APIs and data models; relational and vector data; durable workflows; protocol-based tool execution; multiple model providers; and cloud-native deployment.

We use coding agents, automated tests, traces, evaluations, browser automation, offline replay, cost and latency signals, and production feedback as part of everyday engineering. We are building toward a governed collect, learn, evaluate, and apply loop rather than a single monolithic training system.

This is context, not a checklist. We do not require previous experience with every language, framework, model provider, cloud platform, database, or infrastructure tool in our stack. Comparable experience building distributed runtimes, experimentation platforms, retrieval or recommendation systems, workflow engines, developer platforms, or production AI systems is…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary