Agent Harness Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-09-09
Listing for:
Nickel
Full Time
position Listed on 2026-09-09
Job specializations:
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Company Description
Nickel is building AI infrastructure to enable doctors to deliver personalized, diagnostics-led longevity medicine to people everywhere. Backed by an eight-figure funding round, Nickel is preparing for launch and expanding its team with exceptional software and AI engineering talent. The company focuses on turning complex medical diagnostics into actionable insights that support better long-term health outcomes. Team members join a fast-paced, mission-driven environment with the opportunity to shape core infrastructure at the intersection of healthcare and artificial intelligence.
Role DescriptionBuild Nickel’s proprietary Agent Harness technical stack through the synergy of algorithms, engineering, and modeling:
- Models handle reasoning and generation.
- Harness manages the entire runtime beyond the model: context, memory, tools, control loops, multi-agent scheduling, and self-evolution.
- Establish and close loop with the Model Training effort to drive the co-evolution of models and the Harness.
- Research and Implementation of Cutting-edge Harness Mechanisms
- Context Management & Context Engineering:
Truncation, compression, KV Cache/prefix cache reuse, and cross-turn context strategies. - Long-term Memory:
Episodic/semantic memory, vector stores, session persistence, and replay. - Agent Loop & Tooling:
Agent Loop, Tool Use, Skills, MCP protocol integration, and tool orchestration. - Multi-Agent Systems:
Subagent communication, task decomposition, result aggregation, and hierarchical planning. - Self-Evolving Agents:
Execution and recovery of ultra-long-horizon tasks (multi-step, multi-day, multi-session). - Model–Harness Synergy
- Define model requirements from a Harness perspective (e.g., reasoning formats, tool-call stability, planning signals).
- Surface failure modes from real-world tasks (e.g., planning collapse, tool misuse, context loss) to the training effort and work out solution.
- Evaluation and Research Loop
- Design Harness-level benchmarks, evaluation metrics, and data annotation strategies.
- Iterate through experiments using real-world tasks (internal and external scenarios) rather than relying solely on model benchmark scores.
- System Architecture and Engineering Implementation
- Participate in technical selection and architecture design for Harness products/systems (Agent runtime, scheduling layer, MCP client/server, sandboxing, state machines).
- Translate research prototypes into runnable, observable, and iterable production systems.
- Solve production-grade challenges such as interrupt recovery, streaming, long-running tasks, error fallback, and concurrent sub-agent management.
- Background in Computer Science or a related field. Master’s degree or above for the algorithm track (exceptional candidates may be exempted); bachelor’s degree or above from a reputable university for the engineering track.
- 2 years of research experience for the algorithm track (exceptional candidates may be exempted). For the engineering track: strong technical skills, broad vision, and the ability to independently drive modules.
- Deep user of code-focused and general-purpose Agent products (e.g., Claude Code, Cursor, Codex, Copilot, Manus, Open Claw), integrating Agents into daily workflows.
- Proficiency in LLM and Agent fundamentals: LLM APIs, KV Cache, Agent Loop, Tool Use, Reasoning, Planning, Skills, MCP, Memory, Subagent, and Multi-Agent systems.
- Understanding & hands-on capability of model training workflows (pretraining / SFT / RLHF / DPO, etc.), with the ability to contribute to training data construction, preference labeling, and evaluation loop design.
- Capable of translating Harness-layer observations into interpretable, quantifiable improvement proposals for model training.
- Practical, hands-on understanding of Prompt Engineering, Context Engineering, and Harness Engineering (deep understanding for the research track; solid understanding for the engineering track).
- Strong intuition and judgment regarding model behavior; ability to distinguish between “model defects” and “Harness failure to handle edge cases.”
- Clear communication skills in both Chinese and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×