Applied AI Engineering Intern
Listed on 2026-08-30
-
Software Development
Backend Developer, AI Engineer (Applied/Software)
Applied AI Engineering Intern
Company:
Blockhouse
Location:
Remote (US time zones required); NYC in-person or hybrid preferred
Type:
Internship — full-time, hiring is evergreen
Compensation: $30-40/hour + meal / transportation stipend
About Blockhouse
Blockhouse is an institutional trading infrastructure platform for high and mid frequency trading strategies across centralized and decentralized perpetual venues - across all asset classes.
Role Overview
The internship is on the harness: everything around the model that makes an agent with real credentials safe to run on a real book. You will work directly with the founders and the engineers who own the systems the agent reads.
Where the work is:
Orchestration. A turn runs for minutes: planning, choosing among about thirty tools, long chains of tool calls. Scheduled runs fire on a cron and post to a channel. Sessions and workflow steps are persisted in Postgres, and a reply is posted once, complete.
Isolation. Each conversation runs in its own Kubernetes sandbox. Agents that share any state, such as a package cache, will find it and coordinate through it unprompted, so what sandboxes can see of each other is a design decision.
Credentials. For third-party services the agent holds a placeholder. A proxy at the network edge injects the real key for the one host it is bound to and refuses every host off the allowlist.
Context. Skills load on what people ask; tools carry the facts. The standing prompt has a byte budget enforced in CI, so adding an instruction means removing one. The right change is usually a new tool and a shorter prompt.
Evaluation. A live answer is graded only on evidence it carries: its own tool calls, a lookup against the source of truth, its internal consistency. A flagged answer becomes a ticket whose fix ships with a regression test. Only a reviewed pull request can change the agent.
Cost. Spend scales with context size times loop iterations. Every action the agent takes is logged outside the sandbox.
If you have built an agent harness of your own, even a small one, we would like to see it.
Responsibilities
- Add a data source end to end: the tool that reads it, the skill that teaches the agent to use it, and the eval that checks the answers
- Extend the orchestration layer: scheduled runs, durable multi-step workflows, delivery back into Slack
- Extend the eval layer: grounded checks, validating the LLM judge against human labels, per-workflow token metering
- Move a rule out of the standing prompt into a tool and show the prompt got shorter
- Work with a senior engineer on the isolation and credential boundary
Required Qualifications
- You have built an LLM agent with tool use and multi-step turns that ran for real users, even a handful. You can explain what it does when the process dies mid tool call, how it chooses a tool, and what a turn costs.
- Python and Rust. Everything we have built on the runtime is Python; the runtime itself is Rust, and you will change both.
- Linux, Docker, git, and Kubernetes:
Network Policy, Helm, pod lifecycle. The runtime runs on it and every data source is wired through it.
- Postgres as a durable state store: transactions, idempotent writes, checkpoint and replay.
- Claude Code and the Anthropic API: tool definitions, skills, prompt caching, and how a harness loop is driven from a CLI.
- You understand why an agent should hold a placeholder instead of a key, why a keyword match over live answers is not an eval, and why an LLM judge is unvalidated until a human grades a sample of its verdicts.
- You prefer small, verified changes to rewrites, and you write them so someone who was not there can follow the reasoning.
Preferred Qualifications
- Contributions to an open-source agent runtime or coding-agent harness
- You have built a Slack app:
Events API, request signing, delivery of long-running replies
- Security engineering at a network boundary: TLS-intercepting proxies, credential injection, egress allow lists
- Published research or serious independent work on agent evaluation or multi-agent coordination
- Trading systems:
PnL, ledgers, positions, funding, and why a figure is relayed from the system that owns it
Why Join Blockhouse
- Exposure to live trading book, and the desk reads what the agent writes every weekday
- Immigration support (H-1B, F-1)
- Full-time offers for interns who ship
- Evaluation for equity and PnL split
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).