Senior Software Engineer, Agentic Systems
Listed on 2026-09-04
-
Software Development
Backend Developer, AI Engineer (Applied/Software), Software Engineer, DevOps
Are you looking to build tools that engineers actually keep in their workflow? Do you get excited about AI agents that fix real vulnerabilities instead of writing confident nonsense about them? Are you willing to tolerate ridiculous bird puns? If you said yes to at least two of those, keep reading.
The RoleStack Hawk builds the security layer for AI-assisted engineering. Our platform finds exploitable vulnerabilities in running applications and APIs, then hands the proof and the fix path directly to the coding agent already sitting in the developer's editor. Claude Code, Cursor, Copilot, Codex. Scan, fix, verify, in a loop, without a human copying findings between tabs.
That product is real and shipping. It is also early, which means the architecture decisions that matter most have not been made yet.
We are looking for a senior engineer to own that loop. You will spend most of your time on agentic systems: the tool surfaces our agents call, the context they get, the loop that decides when a fix is actually verified, and the evals that tell us whether any of it got better this week. You will also work in the core platform, because the loop is only as good as the scan engine and APIs underneath it.
This is not a research role. Everything you build goes to customers.
The team is small. Your work will be visible in the product within days, and in customer conversations within weeks.
Want to see what we mean? Our agent skills are open source at Read them.
On DenverThis role is onsite in Denver. We know that narrows the pool and we are doing it on purpose. The hard part of this work is not writing the code. It is the whiteboard argument about why the loop stopped early, the shoulder tap when an eval result looks wrong, and the twenty minutes after a customer call that turns into a design change.
That happens in a room with our small team.
- Design and build the find, fix, and verify loop that powers our agentic security workflows
- Build and own the tool layer coding agents use to reach our platform. Some of this exists and needs a rewrite. You will have a strong opinion about how, and the room to act on it
- Extend our published agent skills, which are open source and in use today
- Do the context engineering work that makes the difference between a useful agent and a plausible one. Decide what the model sees, in what shape, and at what point in the loop
- Build eval harnesses for non-deterministic systems. Define what "better" means, measure it repeatably, and defend the number
- Build core product features and supporting services in a microservices architecture using Kotlin, Java, Rust, gRPC, Postgres, Docker, Kubernetes, and Gradle
- Design and build APIs for both our UI and our customers
- Build developer-facing command line tooling
- Integrate with the platforms engineers already live in:
Git Hub, Git Lab, CI systems, IDEs, and agent runtimes - Work directly with product, security, and go to market. We are small enough that you will hear customer feedback firsthand and ship against it
- Learn more about vulnerability classes than you expected to, and get very good at explaining Remote OS Command Injection at parties
We move fast and this list is not exhaustive.
About You Core experience- 6+ years building and shipping production SaaS software
- Deep in at least one statically typed language. Python, Type Script, Go, or Rust experience is where our agent tooling and CLI live, so that helps immediately
- Our core platform is Kotlin and Java. We do not require you to have written either. If you are strong somewhere else, you will be productive in our codebase inside a month and we are fine with that math
- Experience with microservices in a modern cloud environment, containers, and container orchestration
- Solid API design instincts and the tooling that goes with it
- Obsessive about automation and automated testing
- You have built something with an agent loop, not just called an LLM API. Tool calling, multi step execution, retries, failure handling, knowing when to stop
- You have written evals for non-deterministic output. You can talk about what you measured, why that metric was…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).