Infrastructure Engineer
Listed on 2026-07-25
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
About Elicit
Elicitis an AI research assistant that useslanguage models to help researchers figure out what’s true and make better decisions, starting with common research tasks like literature review.
What we're aiming for:
Elicit radically increases the amount of good reasoning in the world.
For experts, Elicit pushes the frontier forward.
For non-experts, Elicit makes good reasoning more affordable. People who don't have the tools, expertise, time, or mental energy to make well-reasoned decisions on their own can do so with Elicit.
Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes. This expands our collective understanding of safe AGI architectures.
Visit our Twitter to learn more about how Elicit is helping researchers and making progress on our mission.
Why we're hiring for this roleElicit is an AI research platform used by scientists, pharma companies, and decision-makers for high-stakes evidence synthesis. A single session could trigger hundreds of thousands of language model invocations across multiple providers, which means our infrastructure decisions directly impact cost, reliability, and the quality of research outcomes for users making decisions worth millions of dollars.
Our infra is well-architected using best practices — Terraform, Kubernetes, Argo CD, Git Hub Actions. But we're at an inflection point: enterprise contracts are getting larger, single-tenant deployments are multiplying, and the surface area that needs dedicated attention has outgrown what our current team can cover part-time. This is the first dedicated infrastructure hire and you'll define how this function works at Elicit.
James (Head of Engineering, ex-Square) set up the original infrastructure and will be your close partner. This role will own and evolve the infrastructure platform that underpins Elicit's product. You will ensuring it is reliable, secure, cost-efficient, and ready for the demands of a growing enterprise customer base. Under your ownership, our systems will scale gracefully across single-tenant deployments, our SLAs will be backed by real engineering rigor rather than best intentions, and our compliance posture will be a selling point rather than an afterthought.
Whatyou'll own
Own our cloud infrastructure across AWS and GCP — Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
Scale single-tenant deployments from a handful to many — each with distinct data retention, geographic, monitoring, and compliance requirements. Make a private cloud deployment a repeatable, low-overhead operation.
Build our observability and incident response practice — proactive monitoring, alerting, SLA tracking, and structured post-mortems that make the whole team better at diagnosing and resolving issues.
Drive compliance and security operations — ensure we follow through on the policies we've written (SOC 2, NIST AI framework, EU Cyber Resilience). Own disaster recovery exercises, database restoration drills, and security event monitoring (SIEM).
Manage infrastructure cost and capacity — make smart decisions about where we run workloads (AWS, Core Weave, Parasail), optimize spend, and plan capacity as usage grows.
Improve developer experience — CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
Contribute to backend systems where infrastructure and application intersect — circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.
A private cloud deployment is a ~1-day turnkey operation. Playbooks and templated Terraform make standing up Elicit in a customer's cloud routine, which opens up 8-figure enterprise deals.
Our observability signal:noise ratio improves 10-fold. Health monitors cover every endpoint and job, and an alert firing means something needs attention.
Disaster recovery is practiced. We run database restoration drills and provider-outage dry runs on a schedule, with post-mortems that make the whole team better at diagnosis.
Ou…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).