×
Register Here to Apply for Jobs or Post Jobs. X

Research Intern: Interpretability & Reliability; Summer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: CTGT
Seasonal/Temporary, Apprenticeship/Internship position
Listed on 2026-09-09
Job specializations:
  • Software Development
    Python, Backend Developer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 52000 - 68000 USD Yearly USD 52000.00 68000.00 YEAR
Job Description & How to Apply Below
Position: Research Intern: Interpretability & Reliability (Summer 2027)

The Role

Frontier models are now usually right and occasionally confidently wrong, and they cannot tell you which is which. A model that is 95% reliable is useless in the settings we serve. CTGT's research function exists to close that gap. Our founding research stems from feature learning in neural networks, and we use that machinery to extract and steer features at runtime, on open and closed-weight models, without training a new artifact for every behavior.

As a research intern, you will own one hard problem inside this program from end to end. You will take a real research question, like a better way to find what a model represents, intervene on it, or bound how wrong the system can be, design an approach, implement it against real models, and prove or disprove it with evidence that holds up.

You will sit directly with the engineers building the Policy Engine, present in our weekly research review, and be expected to form opinions, ask hard questions, and take problems further than they were handed to you.

We hold interns to the standard of a calculation that has to be right, not a demo that usually works; in practice this means limited ground truth, unverifiable intermediate steps, and failure modes that hide in the tails.

What You Will Do
  • Implement and stress‑test methods for feature extraction and runtime intervention, from control vectors to activation probes, and make them work repeatably across model families
  • Design evaluations that bound error rather than average it: calibration under imbalanced data, reasoning‑trace grading, behavior in verifiable and non‑verifiable task regimes
  • Read the relevant literature, decide what actually matters, reproduce it, and push past it
  • Work with engineering to turn a finding into a Policy Engine capability that ships into audited, high‑stakes environments
  • Present your progress every week and defend your reasoning
Who You Are
  • Pursuing a degree (Bachelor's through PhD) in computer science, mathematics, the sciences, or a similarly unforgiving quantitative field. We care how you think, not your titles.
  • Strong mathematical foundations: linear algebra, probability, optimization, information theory
  • Can read a paper, decide what matters, and implement it
  • Have written real code for real computational systems; fluency with PyTorch and the modern ML stack, or the track record that says you will have it in weeks
  • Drawn to interpretability, model internals, and making systems provably reliable rather than usually fine
  • Self‑directed, and able to make real progress without constant scaffolding
Our Stack
  • Languages:

    Python, Rust, and Node/Type Script, with React on the frontend
  • Data:
    PostgreSQL, vector, and graph databases
  • Infra:
    Docker, Kubernetes, Terraform, across several cloud providers and customer VPCs
  • ML:
    Self‑hosted models on multiple GPU providers and frontier APIs
Logistics
  • Full‑time, in person in San Francisco
  • 10 to 12 weeks between May/June and August/September 2027
  • We sponsor US visas
What We Offer

World‑Class Backing:
You will join a venture‑backed company with institutional investors including Google's Gradient Ventures, General Catalyst, and Y Combinator.

Real Impact:
You will work directly on the core systems that determine how models perform in the wild. Your work ships into real, high‑stakes environments where governance, auditability, and performance are non‑negotiable.

Autonomy & Trust:
We operate with a high degree of trust. You are expected to form strong technical opinions and execute on them.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary