AI Engineer, Internship - Summer 2026 - Applications Open Now
Listed on 2026-08-12
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Who Are We?
Postman is the world's leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration-enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded.
Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at or connect with Postman on X (Use the "Apply for this Job" box below).:
We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman.
We're seeking an AI Engineer Intern to work alongside our AI team on large-scale AI and Agentic systems from data pipeline to production deployment. This role is scoped for someone with foundational experience who wants to deepen it: you'll own discrete pieces of real systems under the mentorship of senior engineers, not shadow work or isolated coursework-style projects.
What You'll Do Benchmarks & EvaluationContribute to APIFlow-Bench , our open-source benchmark for real API-development work: design and review benchmark tasks and their mock API environments, extend the evaluation harness and task-generation pipeline in Python, and help maintain the public multi-model leader board with statistical confidence intervals.
Help build a new action-level AI safety benchmark: instead of grading what a model says, it scores what an agent actually does inside a simulated enterprise API environment. You’ll work on scenario design, threat modeling (prompt injection, data exfiltration, permission overreach), and auditable evaluation design.
Model Training & EfficiencyFine-tune open-weight models for tool calling and agentic tasks (SFT, distillation, and RL) using PyTorch and the open-source training ecosystem, on both managed training platforms and self-managed cloud GPUs. Design and run experiments with rigor: evaluate every training run on our benchmarks, support ablation studies and error analysis, track experiments, and report results honestly, including cost.
Evaluate ultra-low-bit quantized models for on-device use: extend our quantized vs. full-precision benchmark comparisons and analyze where and why they diverge.
Agent Systems & Engineering PracticeHelp build the next generation of Postman’s in-product AI agent (Agent Mode): a deliberately minimal agent architecture that calls LLM APIs directly (tool loops, multi-step execution, checkpointing), primarily in Type Script. No prior Type Script is required; strong Python fundamentals transfer quickly. Read the source code of open-source agent harnesses and turn what you learn into design specs and prototypes. Document experiments, design decisions, and runbooks so your work is legible to the next person;
flag safety, fairness, or privacy concerns you observe in model or agent behavior.
Currently pursuing a BS, MS, or PhD in Computer Science, Data Science, or a related quantitative field. Hands-on experience training or evaluating ML models: course projects, research, hackathons, or a prior internship all count. Solid Python fundamentals: data structures, functions, basic testing; comfortable writing and reviewing code outside of notebooks. Working knowledge of at least one deep-learning framework (PyTorch preferred). Clear written and verbal communication, and a habit of documenting what you build.
Preferred QualificationsExperience fine-tuning open-weight LLMs (SFT, LoRA, RL, or distillation), with the improvement measured on a benchmark. Experience building LLM agents (tool calling, multi-step loops) or LLM evaluation harnesses/benchmarks, and reporting results with statistical rigor. A track record of shipping real software end-to-end: APIs and services, CLIs, Docker, CI/CD, cloud; public code on Git Hub is a big plus. Interest or experience in AI safety…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).