Founding Engineer, Search & Relevance
Listed on 2026-09-17
-
Software Development
About Whym
Whym is an AI-powered app that helps people make time for what matters most. We suggest activities aligned with your values and help you coordinate plans with close friends. We are a two-person founding team with experience from Google, Linked In, Lyft, and earlier-stage startups. Whym is a Public Benefit Corporation with pre‑seed funding.
About the roleYou will own Whym's recommendation quality end to end. Every suggestion we put in front of a user passes through systems you will build: the catalog pipeline that assembles what there is to do, the ranking that orders what we show, and the evaluation that tells us whether it works. Most of those systems do not exist yet. The mandate is scrappy: stand them up from first principles, capture the low-hanging fruit, and let eval help you find the next win.
Quality is central to the product's promise. Users respond to the concept. Whether they stay depends on whether each suggestion lands.
This is a build-heavy role.
Expect to spend most of your time shipping production code
: services, pipelines, indexes, and harnesses, with modeling and analysis in service of what you ship.
This is a tiny startup. Everyone does customer support, including the CEO and CTO. You will too. Some weeks you will debug a ranker. Other weeks you will review user feedback, build a labeling UI, or sit with a tester to understand why a suggestion missed. You should find that appealing, not beneath you.
What you will work on- Ranking and relevance
. Design, build, and ship the systems that decide what Whym suggests to whom, and in what order. Own the full stack: retrieval, candidate generation, ranking, and freshness. Iterate based on eval signals and user feedback. Today's ranker is an LLM self-score placeholder. You will replace it with something learned, one measured win at a time. - Data pipeline and indexing
. Build the infrastructure that turns the sprawl of internet-accessible sources (event listings, venue pages, local calendars) into a coherent, current view of what there is to do in the real world. You will design the index, decide what we own versus fetch on demand, and expand source coverage as we open new markets. Scraping, integration, and data partnerships run through this function. - Evaluation
. Build the system that tells us whether a quality change helped or hurt. Design rubrics. Run human eval. Operate LLM-as-judge ntain offline test sets. Set up AB tests. Create the metrics dashboards the team lives by. Without eval, quality work is guesswork. You will make sure we are not guessing. - Quality research
. Answer the hard questions with data. Where are we failing? Which suggestions convert? What does a good session look like? Then build agentic workflows that ask those questions continuously: agents that sweep for failure patterns, propose ranking changes, and check them against eval. Research sits next to ranking and eval because all three only work when they feed each other.
- 8+ years in engineering, with meaningful time on search, feed, recommender, ads, or ranking systems at consumer scale.
- Strong software engineering foundations. You have designed, built, and operated production systems end to end: services, schemas, pipelines, and the infrastructure under them. Not only models and notebooks.
- You have shipped ranking or retrieval systems that real users depend on. We want to hear about specific decisions you made, what you measured, and how the system improved over time.
- Solid grounding in ranking and experimentation fundamentals. You can design an experiment, reason about metrics, and tell a good signal from a noisy one. You do not need to be a researcher.
- Hands‑on with evaluation. You know how to build test sets, write rubrics, and run…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).