×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer: Skills Remote; US or EU timezones

Remote / Online - Candidates ideally in
New York, New York County, New York, 10261, USA
Listing for: Meetkay
Full Time, Remote/Work from Home position
Listed on 2026-09-28
Job specializations:
  • Software Development
    Backend Developer, AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: Software Engineer: Skills Remote (US or EU timezones)
Location: New York

Engineering Remote (US or EU timezones) Full-time

Yak Labs is building a desktop AI application that turns knowledge workers into AI power users, without docs, API keys, or knowing what to ask for. One surface for agentic AI work: proactive, personalized, progressively powerful. The gap between AI power users and everyone else isn’t talent, it’s setup. We close it.

The team: Alex Goddijn spent four years at Palantir, then founded an agent harness company that was acquired by Poolside AI. Seb Goddijn led internal AI ’re building Kay so that kind of leverage isn’t limited to companies with their own internal AI team.

We have funding. Product and design work in person in New York City, and some engineering roles are remote. Small team, high intensity, no passengers.

The job

Skills are how Kay gets better at helping someone do their work. A skill captures the procedure, judgment, and taste for a kind of task: how this person reviews a contract, preps a board update, or triages their inbox. You own the skill system end to end: how skills are written, found, loaded at the right moment, learned from someone’s own work, shared across a team, and proven to help.

This is where “learns how you work” becomes real or stays a slogan.

Problems you’ll work on

These are examples, not lanes. You will move between them, while taking real ownership of the ones that match your interests.

  • Learning skills from use: Every session is evidence of how someone likes their work done. How do you spot a repeated workflow, turn it into a skill, and propose it at a moment the person will say yes, without burying them in suggestions they didn’t ask for?
  • The right skill at the right time: One person may have hundreds of skills from themselves, their team, and their plugins. How does an agent pick the few that apply to this task, combine them, and ignore the rest, inside a finite context window?
  • Proving a skill helped: A skill that reads well can still make outcomes worse. How do we evaluate skills against real tasks, notice when one goes stale or starts hurting, and fix or retire it?
  • Skills that travel: What someone built for themselves should be easy to share with their team or a marketplace, without leaking private context, and it should keep improving as others use it.
What you’ll do
  • Build the skill runtime: discovery, retrieval, loading, and composition inside the agent loop.
  • Build the pipeline that turns session history into new skills and improvements to existing ones, and the review surfaces where a person accepts, edits, or rejects them.
  • Build evals for skills: task suites, graders, and production signals that tell us whether a change helped.
  • Write first-party skills for the work our users do most, and set the bar for what a good skill looks like.
  • Own sharing: team and marketplace distribution, versioning, and provenance.
Who you are
  • You’ve built with LLMs past the prompt: retrieval, evals, agent loops. You know a prompt change is a code change, and you test it like one.
  • You write very well. A skill is instructions for a model, and most bad skills are bad writing.
  • You care about measurement. You can design an eval that tells you something true, and you distrust one that only ever goes up.
  • You have good product instincts about when to interrupt someone and when to stay quiet.
  • You lean on skills, rules files, or custom instructions heavily in your own work, and you have opinions about what makes them work.
Bonus
  • You’ve built recommendation, ranking, or retrieval systems in production.
  • You’ve built LLM eval infrastructure.
  • You’ve shipped PRs to open-source agentic harness projects.
Who we look for

These apply to every role at Yak Labs.

  • First principles thinking. We’re building in a category that’s being defined in real time. Best practices…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary