×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff; QA Engineer - Agentic Systems

Job in New York, New York County, New York, 10261, USA
Listing for: Solstice
Full Time position
Listed on 2026-06-30
Job specializations:
  • Software Development
    AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 160000 - 300000 USD Yearly USD 160000.00 300000.00 YEAR
Job Description & How to Apply Below
Position: Member of Technical Staff (QA Engineer - Agentic Systems)
Location: New York

About the role

We're hiring our first dedicated QA lead to own quality for the AI that powers Solstice. Our platform generates regulated pharmaceutical marketing content for the brands we work with, so when the output is wrong, say an unsupported claim or a missing safety disclosure, it becomes a real compliance problem and not just a bug to file.

What makes this hard is that the system is probabilistic. The same prompt can return different answers, "correct" is often a judgment call, and the exact-match assertions that traditional QA relies on don't apply. We need someone who can measure quality anyway and build the evals that catch regressions before a client does, so the rest of the team can keep moving quickly.

This role is about more than the models, though. Just as much of what keeps customers happy is ordinary product reliability: that the app does what it should, and that a frontend tweak or a new feature doesn't quietly break something in production. You'll own that side too, with a solid end-to-end suite in Playwright and the hands‑on manual testing that catches what automation misses.

This is a senior, hands‑on engineering job. Most of your time goes to writing code and building test infrastructure, but you'll also dig into manual testing whenever that's the fastest way to find a problem. Your work will reach across the whole product, from the backend services to the frontend customers use every day.

What you'll be working on
  • Build our evaluation systems. Because we can't check an output against a single correct answer, you'll design the evals that score quality instead and decide, with evidence, what is good enough to ship.

  • Make models and prompt changes safe. We swap models and rewrite prompts constantly. Your tooling should flag a drop in quality, a jump in cost, or a latency regression before a customer runs into it.

  • Test the agents for the ways they actually fail. Agents drift off their goal, loop on the same tool call, pick the wrong tool, or get hijacked by a malicious instruction buried in a document we ingest. Those are the cases you'll design for.

  • Protect the compliance‑critical paths. The checks that keep an unsupported claim or a missing disclosure out of a finished asset are the ones that matter most, and you'll own how we test them, including verifying claims against approved source material.

  • Own end‑to‑end testing across the app. Build and maintain a Playwright suite that exercises the real user flows, from login through content creation and review, so a frontend or API change can't quietly break something a customer depends on.

  • Run hands‑on manual and exploratory QA. Automation misses things, especially on new features and messy UI states. You'll test releases by hand, dig for the edge cases, and be the last set of eyes before we ship.

  • Get CI/CD quality gates in place. Today nothing runs automatically when someone opens a pull request: no tests, no linting, no type checks. Building that is yours.

  • Use production as a test bed. We already trace and monitor what the system does once it's live. You'll turn those signals into drift detection and into new regression tests whenever something slips past us.

  • Harden the background jobs. A lot of our work runs in long pipelines, so they need to survive retries, timeouts, and worker crashes without dropping or duplicating work.

  • Set the testing bar. As our first QA hire, you'll define what good testing looks like here and help the rest of the team write code that's easy to trust.

What we're looking for

This is an engineering role first. You should be comfortable building testing and evaluation tools in code, and equally comfortable rolling up your sleeves for hands‑on manual testing when that's what the situation needs.

Must-haves
  • Strong Python, and real experience building test infrastructure and getting it to run automatically in CI/CD.

  • Strong end‑to‑end and UI test automation, especially with Playwright.

  • A genuine manual and exploratory QA discipline. You can test a feature by hand, find the edge cases, and own release sign‑off.

  • Experience testing non‑deterministic, ML, or LLM‑based systems, or the appetite to build that capability from…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary