×
Register Here to Apply for Jobs or Post Jobs. X

Senior Applied AI Engineer, Agent Quality & Evaluations

Job in San Antonio, Bexar County, Texas, 78208, USA
Listing for: Flodesk
Full Time position
Listed on 2026-09-28
Job specializations:
  • Software Development
    AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 150000 - 235000 USD Yearly USD 150000.00 235000.00 YEAR
Job Description & How to Apply Below

Senior Applied AI Engineer, Agent Quality & Evaluations

Remote

Flodesk is recognized in the Inc 5000 as one of the world's fastest-growing email marketing companies, built to help entrepreneurs sell online and design emails that people love to get. We're committed to giving small businesses simple and intuitive tools that help them grow, nurture, and monetize their email list.

We’re a remote-first company headquartered in San Francisco with a globally distributed team, including in-person hubs in Da Nang (Vietnam), Barcelona (Spain), and Menlo Park (California). Our team reflects the diversity and creativity of the people we serve. Join our mission to level the playing field for small business owners through good design.

About the role

Flodesk is building toward a future where small business owners do not need to be expert marketers to grow. Today, our AI helps members create and edit emails, but that is only the starting point. We are building a system that can understand a member's business, brand, audience and past performance, then help turn a goal into a complete marketing campaign across email, workflows, forms, sales pages, segmentation, scheduling, analytics and recommendations.

Reporting to the Head of Product, AI Systems and Core Experience, you will shape how that system behaves and how we know it is good. You will own Flodesk's prompts, agent instructions, quality definitions, evaluation criteria and continuous-improvement loop by building the technical tooling, prototypes and evaluation infrastructure that make this work measurable and scalable. This is not a role for writing clever prompts in isolation.

It is hands-on systems work at that connects engineering, product and natural language. You will need to understand how context, models, tools, product state and structured outputs come together to create a member experience, then identify the right layer to change when that experience fails.

What you’ll do
  • Managemember-facing AI behavior across prompt-to-create, agentic editing and future jobs such as segmentation, scheduling, analytics and recommendations.
  • Create, test, version and document prompts and agent instructions, with a clear record of what changed, why it changed and how behavior improved
  • Define how the system should use context, member data, tools and product state. Build and prototype the agent loops, tools and structured interfaces needed to make those behaviors work
  • Own Flodesk's AI evaluation practice, including representative datasets, behavioral scenarios, scoring rubrics, regression suites and human review
  • Build and operate eval harnesses, monitoring and quality dashboards so the team can measure quality without relying on one-off manual checks
  • Diagnose failures across prompts, context, orchestration, models, tools, data and product code, then fix the right layer or partner with the engineer who owns it.
  • Set quality baselines and release gates for prompt, model and agent changes. Turn production failures, traces and member feedback into durable regression cases
  • Partner with product, design, marketing and copy experts to encode Flodesk's point of view into the system, while bringing your own judgment about how the product should behave
What you bring
  • A track record of building or meaningfully improving production LLM or agentic products, not just prototypes
  • Strong software engineering skills in Python, Type Script or a similar language that let you build reliable internal tools, evaluation systems and prototypes
  • The interaction between prompts, context, tools, agent loops, orchestration and product state is clear to you
  • LLM evaluation, observability and experimental design are areas where you have direct experience, and you know how to combine automated checks with…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary