×
Register Here to Apply for Jobs or Post Jobs. X

Senior Applied Research Scientist

Job in Toronto, Ontario, C6A, Canada
Listing for: ServiceNow
Full Time position
Listed on 2026-07-25
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, Service Now is the AI control tower for business reinvention. Our Service Now AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better.

We’re building an AI‑native culture where technology and talent are unstoppable together. And we’re just getting started. Join us to put AI to work for people.

About the Team
The Agentic Engineering org at Service Now is the customer‑obsessed engineering group that builds a conversational AI experience that turns enterprise intent into completed work. We advance how enterprise AI reasons, remembers, and executes. The Agent Orchestration team — the team you’ll join — owns the execution core: the agent harness, orchestration runtime, multi‑agent coordination, memory management, and the evaluation frameworks that ensure agents behave correctly in production.

Every autonomous action Otto promises depends on what this team ships. By joining our team, you’ll be at the forefront of our AI transformation journey, backed by the global scale of Service Now and the agility of a high‑growth environment. We are looking for world‑class talent to help us extend agentic AI to every employee across every corner of the business.

What You’ll Do
As a Senior Applied Research Scientist, you will own significant parts of the agent harness — the infrastructure layer that enables AI agents to reason over real enterprise data, take action across workflows, and run safely at Fortune 500 scale.

Harness engineering:
Design and build the agent execution harness — the orchestration layer that routes inputs, manages context, invokes tools, handles retries, and surfaces execution state across multi‑step agentic workflows

Reliability at scale:
Own the runtime's fault tolerance, latency, and throughput; design for enterprise workflows that cannot fail silently or non‑deterministically

Observability:
Instrument the harness with tracing, cost attribution, and latency visibility so the team can reason about agent behavior in production and catch failures before customers do

Prompt infrastructure:
Build prompt management systems — versioning, templating, and systematic evaluation — that keep agent behavior stable across model updates and configuration changes

Eval engineering:
Design and own evaluation frameworks (unit evals, integration evals, production monitors) that measure agent quality, catch regressions, and drive data‑informed decisions

LLM integration:
Integrate with and abstract over frontier LLMs, managing model routing, fallback strategies, cost and latency tradeoffs in production

Technical leadership:
Raise the technical bar through architecture decisions, code reviews, and coaching — particularly on agentic design patterns and production AI discipline.

System boundary design:
Define where agent logic lives — what's a tool call, a sub‑agent, a hardcoded path, or a human escalation — and establish those design standards across the team

Qualifications
To be successful in this role you have:

4+ years building production software systems with a strong track record on reliability, performance, and scalability

Hands‑on experience shipping generative AI products — not just integrating LLM APIs or building prototypes, but owning AI‑powered features that production users depend on

Solid depth in how large language models work: failure modes, context constraints, and how prompt design shapes model behavior at scale

Practical prompt engineering experience: systematically designing, versioning, and evaluating prompts across model updates or A/B evaluation cycles

A real track record in eval engineering — not just familiarity, but a portfolio of evaluation suites designed, shipped, and used to drive quality decisions in production AI systems

Cost and efficiency awareness at the system level: experience reasoning about model routing, inference…
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary