Principal Engineer
Listed on 2026-07-09
-
Software Development
AI Engineer (Applied/Software), Software Architect, DevOps, Backend Developer
Principal Engineer
Organizations everywhere struggle under the crushing costs and complexities of "solutions" that promise to simplify their lives. To create a better experience for their customers and employees. To help them grow. Software is a choice that can make or break a business. Create better or worse experiences. Propel or throttle growth. Business software has become a blocker instead of ways to get work done.
There's another option. Freshworks. With a fresh vision for how the world works.
Freshworks Inc. builds uncomplicated service software that delivers exceptional employee and customer experiences. Our people-first approach to AI eliminates friction, helping businesses reduce complexity, lower cost-to-serve, and deliver faster, more human support through enterprise-grade yet easy-to-use CX and IT solutions. Nearly 75,000 companies, including Bridgestone, New Balance, Nucor, S&P Global, and Sony Music, trust Freshworks to power their Employee Experience (EX) and Customer Experience (CX) operations.
Fresh vision. Real impact. Come build it with us.
Job DescriptionWe are seeking a Sr. Staff / Principal AI Systems Engineer to build, scale, and operate the AI Agent Platform that powers reasoning-driven assistants and autonomous agents across Freshworks — and to make that platform operate itself. You will own the systems architecture and engineering of a multi-tenant agent runtime that serves agentic workloads at high throughput and low latency, and you will pioneer an Agentic AIOps approach where autonomous agents monitor, diagnose, and remediate the platform in production.
This is a hands-on systems engineering role at the staff/principal level, with a strong platform-scale and agentic-operations center of gravity. You'll design how thousands of agents are orchestrated, served, observed, and kept healthy at enterprise scale, and you'll set the technical direction which the broader engineering organization builds against. The level (Sr. Staff vs. Principal) will be calibrated to your scope of technical ownership and organizational impact.
Key Responsibilities:
- Architect and build the core AI Agent Platform — agent runtime, orchestration, tool/API invocation, state and memory management, and the retrieval/knowledge services agents reason over
- Design for scale and efficiency: high-throughput multi-tenant serving, concurrency and queueing for agent workloads, model/inference routing, caching, and cost-aware execution
- Build the control plane and systems primitives other teams use to define, deploy, version, and operate agents safely
- Drive latency, throughput, and cost optimization across the agentic request path (planning → retrieval → tool calls → generation)
- Architect agentic operations workflows where autonomous agents observe platform telemetry, reason about anomalies, perform root-cause analysis, and execute remediation — shifting operations from human-driven to agent-driven
- Design multi-agent operational loops (detection, diagnosis, remediation) that collaborate, escalate, and hand off to on-call humans with clear rationale and audit trails
- Build closed-loop self-healing for the platform: auto-detection and repair of failing agents, degraded tools/connectors, stale knowledge, failed ingestion, and retrieval/index drift
- Define guardrails, confidence thresholds, and human-in-the-loop controls that make autonomous remediation safe at multi-tenant scale
- Apply LLMs to operations directly — incident summarization, runbook generation, on-call copilots, and natural-language querying of platform telemetry
- Instrument the platform end to end: distributed tracing across planning-retrieval-tool-generation loops, metrics, structured logging, and event correlation so multi-agent behavior is explainable and debuggable
- Define golden signals for both system health and agent quality — task success rate, tool-call accuracy, grounding/hallucination rates, latency, cost-per-task, throughput — as first-class telemetry the operating agents act on
- Establish SLOs/SLIs and error budgets for agent workflows, with alerting that feeds the agentic-ops layer
- Engineer the platform as a resilient, event-driven, cloud-native…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).