×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer - AIOps

Job in Scottsdale, Maricopa County, Arizona, 85261, USA
Listing for: Early Warning Services, LLC
Full Time position
Listed on 2026-08-28
Job specializations:
  • Software Development
    DevOps, Software Engineer, Software Architect, Cloud Engineer - Software
Job Description & How to Apply Below
Position: Staff Software Engineer - AIOps
At Early Warning, we've powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle , Paze , and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.

Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.

Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.

Staff Software Engineer - AIOps

Overall Purpose

The Staff Software Engineer - AIOps is a senior hands-on technical contributor responsible for designing, building, and operating the AIOps agent platform that enables AI agents to observe events across Early Warning technology environments, generate diagnoses and recommendations, and execute approved actions within defined safety controls. The role focuses on the software systems, control loop, tool interfaces, telemetry, evaluation capabilities, and operator experience required to run agentic capabilities safely and reliably in production;

it is not primarily a model development or research role.

This position executes complex technical work with general direction, independently makes implementation decisions within established architecture and standards, and may seek guidance for novel, highly ambiguous, or cross-enterprise decisions. The Staff Engineer applies software engineering, distributed systems, and platform engineering principles and partners across Engineering, Technology Operations, Security, Risk and Compliance, and other technology teams to establish reusable patterns and guardrails for progressive levels of automation.

Essential Functions
  • Apply mature software engineering practices across the agent runtime, tool layer, APIs, and operator experience, including versioned interfaces, automated testing, code review, release management, observability, and regression coverage.
  • Execute complex AIOps engineering assignments with general direction; break work into deliverable components, identify practical implementation solutions, raise risks and dependencies, and seek guidance when decisions extend beyond established architecture or standards.
  • Contribute to the design and evolution of the agent platform architecture; maintain architecture decision records, define interface standards, and evaluate significant design and build-versus-buy tradeoffs with appropriate technical guidance.
  • Design, build, test, and operate the core agent runtime, including event ingestion, context assembly, planning, tool invocation, verification, escalation, and response handling.
  • Design and maintain scoped, typed, and auditable tool interfaces that enable agents to interact with CI/CD platforms, infrastructure automation, Kubernetes, network services, ITSM platforms, observability systems, and secrets management solutions.
  • Define and implement controls for agent actions, including read-only and advisory capabilities, human approval requirements, narrowly scoped unattended actions, permission boundaries, blast-radius limits, dry-run capabilities, reversibility, and emergency shutdown controls.
  • Define tool contracts, versioning standards, testing requirements, and intended behavior so agent-accessible capabilities are managed as reliable production interfaces.
  • Build event-routing capabilities that receive and prioritize alerts, pipeline failures, tickets, operational requests, and other technology events and provide appropriate context to the agent runtime.
  • Establish telemetry capabilities used by the agent, including logs, metrics, and traces from relevant technology platforms, and contribute to standards for telemetry quality and schema evolution.
  • Design and operate model routing, session and state management, retries, timeouts, and cost and latency controls appropriate for a production service.
  • Implement comprehensive auditability for events received, decisions generated, approvals obtained, and actions executed to support operational, risk, compliance, and examination requirements.
  • Develop evaluation capabilities for agent decision quality, including historical incident replay, controlled or shadow-mode evaluation, measurement of proposed actions, and regression testing.
  • Build feedback mechanisms that incorporate human approvals, overrides, and operational outcomes into measurable improvements to platform quality.
  • Package reusable agent capabilities, tool integrations, APIs, documentation, and implementation patterns so other technology teams can extend and adopt the platform through defined self-service practices.
  • Develop and maintain the operator experience, including approval workflows, agent activity and decision history, telemetry views, and APIs that expose agent state to user…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary