AI Platform Engineer
Listed on 2026-09-14
-
IT/Tech
SRE/Site Reliability, IT Infrastructure, Systems Engineer, Cloud Computing: Infrastructure & Operations
Axos Bank
Target Range: $ /Yr.
- $ /Yr. Actual starting pay will vary based on factors including, but not limited to, geographic location, experience, skills, specialty, and education.
Eligible for an Annual Discretionary Cash Bonus Target: 10%
Eligible for an Annual Discretionary Restricted Stock Units Bonus Target: 10%
These discretionary target bonuses may be awarded semi-annually based upon your achievement of performance goals and targets.
About This JobAxos Bank is hiring an AI Platform Engineer to engineer, operate, and help govern the platforms and infrastructure that power the bank’s AI and automation capabilities. This is a hands‑on, senior individual‑contributor role spanning platform engineering, the supporting data and compute tier, reliability, and platform governance — applied across a growing portfolio of enterprise AI and automation platforms rather than any single product.
The successful candidate is platform‑and‑infrastructure‑focused: equally comfortable deploying and hardening an enterprise platform, tuning a PostgreSQL cluster or Redis tier, and establishing the engineering standards and controls the platform estate runs under. The role ensures that the bank’s AI and automation platforms are reliable, performant, secure, and well‑governed as adoption scales across lines of business. This is the hands‑on engineering seat behind our AI and automation platforms, the person who deploys, hardens, scales, and governs the platforms and the infrastructure they run on.
The role offers deep technical ownership of a fast‑growing platform footprint, close partnership with senior architects in Infrastructure, Security, and Identity, and a central part in scaling the bank’s Automation Center of Excellence.
- Deploy, configure, upgrade, and operate the bank’s portfolio of enterprise AI and automation platforms across Dev/QA/UAT/Prod — including high‑availability topology, scaling, version and patch management, and full platform lifecycle
- Tune platforms for performance, capacity, and cost as adoption grows; plan and execute upgrades and migrations with minimal service impact
- Integrate platforms with enterprise identity (Entra /OIDC), secrets management, and source control; manage platform‑level RBAC and credential lifecycle
- Evaluate and onboard new AI and automation platform capabilities into the supported estate
- Engineer and operate the PostgreSQL tier supporting the platform estate — replication, high availability and automated failover, point‑in‑time recovery, connection pooling, and performance/query tuning
- Engineer and operate the Redis tier — persistence configuration, memory and eviction policy, queue durability and high availability (Sentinel or Cluster)
- Engineer the Linux hosts and Docker/Kubernetes containers running the platforms — hardening, patching, capacity, and performance tuning, in partnership with Infrastructure
- Design, implement, and validate HA/DR for the platform data tier against defined RTO/RPO — failover behavior, backup validation, and recurring recovery exercises
- Instrument the platforms and their infrastructure for service health, replication/failover events, performance, and capacity using the enterprise observability stack
- Build and maintain dashboards, alerts, and runbooks; lead incident response and root‑cause analysis for platform and data‑tier events
- Build and maintain CI/CD pipelines and promotion workflows for platform artifacts and configuration across environments
- Manage infrastructure‑as‑code (Terraform or equivalent) for the platform data and compute tier, and automate routine platform operations
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).