×
Register Here to Apply for Jobs or Post Jobs. X

AI Systems Engineer - Data & State Management - Senior

Job in Salt Lake City, Salt Lake County, Utah, 84190, USA
Listing for: EY
Full Time position
Listed on 2026-08-30
Job specializations:
  • Software Development
    Cloud Engineer - Software, Backend Developer, Database Engineering
Job Description & How to Apply Below
Location:

Anywhere in Country

At EY, we're all in to shape your future with confidence.

We'll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go.  Join EY and help to build a better working world.

** The opportunity*
* We are seeking an AI Systems Engineer to own the stateful backbone of EY's AI-native platform, including the data stores, memory tiers, and event streaming systems that hold and move every piece of state in EY's Hybrid AI Multi-Environment Runtime (HAI). This role ensures that agentic AI workloads have durable, performant, and consistent access to data across cloud, on-prem, edge, client-managed, and air-gapped environments.

Within HAI, this is a distinct discipline from platform operations and model serving. This role owns everything that must remember, persist, or flow: relational and key-value state, vector and graph stores, object storage, durable workflows, and the streaming and change-data-capture pipelines that connect them. It is the layer that makes the platform stateful, reliable, and event-driven.

It is ideal for a data-infrastructure engineer who is equally comfortable operating production databases and high-throughput streaming systems, who treats data durability, consistency, and recoverability as non-negotiable in regulated client contexts, and who understands that state is the hardest part of any distributed platform to get right.

** Your key responsibilities*
* +  
** Own the memory and data stores:
** relational and durable state (PostgreSQL, DBOS durable workflows), caching (Redis/Valkey), vector stores (Qdrant/Milvus/PGVector), knowledge graphs (Neo4j), and object/block storage (MinIO, OpenEBS Mayastor), across every environment and tenant.

+  
** Own event streaming and async messaging:
** Apache Kafka (Strimzi), NATS Jet Stream (agent-to-agent), Debezium (change data capture), Apache Flink (stream processing), and Apicurio/Cloud Events (schema and event contracts).

+ Own data durability, consistency, and recoverability: replication, backup/restore, point-in-time recovery, and cross-environment data movement, tiered by RPO/RTO.

+ Build and operate streaming and CDC pipelines that move data reliably between stores and services, with schema governance and evolution that prevents breaking changes across producers and consumers.

+ Make state multi-tenant and portable, ensuring isolation, performance, and consistent semantics whether running on managed cloud services or self-hosted OSS in an air-gapped environment.

+ Provide the data and lineage substrate that downstream governance, observability, and AI knowledge capabilities depend on, as well as integrating with lineage tooling.

** Skills and attributes for success*
* + Deep expertise operating production databases and data stores at scale, including relational, key-value, vector, graph, and object storage.

+ Strong command of streaming and event-driven architectures (Kafka, NATS, CDC, stream processing) and the consistency tradeoffs they involve.

+ A durability-first mindset: thinking in terms of consistency, recoverability, blast radius, and data correctness under failure.

+ Ability to operate stateful systems consistently across managed cloud and self-hosted OSS in cloud, on-prem, edge, and air-gapped environments.

+ Strong grasp of schema governance and evolution, preventing breaking changes across producers and consumers.

+ Strong communicator able to guide consuming teams toward the right storage and streaming patterns.

+ Orientation toward reliability and toil reduction through automation and infrastructure-as-code for data systems.

** To qualify you must have*
* + Bachelor's or Master's degree in Computer Science or related technical field.

+ 8+ years operating production data infrastructure, streaming systems, or database platforms at scale.

+ Hands-on expertise with relational databases (PostgreSQL) and caching (Redis/Valkey), including HA, replication, and backup/recovery.

+ Exposure to AI/ML data patterns - embeddings, retrieval, feature/state stores for agentic workloads.

+ Production experience with vector and/or graph databases (Qdrant, Milvus, PGVector, Neo4j) in AI/ML contexts.

+ Deep experience with event streaming and messaging (Apache Kafka/Strimzi, NATS) and change data capture (Debezium).

+

Experience with stream processing (Apache Flink) and event/schema governance (Cloud Events, schema registry).

+ Experience running stateful systems on Kubernetes (operators, persistent volumes, object/block storage such as MinIO/OpenEBS).

+ Ability to define clean ownership boundaries and data/schema contracts with platform, trust, runtime, and delivery teams.

** Ideally, you'll also have*
* +

Experience with durable workflow engines (DBOS, Temporal, or equivalents).

+

Experience with data lineage and metadata tooling (Open Lineage, Marquez, or equivalents).

+ Proven track record operating data systems under compliance, security, or regulatory constraints, including data-at-rest and…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary