AI Platform Engineer: Agent & Retrieval Infrastructure
Listed on 2026-09-18
-
Software Development
AI Engineer (Applied/Software), Data Engineering
About Bedrock Ocean
Bedrock Ocean builds and operates autonomous underwater vehicles (AUVs) that collect georeferenced ocean-floor data at commercial scale. We deliver bathymetric and imagery data products to customers through our own platform, and we're scaling toward continuous, around-the-clock data collection campaigns spanning months at a time.
We are building AI agents on Amazon Bedrock to support our ocean data, internal operations, and customer platform. This role owns that architecture.
(One note on names. Amazon Bedrock is the AWS service. Bedrock Ocean is us. They are unrelated, and we are aware it is confusing.)
The RoleWe are looking for a Staff Platform Engineer to lead our AI architecture. This role goes beyond building agents on existing platforms; you will create the infrastructure itself, including the orchestration layer, the data and retrieval pipeline, and the security model required to work with production data. You will also build the tools and abstractions that allow our engineering team to implement AI features independently.
This position combines software engineering, data engineering, and infrastructure operations. You will manage the full lifecycle of our Amazon Bedrock implementation, from initial data chunking to IAM access controls. While some of our data pipelines are already in place, they will require significant expansion, and others will need to be built from scratch.
Security is core to this role, not an afterthought. Because agents with tool access represent a new kind of system actor, you will define their operational boundaries, including what they can access, the actions they can perform autonomously, and the monitoring required to detect issues.
Our roadmap prioritizes internal engineering and operational systems first to ensure a fast feedback loop, followed by our ocean and survey data products. Customer-facing retrieval is the final, high-stakes phase. You will play a key role in defining this sequence. While you will not be responsible for the core data transport design (store-and-forward or hub-and-spoke), you will work closely with that team to ensure it meets our retrieval and data freshness requirements.
WhatYou'll Do
Architect Agent Orchestration: Design the Amazon Bedrock integration, including agent and action group configuration, backend APIs, model access, throughput, and cross-environment deployment.
Manage Retrieval Data Plane: Own the end-to-end retrieval pipeline from ingestion and chunking to embedding and storage in Amazon Open Search Serverless. Focus on optimizing for index design, cost, and capacity.
Extend Data Pipelines: Adapt ingestion pipelines for internal knowledge, ocean data, and customer platforms, addressing challenges specific to geospatial and large-binary datasets.
Secure AI Infrastructure: Implement robust security including Bedrock Guardrails, VPC and Private Link network boundaries, least-privilege IAM, and audit trails to ensure data isolation.
Define Agent Governance: Build the mechanisms to enforce approval boundaries for autonomous actions, ensuring agents are safe and monitored.
Establish LLMOps & Observability: Implement comprehensive monitoring for tracing, tool calls, and retrieval performance, using Cloud Watch and LLM-specific tools like Langfuse or Phoenix.
Build Evaluation Frameworks: Create the infrastructure to run automated evaluations, track results, and manage release gates for model accuracy.
Enable Engineering Productivity: Provide the team with abstraction layers, SDKs, and self-service environments that allow engineers to ship AI features independently.
Operational Excellence: Manage the environment as code across all stages, ensuring deployment safety and participating in incident reviews.
8+ years in software and infrastructure engineering, including deep production backend experience (Python or Type Script preferred, Go fine) and staff-level ownership of technical direction.
Hands-on experience standing up Amazon Bedrock in production: agents, knowledge bases, guardrails, model access, and the throughput and quota decisions that come with them.
Containerized service deployment on ECS, EKS, or Lambda, with CI/CD you have owned rather than inherited. The models are managed, but the backend APIs, tool endpoints, and ingestion jobs still run somewhere real.
Practical RAG and vector search experience: embeddings, chunking strategies, semantic search quality, and operating a managed vector database (Open Search…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).