Senior AI Data Engineer
Listed on 2026-10-09
-
Software Development
Data Engineering
Comply is the leading provider of compliance SaaS and consulting services for the global financial services sector. With more than 5,000 clients and hundreds of employees across the globe, Comply empowers Chief Compliance Officers and their teams to proactively manage regulatory obligations, mitigate risk, and scale with efficiency and confidence.
Comply serves thousands of global financial services clients including broker-dealers, insurers, investment banks, private funds, RIAs, and wealth managers who rely on Comply offerings to power their compliance programs. To learn more about Comply, visit
About the roleComply is looking for a Senior AI Data Engineer to turn our semantic data models into the knowledge graphs, vector search, and LLM-powered pipelines that make Comply’s financial and regulatory data genuinely AI-ready. Comply is the world’s leading aggregator of financial and regulatory data for compliance, ingesting and enriching vast volumes of complex, high-variance data from hundreds of brokers and data providers across the US and beyond.
The Comply Data Platform is a strategic initiative at the heart of that mission: a modern, cloud-native semantic layer built from the ground up—using JSON-LD as its semantic language—to power AI-driven insights, regulatory analytics, and next-generation data products. In this hands-on role at the intersection of knowledge representation, AI infrastructure, and data platform engineering, you’ll own the delivery of semantic layer components, work hand-in-hand with data engineers, architects, and our ontologist, and make sure AI-ready data products are reliable, performant, and actually adopted.
You’ll join a new team being created to enable Comply’s AI ambitions.
you’re an engineer who turns semantic and ontological concepts into concrete, production-grade systems—and you’re energized by greenfield environments where the architecture is still being defined.
What you’ll do- Implement JSON-LD-based semantic models from our ontologist into production data systems, and build knowledge graph structures that reflect canonical domain models.
- Develop and manage graph database schemas, queries, and ingestion pipelines, keeping semantic consistency between ontology definitions and downstream data products.
- Design and implement embedding pipelines that represent Comply’s financial and regulatory data in vector space.
- Build and operate vector database infrastructure for semantic search and similarity retrieval.
- Implement RAG architectures that ground LLM outputs in Comply’s proprietary data, and evaluate and integrate LLM tooling suited to our use cases.
- Build reliable, observable data pipelines that feed the semantic layer from upstream broker and regulatory sources, applying Data Ops practices such as testing, monitoring, lineage tracking, and SLAs.
- Work with Data and Backend Engineers to embed semantic models into APIs and data contracts, ensuring the semantic layer scales with data volume and platform growth.
- Partner closely with the Ontologist so implemented models faithfully reflect domain intent, and support consuming application teams in adopting AI-ready data products.
- You’ve shipped your first semantic layer components into production, turning ontology models into working knowledge graph structures.
- You’ve stood up or extended embedding and vector infrastructure that a consuming team is actively using for semantic search or RAG.
- You’ve put Data Ops practices in place so pipelines are observable and maintain semantic quality from ingestion through to consumption.
- Strong hands-on data engineering experience, with a focus on semantic or AI data infrastructure.
- Experi…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).