Location
United Kingdom;
Argentina;
Brazil;
Bulgaria;
Canada;
Chile;
Colombia;
Cyprus;
Czech Republic;
Hungary;
Ireland;
Lithuania;
Mexico;
Peru;
Poland;
Portugal;
Romania;
South Africa;
Spain;
Sweden;
Switzerland
Full time
Location TypeRemote
DepartmentEngineering, SRE / Devops
The teamYou will join a senior team of DBAs, SREs, and platform engineers responsible for the database and vector data platforms behind Kraken's most critical services. The group combines deep operational expertise with a platform mindset: we design systems that let engineering teams move quickly while preserving strong guarantees around performance, correctness, security, and reliability.
The team is entering an important phase of evolution as Kraken continues to expand globally and as AI-enabled product and internal platform use cases grow. Over the next year, we are focused on strengthening Postgre
SQL operational excellence, building reliable vector data capabilities, preparing for significant future scale, and increasing automation and self-service so database operations become safer and more efficient at scale.
This is not a traditional keep-the-lights-on DBA team. Our work centers on scaling high-throughput Postgre
SQL environments, operating vector search and embedding stores with production discipline, shaping reliable patterns for high-traffic systems, and transforming operational knowledge into automated, auditable workflows that uplift the engineering organization.
- Design, develop, and maintain high-quality applications using React
- Scale and tune high-throughput Postgre
SQL clusters that support continued global expansion and new product initiatives. - Own Postgre
SQL reliability fundamentals: WAL behavior, checkpoints, autovacuum, query planning, locking, replication lag, backup/restore, and capacity planning. - Help build and operate vector data capabilities, including pgvector and dedicated Vector
DB platforms where appropriate. - Define production patterns for embeddings, approximate nearest neighbor indexes, metadata filtering, recall/latency tradeoffs, reindexing, data freshness, and drift management.
- Strengthen high-availability, disaster-recovery, PITR, and backup approaches through sound design and regularly validated procedures.
- Reduce manual operational work by building automation, improving process consistency, and enabling safe, low-friction database workflows.
- Improve observability and alert quality by championing meaningful metrics, reducing noise, and ensuring operational clarity across Postgre
SQL and vector workloads. - Enhance database security through robust access controls, disciplined patching and upgrade practices, encryption, auditability, and secure operational patterns.
- Contribute to modern platform initiatives, including containerized environments, infrastructure-as-code workflows, Git Ops, reproducible deployments, and self-service database operations.
- Partner with service, AI, and platform teams to drive better performance patterns, operational readiness, data hygiene, and safe use of vector search across the engineering organization.
- Participate in on-call rotations with a long-term focus on making on-call predictable, well-instrumented, and shaped by preventative engineering.
- 5+ years operating Postgre
SQL in high-volume production environments, including performance tuning, replication, backup/restore, upgrades, and incident troubleshooting. - Strong understanding of Postgre
SQL internals and operations: MVCC, transaction isolation, locks, WAL, checkpoints, autovacuum, bloat, statistics, query planner behavior, partitioning, and index strategy. - Hands-on experience with high availability and read scaling: streaming replication, replication slots, failover, lag management, PITR, backups, and disaster-recovery drills.
- Experience with connection pooling and traffic management for Postgre
SQL, especially PgBouncer, HAProxy, Kubernetes service routing, or comparable patterns. - Practical Vector
DB or vector search experience, such as pgvector, Qdrant, Milvus, Weaviate, Open Search vector search, or similar systems. - Ability to reason about vector index types and operational tradeoffs, including HNSW/IVFFlat-style indexes, recall, latency, memory, ingestion throughput, metadata filters, rebuilds, and versioned embeddings.
- Practical experience with CI/CD, Git Ops, and Infrastructure-as-Code workflows. Terraform experience is ideal.
- Solid cloud, Linux, storage, and networking fundamentals.
- Experience with containers and orchestration platforms, including building container images and managing Kubernetes workloads at scale.
- Strong security instincts around access control, credential lifecycle, encryption, auditability, upgrade processes, and safe operational workflows.
- Observability expertise: monitoring, alerting hygiene, SLOs, dashboards, query-level visibility, and readiness for incident response.
- Strong communication and collaboration skills with the ability to partner with stakeholders,…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: