Distributed Systems Architect: Bare-Metal and -Scale
Distributed Systems Architect:
Bare-Metal and High-Scale
You will own the architectural blueprints for Analog’s on-premises open-source infrastructure, designing the clustering, replication, and failover topologies across messaging, streaming, and database layers that our Dev Ops team then automates at scale across client deployments. This is a hands-on, implementation-grade architecture role where the quality of your output directly determines whether a physically isolated or air-gapped client environment survives hardware failure at petabyte scale.
WhatYou’ll Do Open-Source Infrastructure Mapping
- Map Analog’s Azure PaaS components to open-source equivalents:
Event Hub to Kafka/Redpanda, IoT Hub to EMQX, ADX to Click House or Apache Druid, and Blob Storage to Ceph/MinIO - Produce side-by-side equivalence assessments documenting feature gaps, operational differences, and migration risk for each component transition
- Validate client hardware specifications and assess compliance with air-gap security requirements prior to each deployment
- Design multi-node clustering, rack‑aware replication, and quorum topologies for bare‑metal Kafka/Redpanda and PostgreSQL (Patroni) clusters
- Design resilience protocols for hardware and network switch failures in physically isolated or air‑gapped client environments
- Architect RocksDB state backend tuning for stateful Apache Flink workloads, including compaction strategy, block cache sizing, and write‑ahead log configuration
- Tune Linux kernel parameters, NUMA bindings, network ring buffers, and NVMe I/O scheduler configurations to eliminate hardware bottlenecks under petabyte‑scale write workloads
- Define CPU affinity, IRQ balancing, and huge page configurations for latency‑sensitive broker and database processes
- Author reproducible benchmark harnesses to validate configuration changes against client hardware before production rollout
- Deliver precise, implementation‑ready configuration blueprints and runbooks for Terraform and Ansible automation, documents the Dev Ops team can execute without architectural interpretation
- Maintain a library of parameterized reference architectures covering single‑rack, multi‑rack, and geographically distributed bare‑metal topologies
- Conduct pre‑deployment reviews of client hardware specs and provide go/no‑go assessments with remediation guidance
- 10+ years architecting and operating distributed infrastructure at TB to PB scale in production environments
- Expert‑level Kafka or Redpanda clustering: partitioning strategy, ISR tuning, consumer group management, compaction policies, and operational lifecycle at scale
- Deep Linux systems engineering: kernel networking subsystems (TCP buffer tuning, interrupt coalescing), storage fabrics, and NUMA‑aware process binding
- Proven track record deploying HA database clusters without cloud load balancers:
Patroni, Pacemaker, or equivalent in production - Experience deploying software‑defined storage (Ceph, MinIO) on physical hardware in production, including CRUSH maps, erasure coding, and performance tuning
- Ability to produce implementation‑ready blueprints and runbooks; your output must be directly actionable by a Dev Ops automation team without architectural interpretation
Analog builds industrial intelligence infrastructure for the physical world. We deliver high‑throughput data pipelines, real‑time analytics, and edge‑to‑cloud connectivity for mission‑critical environments where cloud dependency is not an option. Our clients operate in regulated, air‑gapped, and physically demanding settings and they depend on us to get the infrastructure right from day one.
This is a full‑time, on‑site role.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).