Senior Kafka DevOps Administrator
Listed on 2026-08-06
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Disaster Recovery IT, IT Infrastructure
Job Summary
We are seeking a Senior Kafka Dev Ops Administrator to manage, optimize, and modernize enterprise-scale Confluent Platform and Apache Kafka environments. This role will focus on operating highly available on-premises Kafka clusters, leading hardware refresh initiatives, improving platform reliability, implementing automation, and ensuring high availability, disaster recovery, and operational excellence. Experience with AWS MSK is desirable to support hybrid and multi-environment deployments.
Key ResponsibilitiesDesign, deploy, administer, and optimize highly available Kafka clusters across on-premises and cloud environments.
Lead Kafka infrastructure upgrades, hardware refreshes, cluster migrations, and cutover activities with minimal downtime.
Configure and manage Kafka topics, partitions, replication, retention policies, quotas, and consumer groups.
Administer Kafka ecosystem components including Kafka Connect, Schema Registry, Mirror Maker/Confluent Replicator, and REST Proxy.
Perform Kafka performance tuning, capacity planning, benchmarking, and cluster right-sizing.
Implement automation for provisioning, deployment, monitoring, and operational tasks using scripting.
Monitor Kafka infrastructure using Data Dog, Grafana, JMX exporters, and centralized logging solutions.
Develop disaster recovery strategies, backup/restore procedures, multi-region deployment plans, and incident response processes.
Troubleshoot Linux, networking, and Kafka platform issues to ensure maximum availability and performance.
Produce comprehensive technical documentation, operational runbooks, and knowledge transfer materials.
5+ years of experience in Systems Engineering, Dev Ops, Platform Engineering, or Site Reliability Engineering (SRE).
4+ years of hands‑on experience managing Apache Kafka in large-scale production environments.
Strong expertise in Kafka internals, including partitions, replication, retention, compaction, ISR, and consumer group rebalancing.
Hands‑on experience with Kafka Connect, Schema Registry, Mirror Maker, and Confluent Replicator.
Strong Linux administration, networking (TCP/IP, DNS, load balancing), and performance troubleshooting skills.
Experience with automation and scripting for infrastructure management.
Hands‑on experience with monitoring and observability tools such as Data Dog, Grafana, JMX exporters, and log aggregation platforms.
Experience implementing disaster recovery, multi‑region architectures, and incident management processes.
Excellent technical documentation and communication skills.
Experience with Apache Kafka on AWS MSK.
Experience with Confluent Platform administration and operations.
Knowledge of Kafka Streams and ksqlDB.
Experience performing hardware refreshes, cluster rebuilds, or large-scale Kafka migrations with minimal downtime.
Bachelor's or Master's Degree in Engineering or a related technical discipline.
Strong analytical, troubleshooting, and problem‑solving skills.
Ability to work effectively in a collaborative Dev Ops and SRE environment.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).