Principal Kafka Engineer
Listed on 2026-09-02
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Data Engineering
Principal Kafka Engineer
Required
Skills & Experience:
10+ years of software engineering experience with strong Java Object-Oriented Programming (OOP) and backend development (Spring Boot). 3-5+ years of Kafka administration/architecture experience, including Apache Kafka or Confluent Platform, Schema Registry, Kafka Connect, Avro, KSQL, and cluster management (Zoo Keeper/KRaft). Strong software architecture background with experience designing scalable, enterprise integration solutions and creating architecture artifacts. Vulnerability Management and platform security expertise, including secure system design and operational best practices.
Experience with modern cloud-native technologies including Docker, Kubernetes, CI/CD, Git Hub Actions, and Infrastructure as Code (Terraform, Ansible, or Salt).
Nice to Have
Skills & Experience:
Front-end development experience with Angular, React (preferred), or other Type Script-based frameworks on a platform team Expertise in Fin Ops. Cloud provider certifications, such as Azure Certified Solutions Architect. Confluent Certified Administrator for Apache Kafka (CCAAK). TOGAF
Job Description:
An employer is seeking a Principal Kafka Engineer to lead the management, scaling, and optimization of our enterprise event-streaming ecosystem. The ideal candidate brings deep expertise in Apache/Confluent Kafka, applies object-oriented programming principles to advanced automation, and uses ITSM practices to support reliable, high-quality service delivery.
Kafka Administration:
Deploy, configure, and maintain highly available Kafka clusters across on-premises and cloud environments, including Azure and GCP. Manage topics, partitions, and replication to ensure platform reliability. OOP-Based Development and Automation:
Design and build reusable frameworks, automation tools, and APIs using Java to simplify cluster provisioning, monitoring, and self-service onboarding. ITSM Integration:
Manage the platform lifecycle through ITSM processes, including incident, problem, and change management, to reduce downtime and maintain SLA compliance. Security and Compliance:
Implement and maintain security controls and authentication mechanisms. Performance Optimization:
Monitor cluster health with tools such as dyna Trace, Prometheus, and Grafana. Perform capacity planning and tuning to support high-throughput data pipelines.
Collaboration:
Partner with Dev Ops and application teams to provide guidance on event-driven architecture and producer/consumer performance best practices.
30% of time spent on Kafka Platform Administrative work 30% of time spent on Infrastructure or Operational Concerns which potentially includes Terraform, Kubernetes, Security, Storage, Networking etc 30% of time spent on Software Development type of work 10% of the time spent resolving customer requests (understanding the service requests submitted and helping them to be completed)
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).