Oracle Senior Python and Kafka Engineer
Listed on 2026-09-07
-
Software Development
Python, Backend Developer
Oracle Senior Python and Kafka Engineer
We are seeking a hands-on Senior Python and Kafka Engineer to join an established engineering team supporting a correctness-critical, event-driven digital repository processing platform. The platform performs downstream processing after digital objects have been deposited and archived. Its services create persistent identifiers, update operational metadata and search stores, generate viewable and streamable derivatives, distribute content to image, media and full-text services, and notify depositors.
The pipeline is event-driven, with an initial archival event being processed by multiple independent services that publish and consume subsequent events. The successful candidate will work directly within the existing team to improve structured logging, end-to-end event visibility, reconciliation, completion verification, and Kafka delivery reliability. This is an implementation-focused engineering position. The architecture is owned internally, and the selected engineer will build enhancements into a live system rather than redesigning or replacing it.
Key Responsibilities:
- Implement a consistent structured JSON logging standard across Python services, including standardized fields, logging levels, contextual information, and actionable failure details.
- Improve Kafka producer and consumer reliability, with particular attention to: offset commit behavior, consumer-group stability and rebalancing, retry processing, dead-letter topic handling, message-loss investigation, duplicate-processing and idempotency considerations.
- Instrument distributed services to provide greater visibility into how individual digital objects move through the processing pipeline.
- Implement or contribute to an event-lineage capability that records: the processing stages an object passed through, the order in which stages occurred, the outcome of each stage, where processing stopped or failed.
- Develop reconciliation and completion-tracking capabilities that compare expected workflow stages with recorded events.
- Help produce an actionable completion verdict for processed objects, including the ability to identify incomplete processing and support alerting or delivery holds.
- Implement telemetry using Open Telemetry and contribute to an Open Lineage-based lineage solution, with Marquez currently identified as the leading implementation option.
- Develop and maintain Python services that interact with Kafka, PostgreSQL, MongoDB, object storage, and external APIs.
- Work with shared internal Python packages that provide message-envelope models, Kafka producer and consumer wrappers, retry and dead-letter handling, shared-state storage, and operational-metadata access.
- Write automated tests and support safe delivery through the existing Kubernetes, Kustomize, and ArgoCD environment.
- Troubleshoot complex distributed-processing issues across services, messages, data stores, and external dependencies.
- Pair closely with the internal engineer leading this work and contribute directly to implementation, testing, technical documentation, and operational readiness.
The principal areas of work are structured logging, event lineage and tracking, reconciliation and completion verification, and Kafka delivery semantics.
Required Qualifications:
- Apache Kafka:
Significant hands-on experience operating or developing Kafka-based production applications. Deep understanding of producer and consumer delivery semantics. Experience with Kafka offsets, commit strategies, consumer groups, partition assignment, and rebalancing. Demonstrated ability to diagnose missing, delayed, duplicated, or repeatedly retried messages. Practical experience implementing retry and dead-letter processing patterns. - Python:
Advanced Python software-engineering experience. Strong understanding of concurrency, threading, and runtime behavior. Experience building and supporting production services and shared Python libraries. Ability to write maintainable, testable, and operationally supportable code. - Observability and Structured Logging:
Experience defining and implementing structured logging standards, not only configuring logging products. Ability to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).