Python/Spark/AI Developer
Listed on 2026-09-04
-
Software Development
Python, SQL Developer, Data Engineering
Location: Hybrid / Remote
Employment Type: Full-Time
Industry: Data Engineering | AI | Software Development | Data Platforms
Waters Edge Solutions is partnering with a number of leading South African and international organisations to recruit an experienced Python/Spark/AI Developer to join a technically driven team working on the modernisation of large-scale data platforms.
This role will focus on replatforming legacy T-SQL workloads into modern Spark and Delta Lake pipelines
, while building the APIs and AI capabilities that sit around them. You'll work primarily in Python, with a strong emphasis on type-safe, tested and spec-driven development.
This is an excellent opportunity for an intermediate-to-senior developer who enjoys solving complex data engineering problems and wants to work across Python, distributed data processing, lakehouse architecture, APIs and emerging AI technologies
.
As Python/Spark/AI Developer, you'll play a key role in migrating legacy data workloads onto a modern Spark-based architecture.
You'll translate existing T-SQL logic into Spark SQL and PySpark pipelines, validate migrated workloads against legacy outputs, and ensure the new platform delivers reliable and demonstrable parity.
The engineering environment is heavily spec-driven. You'll work autonomously from written requirements, design documentation and Architecture Decision Records (ADRs), with changes expected to include appropriate testing and documentation rather than code alone.
Alongside the core data engineering work, you'll build REST APIs around pipeline jobs and contribute to AI capabilities, including provider-neutral LLM integrations and structured AI outputs.
Key ResponsibilitiesBuild production-grade Spark and PySpark pipelines on Delta Lake.
Replatform legacy T-SQL workloads and stored-procedure-era logic into Spark SQL.
Translate existing data logic accurately while maintaining expected business and data semantics.
Write modern, type-hinted Python with strict static analysis.
Develop and maintain automated pytest suites covering unit, integration and end-to-end testing.
Validate migrated pipelines against legacy systems and provide clear parity and regression evidence.
Build REST APIs using FastAPI or equivalent frameworks.
Implement API request validation, authentication, idempotency and job-status semantics.
Work within Docker-based development environments using Compose stacks.
Follow Git Hub flow and PR-driven development practices.
Maintain linting, type-checking and automated test gates within CI pipelines.
Work from design documents, technical specifications and ADRs.
Document technical changes and decisions as part of the development process.
Contribute to AI and LLM-enabled functionality around data workflows.
Work independently within a hybrid or remote engineering environment.
4+ years of professional Python development experience.
2+ years of production experience building Spark or comparable distributed data pipelines.
Strong modern Python development skills, ideally using Python 3.12
.Experience writing type-hinted Python that passes strict static analysis using tools such as pyright or mypy.
Experience with Pydantic models and structured data validation.
Understanding of ABC-based provider patterns.
Experience with modern Python packaging tools such as uv or Poetry.
Familiarity with Click or similar CLI frameworks.
Strong production experience with Apache Spark and Py Spark .
Experience working with open-source Spark environments rather than solely managed vendor platforms.
Strong knowledge of the Data Frame API and Spark SQL.
Understanding of Spark partitioning, performance tuning and driver/executor architecture.
Familiarity with Spark Connect.
Experience with
Delta Lake or an equivalent lakehouse table format such as Apache Iceberg or Hudi.Understanding of MERGE INTO, schema evolution, time travel and idempotent write patterns.
Strong, dialect-portable SQL skills.
Ability to understand legacy T-SQL and faithfully reproduce its behaviour in Spark SQL.
Strong understanding of hashing, surrogate keys, deduplication and set-based data logic.
Strong automated testing experience using pytes…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: