×
Register Here to Apply for Jobs or Post Jobs. X

Technology Lead - PySpark Developer

Job in Mississauga, Ontario, Canada
Listing for: Infosys Limited
Full Time position
Listed on 2026-09-03
Job specializations:
  • Software Development
    Data Engineering, SQL Developer
Job Description & How to Apply Below
Technology|Big Data - Data Processing|Spark, Technology|Functional Programming|Scala

Domain

Delivery

Interest Group

Company

ITL Canada

Requisition

152697

Infosys is seeking an experienced PySpark Developer  to design, develop, and optimize scalable bigdata solutions. The candidate will work on building high-performance batch and real-time data pipelines leveraging the Hadoop ecosystem and distributed computing frameworks(Spark). The role involves working closely with data engineers, architects, and business stakeholders to deliver robust, scalable, and efficient data processing systems.

Required Qualifications

Candidates authorized to work for any employer in Canada without employer-based visa sponsorship are welcome to apply. Infosys is unable to provide immigration sponsorship for this role at this time.

Candidate must be located within commuting distance of  Mississauga, Ontario  or be willing to relocate to the areas.

Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.

At least 4 years of Information Technology experience

At least 4 years of experience in Big Data technologies.

Strong expertise in:

Apache Spark (Core, SQL, Data Frames, RDDs)

Scala programming

Py Spark

Hands-on experience with:

Kafka (real-time streaming)

Hadoop ecosystem (HDFS, Hive, Impala)

No

SQL Databases (HBase, MongoDB, Couchbase)

Strong understanding of distributed computing concepts and data processing frameworks.

Experience in building ETL/data pipelines for large-scale datasets.

Proficiency in SQL and data modeling.

Preferred Qualifications

Hands-on experience with data lakes, data warehouses, and scalable ETL pipeline design, including batch and real-time processing architecture.

Strong understanding and practical exposure to Agile software development methodologies (Scrum) and SDLC practices.

Proven experience in Banking domain, supporting use cases such as fraud detection, risk analytics, regulatory reporting, and customer insights.

Excellent analytical, problem-solving, and communication skills, with the ability to translate business requirements into scalable technical solutions.

Demonstrated ability to work effectively in cross-functional, multi-stakeholder environments, collaborating with Business, Data Engineering, and Architecture teams.

Experience with real-time data streaming frameworks such as Kafka and Spark Streaming for low-latency processing.

Understanding data modeling concepts (dimensional modeling, snowflake schemas) to support analytics workloads.

Experience and desire to work in a global delivery environment.

Key Responsibilities

Design and develop large-scale data processing pipelines using Apache Spark (Scala & PySpark)

Build and optimize batch and real-time data processing workflows using Spark, Kafka, and Hadoop ecosystem

Develop Spark applications using RDDs, Data Frames, and Spark SQL for complex transformations

Develop and optimize PySpark applications leveraging joins, Spark DAG execution flow, stage optimization, transformation techniques, and streaming with dynamic allocation and failover handling.

Implement streaming pipelines using Kafka and Spark Streaming / Structured Streaming

Develop and maintain HDFS, Hive, No

Sql and Impala-based data lake solutions

Convert existing SQL/Hive workloads into optimized Spark jobs for improved performance

Work with ETL pipelines to ingest, cleanse, transform, and process large datasets

Optimize performance through partitioning, caching, serialization, and tuning techniques

Handle data formats such as Parquet, ORC, Avro, JSON

Integrate multiple data sources including streaming systems, flat files RDBMS, and APIs

Collaborate with cross-functional teams to understand business requirements and translate them into scalable technical solutions

Ensure data quality, reliability, and performance monitoring across pipelines

Participate in code reviews, design discussions, and best practices implementation

Key Skills

Distributed Data Processing.

Spark Optimization & Performance Tuning.

Real-time Data Streaming.

Data Modeling & ETL Design.

Problem-solving and…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary