Apache Spark Developer
Listed on 2026-09-25
-
Software Development
Data Engineering
Apache Spark Developer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Apache Spark Developer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $125,000–$185,000 Annually
Experience
Required:
6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job SummaryWe are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines.
You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems.
Key Responsibilities- Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.
- Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
- Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
- Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
- Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
- Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
- Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
- Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
- Integrate Spark applications with enterprise data warehouses, lake houses, and reporting platforms.
- Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
- Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
- Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures.
- Six or more years of professional software or data engineering experience.
- Four or more years of hands-on Apache Spark development experience in enterprise production environments.
- Strong proficiency in Py Spark ,
Scala
, or Spark SQL for distributed data processing. - Deep understanding of Apache Spark architecture including RDDs, Data Frames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
- Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
- Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
- Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
- Experience processing streaming data using Spark…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).