Softwr
Apache Spark logo

Apache Spark

A distributed engine for batch, SQL, streaming and machine learning workloads over data that does not fit on one machine.

As of 30 August 2026, Apache Spark is free to use. Spark is the default engine for large-scale data processing on the JVM, with APIs in SQL, Python, Scala, Java and R. Softwr lists it under Technology.

Overview

What Apache Spark does

Apache Spark is a distributed data processing engine hosted by the Apache Software Foundation and licensed under Apache 2.0. It runs on the JVM and exposes the same execution engine through Spark SQL, DataFrames, Structured Streaming and MLlib, with language bindings for Scala, Java, Python, R and SQL. It runs on Kubernetes, YARN and its own standalone scheduler, and reads from object storage, HDFS, JDBC sources and table formats including Delta Lake, Apache Iceberg and Apache Hudi. Spark 4.0, released in 2025, made ANSI SQL mode the default and expanded Spark Connect, which separates the client from the cluster. What distinguishes it is that one engine covers batch ETL, interactive SQL, stream processing and model training with the same code and the same cluster. That consolidation is the commercial argument: an organisation does not have to run and staff separate systems for its nightly pipelines, its analytics queries and its feature engineering, and a data engineer who learns the DataFrame API can move across all three. It is also why Spark is the assumed engine on every managed data platform, from Databricks and AWS EMR to Google Dataproc and Microsoft Fabric, which makes it the skill most easily hired for in data engineering. It is bought by organisations with genuinely large datasets and a data platform team, usually as a managed service rather than a self-run cluster. The trade-off arrives in two forms. First, the failure modes are distributed ones: skewed shuffles, executor out-of-memory errors and serialisation costs, none of which are visible in the code and all of which require somebody who reads the Spark UI. Second, the fastest Spark is not the Spark you can download, because Databricks' Photon engine and comparable vendor accelerations are proprietary, so published performance figures often describe a fork you can only rent.

What people use it for

  • Nightly ETL over terabytes in object storage, where a single machine would take longer than the batch window allows
  • Building and maintaining a lakehouse on Iceberg or Delta Lake, where Spark handles both the writes and the compaction
  • Feature engineering and model training across datasets too large to fit in pandas on one node
  • Migrating legacy MapReduce or Hive workloads onto an engine that is still actively developed and widely supported by cloud vendors

The honest half

Where it falls short

Concrete and checkable, so you can decide whether any of them matter to you. This is the half of a review a vendor will not write about Apache Spark.

  • Running it well is JVM operations work: executor sizing, shuffle partition counts, off-heap memory and serialisation all have to be tuned, and the failures you actually get are out-of-memory errors and skewed shuffles rather than wrong answers, so you need somebody who can read the Spark UI or you will scale the cluster instead of fixing the query.
  • The fastest Spark is not open source. Databricks' Photon engine and comparable vendor accelerations are proprietary, so benchmark numbers quoted for Spark frequently describe a fork you can only rent, and moving off that vendor loses the performance you sized your pipelines around.
  • It is a distributed system with distributed overheads, and modern single-node tools such as DuckDB and Polars finish faster on datasets up to hundreds of gigabytes with no cluster to start, so a Spark job below that threshold is paying coordination cost for nothing.
  • Structured Streaming is micro-batch, which puts an end-to-end latency floor in the range of hundreds of milliseconds to seconds; workloads that need genuine per-event latency go to Flink instead, and discovering this after building on Spark means a rewrite.
  • Major upgrades deliberately break jobs: Spark 4.0 turns ANSI SQL mode on by default, so silent overflow and invalid casts that previously produced nulls now raise runtime errors, and a pipeline that worked for years can start failing purely on upgrade.
  • PySpark hides a process boundary, and Python UDFs serialise every row between the JVM and a Python worker; a direct translation of pandas code into PySpark UDFs can run an order of magnitude slower than the equivalent built-in expressions.

Cross-shopped

What people choose instead of Apache Spark

Each pairing was judged by two reviewers asking whether a buyer would genuinely weigh the two against each other. The ones that failed were deleted rather than published.

Capabilities

Features

  • Unified engine

    Batch, SQL, streaming and machine learning share one execution engine, one cluster and one API surface

  • Catalyst optimiser

    Rewrites and plans queries, with adaptive query execution that changes join strategies and partition counts at runtime

  • DataFrame and SQL APIs

    The same logical plan whether expressed as SQL or as DataFrame code in Python, Scala, Java or R

  • Structured Streaming

    Incremental micro-batch processing with exactly-once semantics against the same DataFrame API as batch code

  • Spark Connect

    A thin client protocol that decouples the application from the cluster, so clients no longer need a co-located JVM driver

  • Kubernetes and YARN support

    Runs as a native Kubernetes scheduler or under YARN, letting clusters be created per job rather than left running

  • Table format integration

    Reads and writes Delta Lake, Apache Iceberg and Apache Hudi tables with transactional semantics on object storage

  • MLlib

    Distributed implementations of common algorithms and pipeline abstractions for training over data larger than one machine

Answered, with sources

Questions people ask

Each answer names the page it came from, so you can check it rather than take our word for it.

When is Spark the wrong choice?

When your data fits comfortably on one machine. DuckDB or Polars will process hundreds of gigabytes on a single large node faster than a Spark cluster, without a scheduler, a driver or a shuffle. Spark earns its overhead when the data genuinely does not fit.

Is Spark the same on Databricks as the open source version?

No. Databricks runs its own runtime including the proprietary Photon engine and its own optimisations, so performance figures and some behaviours do not carry over to open source Spark on EMR, Dataproc or your own Kubernetes cluster.

Can I use Spark for real-time processing?

For near-real-time, yes, with Structured Streaming's micro-batch model, which lands in the sub-second to seconds range. For true per-event latency in the low milliseconds, Flink is the usual choice.

Does upgrading between major versions break things?

Yes, by design in some cases. Spark 4.0 makes ANSI SQL mode the default, which converts previously silent overflow and cast failures into runtime errors. Upgrades need a testing pass over production pipelines rather than a version bump.

Do I need to know Scala?

No. Python covers the vast majority of work and PySpark is the most common interface. Scala still helps when reading the source, writing custom data sources or diagnosing errors that surface as JVM stack traces.

Share

Keep looking

Where to go from Apache Spark

Best Technology software for

Compare Apache Spark with

Other Technology software

  • Manage your team's work, projects, & tasks online

    Free, then $10.99/mo11 researched notes
  • One app to replace them all

    Free, then $7/user/month (annual)13 researched notes
  • The collaborative interface design tool

    Free, then $12/mo15 researched notes
  • The issue tracking tool you'll enjoy using

    Free, then $10/mo11 researched notes
  • The original open source framework for distributed storage and batch processing on commodity servers, now largely a legacy platform.

    Free plan15 researched notes
  • The real-time data platform

    Free plan7 researched notes
  • Cloud native core banking where products are written as smart contracts

    Pricing on request11 researched notes
  • Enterprise-grade monitoring solution with advanced visualization and reporting

    Free plan11 researched notes
  • A platform built for a new way of working

    Free, then $9/mo12 researched notes
  • Apple's WebKit browser, available only on Apple operating systems and updated only with them.

    Free plan15 researched notes
  • The data streaming platform built on Apache Kafka, delivered as a fully managed cloud service

    Free plan12 researched notes
  • Microsoft's email and calendar client, in the middle of a transition from the classic Windows application to a web-based replacement.

    Free, then $1.99/mo15 researched notes
  • The Conversation Cloud

    From $2,500/mo11 researched notes
  • Replay what users do on your site to find bugs faster

    Free, then $99/mo12 researched notes
  • hyperextensible Vim-based text editor

    Free plan10 researched notes
  • Take back control of your time

    From $12/mo10 researched notes
  • Uptime monitoring for hobby and non-profit projects, up to enterprise

    Free plan11 researched notes

Softwr does not host reviews and shows no star rating for Apache Spark, because a rating we did not collect is not ours to publish. What is here is the pricing and platform detail from the vendor’s own pages, limitations we could state concretely, and alternatives a reviewer confirmed people weigh against it. Tell us if any of it is wrong.

More on Apache Spark

Best Technology software alternatives

Cloud native credit card processing and core banking from Bhavin Turakhia's Zeta

quote

Cloud native core banking where products are written as smart contracts

quote

Data driven personalisation and money insights inside a bank's existing app

quote

Cloud native core banking, sold as Finxact from Fiserv since the 2022 acquisition

quote

Digital banking platform for United States banks and credit unions

quote

Secure access for everyone

The digital analytics platform to understand your users

Apple's WebKit browser, available only on Apple operating systems and updated only with them.

free

Compare Apache Spark with alternatives