Softwr
Apache Hadoop logo

Apache Hadoop

The original open source framework for distributed storage and batch processing on commodity servers, now largely a legacy platform.

As of 30 August 2026, Apache Hadoop is free to use. Hadoop provides HDFS for distributed storage and YARN for cluster scheduling, and still runs large on-premises estates. Softwr lists it under Technology.

Overview

What Apache Hadoop does

Apache Hadoop is an Apache Software Foundation project licensed under Apache 2.0, made up of four parts: HDFS, a distributed filesystem that replicates blocks across nodes; YARN, a cluster resource manager and scheduler; MapReduce, the original batch processing model; and Hadoop Common, the shared libraries. It is written in Java and designed to run on racks of commodity servers with directly attached disks, tolerating individual machine failures through replication. A large surrounding ecosystem grew around it, including Hive, HBase, Ranger, Oozie and Ozone, and Spark and Flink both still run under YARN. The distinguishing property today is not technical, it is economic and it runs the wrong way. Hadoop's original argument was that storage and compute on your own commodity hardware were far cheaper than a proprietary appliance. Object storage inverted that by letting storage and compute scale and be paid for separately, which is why new data platforms start on S3 or its equivalents rather than on HDFS. The commercial consequence is severe: Cloudera and Hortonworks merged, MapR was sold, and the free distributions CDH and HDP reached end of support, so the packaged, patched, freely available Hadoop that most organisations actually ran no longer exists. What remains is a subscription product from Cloudera or assembling Apache tarballs yourself. Who still runs it are organisations with an existing multi-petabyte on-premises estate, often for data residency, egress cost or regulatory reasons, where HDFS at rack scale is genuinely cheaper than cloud object storage. The trade-off is staffing and direction. Operating it is a specialism covering Kerberos, YARN queues, JVM tuning and a stack of interdependent project versions, the people who hold that skill have been moving to cloud data platforms for a decade, and MapReduce itself is maintained for compatibility rather than developed, so new work written against it is written against a dead end.

What people use it for

  • Operating an existing multi-petabyte on-premises estate where data residency or egress costs rule out moving to cloud object storage
  • Running Spark or Flink under YARN on hardware you already own, using HDFS as the storage layer
  • Keeping long-lived regulated archives on infrastructure entirely within your own data centres and legal jurisdiction
  • Maintaining legacy Hive and MapReduce workloads during a staged migration to a lakehouse or cloud platform

The honest half

Where it falls short

Concrete and checkable, so you can decide whether any of them matter to you. This is the half of a review a vendor will not write about Apache Hadoop.

  • The free vendor distributions no longer exist: Cloudera's CDH and Hortonworks' HDP have reached end of support and the successor CDP is subscription-only, so running Hadoop without paying now means assembling, testing and security-patching Apache releases yourself.
  • HDFS couples storage to compute, so adding capacity means buying whole nodes with CPU and memory you may not need, and the entire industry moved to object storage precisely because it lets the two be bought separately.
  • The NameNode holds all filesystem metadata in memory, so a cluster with tens of millions of small files exhausts heap long before it exhausts disk, and the remedy is a file compaction job that somebody has to write, schedule and own indefinitely.
  • Operating it is a distinct specialism covering Kerberos, YARN queue tuning, JVM garbage collection and the compatibility matrix between Hive, HBase, Ranger, Oozie and the core, and an upgrade touches all of them at once rather than one at a time.
  • MapReduce is maintained for compatibility rather than actively developed, and new work goes to Spark or Flink, so a job written against MapReduce today is written against an API that will not gain anything further.
  • Hiring is against you: the talent pool has moved to cloud data platforms over the past decade, so a Hadoop estate increasingly depends on a small number of individuals, which makes it a succession risk before it is a technical one.

Cross-shopped

What people choose instead of Apache Hadoop

Each pairing was judged by two reviewers asking whether a buyer would genuinely weigh the two against each other. The ones that failed were deleted rather than published.

Capabilities

Features

  • HDFS

    Distributed filesystem that splits files into blocks and replicates them across nodes, with erasure coding as a lower-overhead alternative

  • YARN

    Cluster resource manager that schedules containers across nodes, with queues, capacity guarantees and multi-tenancy

  • MapReduce

    The original batch processing model, retained for compatibility with existing jobs

  • HDFS federation and high availability

    Multiple namespaces and standby NameNodes with automatic failover to remove the single metadata point of failure

  • Kerberos security

    Cluster-wide authentication for users and services, with Ranger commonly layered on for authorisation and audit

  • Rack awareness

    Places replicas across racks so a rack failure does not take out every copy of a block

  • S3A and object store connectors

    Reads and writes cloud object storage through the Hadoop filesystem API, which is how many non-Hadoop tools inherited S3 support

  • Ecosystem compatibility

    Spark, Flink, Hive and HBase all run on YARN and read HDFS, so the cluster can host engines newer than MapReduce

Answered, with sources

Questions people ask

Each answer names the page it came from, so you can check it rather than take our word for it.

Is Hadoop dead?

No, but it is legacy. Large on-premises HDFS estates still run and are still supported, and Spark and Flink still run on YARN. What has ended is Hadoop as a default choice for new platforms, which now start on object storage.

Can I still get a free packaged distribution?

Not a maintained one. CDH and HDP reached end of support and Cloudera's CDP is a paid subscription. The remaining free route is building and patching Apache releases yourself, which is a real engineering commitment.

Do I need Hadoop to run Spark?

No. Spark runs standalone, on Kubernetes and on managed cloud services, and reads object storage directly. Many Spark deployments include Hadoop client libraries for the filesystem connectors without running a Hadoop cluster at all.

What replaced HDFS?

Object storage, typically S3 or a compatible system, combined with an open table format such as Apache Iceberg or Delta Lake. Apache Ozone exists as an object store within the Hadoop ecosystem for organisations staying on-premises.

Is it cheaper than the cloud?

It can be at multi-petabyte scale with steady, predictable utilisation, particularly where egress charges would be large. Include the staffing cost honestly, because the specialist operators a Hadoop cluster requires are scarce and therefore expensive.

Share

Keep looking

Where to go from Apache Hadoop

Best Technology software for

Compare Apache Hadoop with

Other Technology software

  • Manage your team's work, projects, & tasks online

    Free, then $10.99/mo11 researched notes
  • One app to replace them all

    Free, then $7/user/month (annual)13 researched notes
  • The collaborative interface design tool

    Free, then $12/mo15 researched notes
  • The issue tracking tool you'll enjoy using

    Free, then $10/mo11 researched notes
  • A distributed engine for batch, SQL, streaming and machine learning workloads over data that does not fit on one machine.

    Free plan15 researched notes
  • The open source Firebase alternative

    Free, then $25/mo10 researched notes
  • Cloud native core banking where products are written as smart contracts

    Pricing on request11 researched notes
  • A distributed SQL engine that queries data where it already lives, across object storage, warehouses and operational databases.

    Free plan15 researched notes
  • Build component driven UIs faster

    Free plan10 researched notes
  • Production-grade container orchestration

    Free plan10 researched notes
  • The single platform to analyze, test, observe, and deploy new features

    Free plan13 researched notes
  • The real-time data platform

    Free plan7 researched notes
  • Application monitoring platform built by developers for developers

    Free, then $26/mo15 researched notes
  • Roadmapping software for product builders

    From $59/mo14 researched notes
  • Capture, organize, and analyze product feedback

    Free, then $400/mo15 researched notes
  • The inside sales CRM of choice for startups & SMBs

    From $9/mo15 researched notes
  • The doc that brings it all together

    Free, then $10/mo12 researched notes
  • A distributed SQL engine that queries data where it already lives, across object storage, warehouses and operational databases.

    Free plan15 researched notes

Softwr does not host reviews and shows no star rating for Apache Hadoop, because a rating we did not collect is not ours to publish. What is here is the pricing and platform detail from the vendor’s own pages, limitations we could state concretely, and alternatives a reviewer confirmed people weigh against it. Tell us if any of it is wrong.

More on Apache Hadoop

Best Technology software alternatives

Cloud native credit card processing and core banking from Bhavin Turakhia's Zeta

quote

Cloud native core banking where products are written as smart contracts

quote

Data driven personalisation and money insights inside a bank's existing app

quote

Cloud native core banking, sold as Finxact from Fiserv since the 2022 acquisition

quote

Digital banking platform for United States banks and credit unions

quote

Secure access for everyone

The digital analytics platform to understand your users

Apple's WebKit browser, available only on Apple operating systems and updated only with them.

free

Compare Apache Hadoop with alternatives