Apache Spark vs Cloudera Distribution for Hadoop vs Outerthought Lily comparison

Apache and Cloudera are both solutions in the Hadoop category. Apache is ranked #1 with an average rating of 8.7, while Cloudera is ranked #2 with an average rating of 8.4. Apache holds a 13.4% mindshare in H, compared to Cloudera’s 14.0% mindshare. Additionally, 90% of Apache users are willing to recommend the solution, compared to 92% of Cloudera users who would recommend it.

Apache Spark

Read 68 Apache Spark reviews

3,663 Views
879 Comparison Views

90% willing to recommend

Cloudera Distribution for H...

Read 51 Cloudera Distribution for Hadoop reviews

2,002 Views
1,181 Comparison Views

92% willing to recommend

Outerthought Lily

125 Views
123 Comparison Views

Apache Spark

Cloudera Distribution for H...

Outerthought Lily

Comparison Buyer's Guide

Download the report

Executive Summary

We performed a comparison between Apache Spark, Cloudera Distribution for Hadoop, and Outerthought Lily based on real PeerSpot user reviews.

Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop.

To learn more, read our detailed Hadoop Report (Updated: January 2026).

Buyer's Guide

Hadoop

January 2026

Download the complete report

Helped 881,665 peers since 2012

Review summaries and opinions

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:

Mindshare comparison

As of February 2026, in the Hadoop category, the mindshare of Apache Spark is 13.4%, down from 18.4% compared to the previous year. The mindshare of Cloudera Distribution for Hadoop is 14.0%, down from 27.4% compared to the previous year. The mindshare of Outerthought Lily is 2.7%, up from 1.3% compared to the previous year. It is calculated based on PeerSpot user engagement data.

Hadoop Market Share Distribution
Product	Market Share (%)
Apache Spark	13.4%
Cloudera Distribution for Hadoop	14.0%
Outerthought Lily	2.7%
Other	69.9%

Hadoop

Featured Reviews

Devindra Weerasooriya

Data Architect at Devtech

Provides a consistent framework for building data integration and access solutions with reliable performance

The in-memory computation feature is certainly helpful for my processing tasks. It is helpful because while using structures that could be held in memory rather than stored during the period of computation, I go for the in-memory option, though there are limitations related to holding it in memory that need to be addressed, but I have a preference for in-memory computation. The solution is beneficial in that it provides a base-level long-held understanding of the framework that is not variant day by day, which is very helpful in my prototyping activity as an architect trying to assess Apache Spark, Great Expectations, and Vault-based solutions versus those proposed by clients like TIBCO or Informatica.

Read full review

Rok Dolinsek

Manager, Bussines Development & Co Owner at Troia d.o.o.

Enables on-premise implementation with powerful data processing capabilities

This is the only solution that is possible to install on-premise. Cloudera provides a hybrid solution that combines compute on cloud or on-premises. It includes all machine learning algorithms in the Spark machine learning library. All functionalities needed for a big data platform and ETL are on the platform, eliminating the need for other tools. It is scalable, ready for vertical scaling, and very powerful, offering numerous functionalities and configurations for generative AI.

Read full review

Use Outerthought Lily?

Leave a review

See which vendors are best for you

Use our free recommendation engine to learn which Hadoop solutions are best for your needs.

See recommendations

881,665 professionals have used our research since 2012.

Top Industries

By visitors reading reviews

Financial Services Firm

25%

Computer Software Company

Manufacturing Company

University

Financial Services Firm

21%

Marketing Services Firm

Computer Software Company

Comms Service Provider

No data available

Company Size

By reviewers

Large Enterprise

Midsize Enterprise

Small Business

By reviewers
Company Size	Count
Small Business	28
Midsize Enterprise	15
Large Enterprise	32

By reviewers
Company Size	Count
Small Business	16
Midsize Enterprise	9
Large Enterprise	31

No data available

Questions from the Community

What do you like most about Apache Spark?

We use Spark to process data from different data sources.

See all answers

What is your experience regarding pricing and costs for Apache Spark?

Apache Spark is open-source, so it doesn't incur any charges.

See all answers

What needs improvement with Apache Spark?

Areas for improvement are obviously ease of use considerations, though there are limitations in doing that, so while ...

See all answers

What do you like most about Cloudera Distribution for Hadoop?

The tool can be deployed using different container technologies, which makes it very scalable.

See all answers

What is your experience regarding pricing and costs for Cloudera Distribution for Hadoop?

The price for Cloudera is average, yet it is very good compared to other solutions. It can be deployed on-premises, u...

See all answers

What needs improvement with Cloudera Distribution for Hadoop?

If they could support modifying the data more easily than the current implementation, it would be beneficial.

See all answers

Ask a question

Earn 20 points

Comparisons

Spring Boot vs Apache Spark

Compared 15% of the time

AWS Lambda vs Apache Spark

Compared 7% of the time

AWS Batch vs Apache Spark

Compared 6% of the time

Apache NiFi vs Apache Spark

Compared 6% of the time

HPE Data Fabric vs Apache Spark

Compared 4% of the time

More Apache Spark Competitors

HPE Data Fabric vs Cloudera Distribution for Hadoop

Compared 12% of the time

MongoDB Enterprise Advanced vs Cloudera Distribution for Hadoop

Compared 10% of the time

Amazon EMR vs Cloudera Distribution for Hadoop

Compared 9% of the time

Couchbase Enterprise vs Cloudera Distribution for Hadoop

Compared 9% of the time

OpenText Analytics Database (Vertica) vs Cloudera Distribution for Hadoop

Compared 3% of the time

More Cloudera Distribution for Hadoop Competitors

Argyle Data vs Outerthought Lily

Compared 36% of the time

More Outerthought Lily Competitors

Product Reports

Buyer's Guide

Apache Spark

January 2026

Download Apache Spark product report

Buyer's Guide

Cloudera Distribution for Hadoop

January 2026

Download Cloudera Distribution for Hadoop product report

Buyer's Guide

Hadoop

January 2026

Download Outerthought Lily product report

Also Known As

No data available

Lily

Overview

Spark provides programmers with an application programming interface centered on a data structure called the resilient distributed dataset (RDD), a read-only multiset of data items distributed over a cluster of machines, that is maintained in a fault-tolerant way. It was developed in response to limitations in the MapReduce cluster computing paradigm, which forces a particular linear dataflowstructure on distributed programs: MapReduce programs read input data from disk, map a function across the data, reduce the results of the map, and store reduction results on disk. Spark's RDDs function as a working set for distributed programs that offers a (deliberately) restricted form of distributed shared memory

Apache

Cloudera Distribution for Hadoop is the world's most complete, tested, and popular distribution of Apache Hadoop and related projects. CDH is 100% Apache-licensed open source and is the only Hadoop solution to offer unified batch processing, interactive SQL, and interactive search, and role-based access controls. More enterprises have downloaded CDH than all other such distributions combined.

Cloudera

Lily Enterprise sits on top of the Cloudera or Hortonworks Hadoop platforms, and considers the entire Hadoop stack and all data streams, including data warehouses, data lakes, CRM systems, mobile/online or social media activities, contact center applications and POS systems. Lily combines them into one cumulative, organized and actionable solution, and delivers real-time execution of predictive analytical models to help create individual customer profiles, Lily Customer DNA, to give organizations the most accurate, atomic-level view of their customers.

Outerthought

Sample Customers

NASA JPL, UC Berkeley AMPLab, Amazon, eBay, Yahoo!, UC Santa Cruz, TripAdvisor, Taboola, Agile Lab, Art.com, Baidu, Alibaba Taobao, EURECOM, Hitachi Solutions

37signals, Adconion,adgooroo, Aggregate Knowledge, AMD, Apollo Group, Blackberry, Box, BT, CSC

ING, Orange, France Telecom, Alpha Credit, Turkcell, Eni, Zain Group, AXA, Rogers, Toyota, Belfius

Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop. Updated: January 2026.

DOWNLOAD NOW

881,665 professionals have used our research since 2012.