Try our new research platform with insights from 80,000+ expert users

Amazon EMR vs Cloudera Data Platform vs HPE Ezmeral Data Fabric comparison

 

Comparison Buyer's Guide

Executive Summary

Review summaries and opinions

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
 

Featured Reviews

Prashant  Singh - PeerSpot reviewer
Seamless data integration enhances reporting efficiency and an easy setup
Amazon EMR has multiple connectors that can connect to various data sources. The service charges are based on processing only, depending on the resources used, which can help save money. It is easy to integrate with other services for storage, allowing data to be shifted to cheaper storage based on usage.
Miodrag-Stanic - PeerSpot reviewer
Distributed computing improves data processing while upgrade complexity needs addressing
There are challenges with upgrading or updating various services like Spark, Impala, and Hive on on-premise and bare metal solutions. We aim to address these issues with a Kubernetes-based platform that will simplify the task of upgrading services. We also wish to implement lakehouse capabilities with Iceberg or Delta Lake frameworks.
Hamid M. Hamid - PeerSpot reviewer
A stable and scalable tool that serves as a great database
The initial setup of HPE Ezmeral Data Fabric is easy. I am not sure how long it took to deploy HPE Ezmeral Data Fabric, but I haven't heard about any disadvantages when it comes to the time taken for the deployment. I remember that one of our company's clients who had purchased the product never mentioned the product's setup phase being complex. One of the drawbacks with HPE Ezmeral Data Fabric stems from the fact that the product's upgrade was not straightforward, and it was a complex process since one of the teams in my company who deals with the tool found the upgrade part to be tough. The solution is deployed on an on-premises model. My company has two dedicated staff members to look after the deployment and maintenance phases of HPE Ezmeral Data Fabric.

Quotes from Members

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
 

Pros

"Amazon EMR is a good solution that can be used to manage big data."
"One of the valuable features about this solution is that it's managed services, so it's pretty stable, and scalable as much as you wish. It has all the necessary distributions. With some additional work, it's also possible to change to a Spark version with the latest version of EMR. It also has Hudi, so we are leveraging Apache Hudi on EMR for change data capture, so then it comes out-of-the-box in EMR."
"This is the best tool for hosts and it's really flexible and scalable."
"The solution is scalable."
"It allows users to access the data through a web interface."
"The initial setup is straightforward."
"I rate Amazon EMR as ten out of ten."
"In Amazon EMR it is easy to rebuild anything, easy to upgrade and has good fault tolerance."
"Cloudera Data Platform has significantly improved our data management."
"It is a scalable platform."
"The product offers a fairly easy setup process."
"Ranger for security; with Ranger we can manager user’s permissions/access controls very easily."
"The data platform is pretty neat. The workflow is also really good."
"The upgrades and patches must come from Hortonworks."
"Ambari Web UI: user-friendly."
"Integration with other tools works well for us and we successfully scaled the solution after two to three years without any issues."
"I like the administration part."
"My customers find the product cheaper compared to other solutions. The previous solution that we used did not have unified analytics like the runtime or the analog."
"It is a stable solution...It is a scalable solution."
"HPE Ezmeral Data Fabric can be accessed from any namespace globally as you would access it from a machine using an NFS."
"The model creation was very interesting, especially with the libraries provided by the platform."
 

Cons

"Amazon EMR is continuously improving, but maybe something like CI/CD out-of-the-box or integration with Prometheus Grafana."
"The problem for us is it starts very slow."
"There were times where they would release new versions and it seemed to end up breaking old versions, which is very strange."
"The initial setup was time-consuming."
"The most complicated thing is configuring to the cluster and ensure it's running correctly."
"The product must add some of the latest technologies to provide more flexibility to the users."
"As people are shifting from legacy solutions to other technologies, Amazon EMR needs to add more features that give more flexibility in managing user data."
"There is no need to pay extra for third-party software."
"For on-premise use, I would not recommend Cloudera Data Platform as it is expensive and complicated to upgrade."
"Security and workload management need improvement."
"I would like to see more support for containers such as Docker and OpenShift."
"It's at end of life and no longer will there be improvements."
"It would also be nice if there were less coding involved."
"Hive performance. If Hive performance increased, Hadoop would replace (not everywhere) traditional databases."
"Deleting any service requires a lot of clean up, unlike Cloudera."
"I work a lot with banking, IT and communications customers. Hortonworks must improve or must upgrade their services for these sectors."
"The product is not user-friendly."
"Upgrading Ezmeral to a new version is a pain. They're trying to make the solution more container-friendly, so I think they're going in the right direction. The only problem we've had in the past was the upgrades. The process isn't smooth due to how the Red Hat operating system upgrades currently work."
"The deployment could be faster. I want more support for the data lake in the next release."
"Having the ability to extend the services provided by the platform to an API architecture, a micro-services architecture, could be very helpful."
"HPE Ezmeral Data Fabric is not compatible with third-party tools."
 

Pricing and Cost Advice

"There is a small fee for the EMR system, but major cost components are the underlying infrastructure resources which we actually use."
"The product is not cheap, but it is not expensive."
"The cost of Amazon EMR is very high."
"I rate the tool's pricing a five out of ten. It can be expensive since it's a managed service, and if you are not careful, you can run into unexpected charges. You can make a mistake that costs you tens of thousands of dollars. That's happened to us twice, so I'm sensitive to it. We're still trying to work on that. Our smallest client probably spends a hundred thousand dollars yearly on licensing, while our largest is well over a million."
"There is no need to pay extra for third-party software."
"Amazon EMR is not very expensive."
"Amazon EMR's price is reasonable."
"The price of the solution is expensive."
"Currently, we are using the product in a sandbox environment, and there is no licensing. We might choose a licensing option once we get the results."
"It is priced well and it is affordable"
"The tool's price is cheap and based on a usage basis. The solution's licensing costs are yearly and there are no extra costs."
"HPE is flexible with you if you are an existing customer. They offer different models that might be beneficial for your organization. It all depends on how you negotiate."
"There is a need for my company to pay for the licensing costs of the solution."
report
Use our free recommendation engine to learn which Hadoop solutions are best for your needs.
859,957 professionals have used our research since 2012.
 

Top Industries

By visitors reading reviews
Financial Services Firm
25%
Computer Software Company
13%
Educational Organization
10%
Manufacturing Company
8%
Financial Services Firm
14%
Real Estate/Law Firm
12%
Computer Software Company
10%
Performing Arts
8%
Financial Services Firm
20%
Computer Software Company
14%
Comms Service Provider
11%
Healthcare Company
6%
 

Company Size

By reviewers
Large Enterprise
Midsize Enterprise
Small Business
 

Questions from the Community

What do you like most about Amazon EMR?
Amazon EMR is a good solution that can be used to manage big data.
What is your experience regarding pricing and costs for Amazon EMR?
Compared to others, Amazon seems efficient and is considered good for Big Data workloads. Costs are involved based on...
What needs improvement with Amazon EMR?
There is room for improvement with respect to retries, handling the volume of data on S3 ( /products/amazon-s3-review...
What do you like most about Hortonworks Data Platform?
Distributed computing, secure containerization, and governance capabilities are the most valuable features.
What is your experience regarding pricing and costs for Hortonworks Data Platform?
The pricing model for Cloudera Data Platform is complex and has increased significantly compared to CDH. Initially, C...
What needs improvement with Hortonworks Data Platform?
Cloudera Data Platform should include additional capabilities and features similar to those offered by other data man...
What do you like most about HPE Ezmeral Data Fabric?
It is a stable solution...It is a scalable solution.
What needs improvement with HPE Ezmeral Data Fabric?
There are some drawbacks in HPE Ezmeral Data Fabric when it comes to the interoperability part. HPE Ezmeral Data Fabr...
What is your primary use case for HPE Ezmeral Data Fabric?
The main purpose of HPE Ezmeral Data Fabric for me is that it acts as a database. In my company, we store our data wi...
 

Also Known As

Amazon Elastic MapReduce
No data available
MapR, MapR Data Platform
 

Overview

 

Sample Customers

Yelp
Information Not Available
Valence Health, Goodgame Studios, Pico, Terbium Labs, sovrn, Harte Hanks, Quantium, Razorsight, Novartis, Experian, Dentsu ix, Pontis Transitions, DataSong, Return Path, RAPP, HP
Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop. Updated: June 2025.
859,957 professionals have used our research since 2012.