Apache Spark vs Cloudera Distribution for Hadoop vs HPE Ezmeral Data Fabric comparison

Cancel
You must select at least 2 products to compare!
Apache Logo
2,430 views|1,869 comparisons
89% willing to recommend
Cloudera Logo
2,881 views|2,224 comparisons
91% willing to recommend
Hewlett Packard Enterprise Logo
1,637 views|1,016 comparisons
100% willing to recommend
Comparison Buyer's Guide
Executive Summary

We performed a comparison between Apache Spark, Cloudera Distribution for Hadoop, and HPE Ezmeral Data Fabric based on real PeerSpot user reviews.

Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop.
To learn more, read our detailed Hadoop Report (Updated: April 2024).
769,599 professionals have used our research since 2012.
Featured Review
Quotes From Members
We asked business professionals to review the solutions they use.
Here are some excerpts of what they said:
Pros
"The most valuable feature of this solution is its capacity for processing large amounts of data.""The solution is scalable.""The most crucial feature for us is the streaming capability. It serves as a fundamental aspect that allows us to exert control over our operations.""With Hadoop-related technologies, we can distribute the workload with multiple commodity hardware.""The most valuable feature is the Fault Tolerance and easy binding with other processes like Machine Learning, graph analytics.""The solution has been very stable.""The deployment of the product is easy.""The most valuable feature of Apache Spark is its flexibility."

More Apache Spark Pros →

"The solution is reliable and stable, it fits our requirements.""It has the best proxy, security, and support features compared to open-source products.""The product provides better data processing features than other tools.""The scalability of Cloudera Distribution for Hadoop is excellent.""The solution's most valuable feature is the enterprise data platform.""The product is completely secure.""The search function is the most valuable aspect of the solution.""With a cluster available, you can manage the security layer using the shared SDX - it provides flexibility."

More Cloudera Distribution for Hadoop Pros →

"HPE Ezmeral Data Fabric can be accessed from any namespace globally as you would access it from a machine using an NFS.""My customers find the product cheaper compared to other solutions. The previous solution that we used did not have unified analytics like the runtime or the analog.""It is a stable solution...It is a scalable solution.""The model creation was very interesting, especially with the libraries provided by the platform.""I like the administration part."

More HPE Ezmeral Data Fabric Pros →

Cons
"Spark could be improved by adding support for other open-source storage layers than Delta Lake.""They could improve the issues related to programming language for the platform.""Apache Spark could potentially improve in terms of user-friendliness, particularly for individuals with a SQL background. While it's suitable for those with programming knowledge, making it more accessible to those without extensive programming skills could be beneficial.""In data analysis, you need to take real-time data from different data sources. You need to process this in a subsecond, do the transformation in a subsecond, and all that.""When using Spark, users may need to write their own parallelization logic, which requires additional effort and expertise.""The setup I worked on was really complex.""One limitation is that not all machine learning libraries and models support it.""It requires overcoming a significant learning curve due to its robust and feature-rich nature."

More Apache Spark Cons →

"The solution does not support multiple languages very well and this means users need to create work-arounds to implement some solutions.""Currently, we are using many other tools such as Spark and Blade Job to improve the performance.""They should focus on upgrading their technical capabilities in the market.""While the deployed product is generally functional, there are instances where it presents difficulties.""The price of this solution could be lowered.""Cloudera Distribution for Hadoop is not always completely stable in some cases, which can be a concern for big data solutions.""I would like to see an improvement in how the solution helps me to handle the whole cluster.""It would be useful if Cloudera had more tools like SQL Engines that offer the traditional relational database. We have to do a lot of work preparing the data outside Cloudera before getting it into the platform."

More Cloudera Distribution for Hadoop Cons →

"The deployment could be faster. I want more support for the data lake in the next release.""Having the ability to extend the services provided by the platform to an API architecture, a micro-services architecture, could be very helpful.""HPE Ezmeral Data Fabric is not compatible with third-party tools.""The product is not user-friendly.""Upgrading Ezmeral to a new version is a pain. They're trying to make the solution more container-friendly, so I think they're going in the right direction. The only problem we've had in the past was the upgrades. The process isn't smooth due to how the Red Hat operating system upgrades currently work."

More HPE Ezmeral Data Fabric Cons →

Pricing and Cost Advice
  • "Since we are using the Apache Spark version, not the data bricks version, it is an Apache license version, the support and resolution of the bug are actually late or delayed. The Apache license is free."
  • "Apache Spark is open-source. You have to pay only when you use any bundled product, such as Cloudera."
  • "We are using the free version of the solution."
  • "Apache Spark is not too cheap. You have to pay for hardware and Cloudera licenses. Of course, there is a solution with open source without Cloudera."
  • "Apache Spark is an expensive solution."
  • "Spark is an open-source solution, so there are no licensing costs."
  • "On the cloud model can be expensive as it requires substantial resources for implementation, covering on-premises hardware, memory, and licensing."
  • "It is an open-source solution, it is free of charge."
  • More Apache Spark Pricing and Cost Advice →

  • "When comparing with Oracle Sybase and SQL, it's cheaper. It's not expensive."
  • "The price could be better for the product."
  • "I haven't bought a license for this solution. I'm only using the Apache license version."
  • "Cloudera requires a license to use."
  • "Cloudera Distribution for Hadoop is expensive, with support costs involved."
  • "I wouldn't recommend CDH to others because of its high cost."
  • "The price is very high. The solution is expensive."
  • "The solution is expensive."
  • More Cloudera Distribution for Hadoop Pricing and Cost Advice →

  • "HPE is flexible with you if you are an existing customer. They offer different models that might be beneficial for your organization. It all depends on how you negotiate."
  • "The tool's price is cheap and based on a usage basis. The solution's licensing costs are yearly and there are no extra costs."
  • "There is a need for my company to pay for the licensing costs of the solution."
  • More HPE Ezmeral Data Fabric Pricing and Cost Advice →

    report
    Use our free recommendation engine to learn which Hadoop solutions are best for your needs.
    769,599 professionals have used our research since 2012.
    Questions from the Community
    Top Answer:We use Spark to process data from different data sources.
    Top Answer:In data analysis, you need to take real-time data from different data sources. You need to process this in a subsecond… more »
    Top Answer:The tool can be deployed using different container technologies, which makes it very scalable.
    Top Answer:The tool is expensive. Overall, it's not a cheap software tool, and that is why only large enterprises who are mature… more »
    Top Answer:The tool's ability to be deployed on a cloud model is an area of concern where improvements are required. The tool works… more »
    Top Answer:It is a stable solution...It is a scalable solution.
    Top Answer:There are some drawbacks in HPE Ezmeral Data Fabric when it comes to the interoperability part. HPE Ezmeral Data Fabric… more »
    Top Answer:The main purpose of HPE Ezmeral Data Fabric for me is that it acts as a database. In my company, we store our data with… more »
    Ranking
    1st
    out of 22 in Hadoop
    Views
    2,430
    Comparisons
    1,869
    Reviews
    26
    Average Words per Review
    444
    Rating
    8.7
    2nd
    out of 22 in Hadoop
    Views
    2,881
    Comparisons
    2,224
    Reviews
    14
    Average Words per Review
    443
    Rating
    8.1
    5th
    out of 22 in Hadoop
    Views
    1,637
    Comparisons
    1,016
    Reviews
    4
    Average Words per Review
    550
    Rating
    7.8
    Comparisons
    Also Known As
    MapR, MapR Data Platform
    Learn More
    Overview

    Spark provides programmers with an application programming interface centered on a data structure called the resilient distributed dataset (RDD), a read-only multiset of data items distributed over a cluster of machines, that is maintained in a fault-tolerant way. It was developed in response to limitations in the MapReduce cluster computing paradigm, which forces a particular linear dataflowstructure on distributed programs: MapReduce programs read input data from disk, map a function across the data, reduce the results of the map, and store reduction results on disk. Spark's RDDs function as a working set for distributed programs that offers a (deliberately) restricted form of distributed shared memory

    Cloudera Distribution for Hadoop is the world's most complete, tested, and popular distribution of Apache Hadoop and related projects. CDH is 100% Apache-licensed open source and is the only Hadoop solution to offer unified batch processing, interactive SQL, and interactive search, and role-based access controls. More enterprises have downloaded CDH than all other such distributions combined.

    Forward-leaning companies win market share because they leverage data more effectively than their competitors. Unlock the potential of your data assets with HPE Ezmeral Data Fabric (formerly MapR Data Platform). Empower your data science, analytics, and business teams by simplifying data management on a globally distributed scale. All with enterprise-grade reliability, security, and performance.

    Sample Customers
    NASA JPL, UC Berkeley AMPLab, Amazon, eBay, Yahoo!, UC Santa Cruz, TripAdvisor, Taboola, Agile Lab, Art.com, Baidu, Alibaba Taobao, EURECOM, Hitachi Solutions
    37signals, Adconion,adgooroo, Aggregate Knowledge, AMD, Apollo Group, Blackberry, Box, BT, CSC
    Valence Health, Goodgame Studios, Pico, Terbium Labs, sovrn, Harte Hanks, Quantium, Razorsight, Novartis, Experian, Dentsu ix, Pontis Transitions, DataSong, Return Path, RAPP, HP
    Top Industries
    REVIEWERS
    Computer Software Company30%
    Financial Services Firm15%
    University9%
    Marketing Services Firm6%
    VISITORS READING REVIEWS
    Financial Services Firm25%
    Computer Software Company13%
    Manufacturing Company7%
    Comms Service Provider6%
    REVIEWERS
    Financial Services Firm25%
    Computer Software Company21%
    Insurance Company14%
    Comms Service Provider11%
    VISITORS READING REVIEWS
    Financial Services Firm21%
    Computer Software Company16%
    Educational Organization8%
    Manufacturing Company8%
    VISITORS READING REVIEWS
    Financial Services Firm17%
    Computer Software Company17%
    Manufacturing Company8%
    Comms Service Provider7%
    Company Size
    REVIEWERS
    Small Business40%
    Midsize Enterprise18%
    Large Enterprise42%
    VISITORS READING REVIEWS
    Small Business17%
    Midsize Enterprise12%
    Large Enterprise71%
    REVIEWERS
    Small Business28%
    Midsize Enterprise17%
    Large Enterprise55%
    VISITORS READING REVIEWS
    Small Business16%
    Midsize Enterprise9%
    Large Enterprise75%
    REVIEWERS
    Small Business36%
    Large Enterprise64%
    VISITORS READING REVIEWS
    Small Business23%
    Midsize Enterprise11%
    Large Enterprise66%
    Buyer's Guide
    Hadoop
    April 2024
    Find out what your peers are saying about Apache, Cloudera, Amazon Web Services (AWS) and others in Hadoop. Updated: April 2024.
    769,599 professionals have used our research since 2012.