No more typing reviews! Try our Samantha, our new voice AI agent.
Snr Security Engineer at a tech vendor with 201-500 employees
Real User
Jul 17, 2019
Provides security analytics and has good scalability
Pros and Cons
  • "The scalability has been the most valuable aspect of the solution."
  • "The 2.3 version is quite stable, all of our customers use it, there are around 100,000+ users, and it runs 24/7."
  • "The management tools could use improvement. Some of the debugging tools need some work as well. They need to be more descriptive."

What is our primary use case?

We primarily use the solution for security analytics.

What is most valuable?

The scalability has been the most valuable aspect of the solution.

What needs improvement?

The management tools could use improvement. Some of the debugging tools need some work as well. They need to be more descriptive. 

For how long have I used the solution?

I've been using the solution for three years.
Buyer's Guide
Apache Spark
August 2026
Learn what your peers think about Apache Spark. Get advice and tips from experienced pros sharing their opinions. Updated: August 2026.
908,745 professionals have used our research since 2012.

What do I think about the stability of the solution?

The 2.3 version is quite stable. All of our customers use it, there are around 100,000+ users, and it runs 24/7.

What do I think about the scalability of the solution?

The scalability is very good.

How are customer service and support?

You actually buy Cloudera along with it. You don't really get any support, except you need support.

Which solution did I use previously and why did I switch?

In previous companies, we used MySQL platform and solutions like ArcSight and Splunk. We switched for scalability. MySQL wasn't going to scale, and we don't use Splunk at this company.

How was the initial setup?

The initial setup was complex. It is a complex tool. It's a lot to do with how you will use it. There is a lot to set up. They need to put a lot of scripts to it. There's nearly 60 to set up. When you set up the cloud, it takes about a day to set up. If you set it up on-premise, you know, on hardware, it only takes about a week.

What other advice do I have?

I would rate this solution eight out of 10. 

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user1059558 - PeerSpot reviewer
Portfolio Manager, Enterprise Solutions Architect at Capgemini
Real User
Apr 11, 2019
Supports streaming and micro-batch
Pros and Cons
  • "It is a better MR, supports streaming and micro-batch, and supports Spark ML and Spark SQL."
  • "Better data lineage support."

What is our primary use case?

Streaming telematics data.

How has it helped my organization?

It's a better MR, supports streaming and micro-batch, and supports Spark ML and Spark SQL.

What is most valuable?

It supports streaming and micro-batch.

What needs improvement?

Better data lineage support.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Apache Spark
August 2026
Learn what your peers think about Apache Spark. Get advice and tips from experienced pros sharing their opinions. Updated: August 2026.
908,745 professionals have used our research since 2012.
Director - Data Management, Governance and Quality at Hilton Worldwide
Real User
Mar 19, 2019
Powerful language but complicated coding
Pros and Cons
  • "Powerful language."
  • "It is like going back to the '80s for the complicated coding that is required to write efficient programs."

What is our primary use case?

Ingesting billions of rows of data all day.

How has it helped my organization?

Spark on AWS is not that cost-effective as memory is expensive and you cannot customize hardware in AWS. If you want more memory, you have to pay for more CPUs too in AWS.

What is most valuable?

Powerful language.

What needs improvement?

It is like going back to the '80s for the complicated coding that is required to write efficient programs.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
reviewer894894 - PeerSpot reviewer
Solutions Architect at a computer software company with 51-200 employees
User
Jul 11, 2018
Features include machine learning, real time streaming, and data processing. It doesn't enable spark job scheduling with monitoring capability.
Pros and Cons
  • "The fault tolerant feature is provided."
  • "It provides a scalable machine learning library."
  • "Machine learning, real time streaming, and data processing are fantastic, as well as the resilient or fault tolerant feature."
  • "It should support more programming languages."
  • "Needs to provide an internal schedule to schedule spark jobs with monitoring capability."

What is our primary use case?

Used for building big data platforms for processing huge volumes of data. Additionally, streaming data is critical.

How has it helped my organization?

It provides a scalable machine learning library so that we can train and predict user behavior for promotion purposes.

What is most valuable?

Machine learning, real time streaming, and data processing are fantastic, as well as the resilient or fault tolerant feature.

What needs improvement?

I would suggest for it to support more programming languages, and also provide an internal scheduler to schedule spark jobs with monitoring capability.

For how long have I used the solution?

Trial/evaluations only.
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user786777 - PeerSpot reviewer
Manager | Data Science Enthusiast | Management Consultant at a consultancy with 5,001-10,000 employees
Real User
Dec 10, 2017
We can now harness richer data sets and benefit from use cases
Pros and Cons
  • "With Hadoop-related technologies, we can distribute the workload with multiple commodity hardware."
  • "Organisations can now harness richer data sets and benefit from use cases, which add value to their business functions."
  • "Include more machine learning algorithms and the ability to handle streaming of data versus micro batch processing."
  • "At times when users do not know how to use Spark and request a lot of resources, then the underlying JVMs can crash, which is a big sense of worry."

How has it helped my organization?

Organisations can now harness richer data sets and benefit from use cases, which add value to their business functions.

What is most valuable?

Distributed in memory processing. Some of the algorithms are resource heavy and executing this requires a lot of RAM and CPU. With Hadoop-related technologies, we can distribute the workload with multiple commodity hardware.

What needs improvement?

Include more machine learning algorithms and the ability to handle streaming of data versus micro batch processing.

For how long have I used the solution?

Three to five years.

What do I think about the stability of the solution?

At times when users do not know how to use Spark and request a lot of resources, then the underlying JVMs can crash, which is a big sense of worry. 

What do I think about the scalability of the solution?

No issues.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user746943 - PeerSpot reviewer
Big Data and Cloud Solution Consultant at a financial services firm with 10,001+ employees
Real User
Oct 2, 2017
Provides flexibility for application creation with less coding effort
Pros and Cons
  • "DataFrame: Spark SQL gives the leverage to create applications more easily and with less coding effort."
  • "Dynamic DataFrame options are not yet available."

What is most valuable?

DataFrame: Spark SQL gives the leverage to create applications more easily and with less coding effort.

How has it helped my organization?

We developed a tool for data ingestion from HDFS->Raw->L1 layer with data quality checks, putting data to elastic search, performing CDC.

What needs improvement?

Dynamic DataFrame options are not yet available.

For how long have I used the solution?

One and a half years.

What do I think about the stability of the solution?

No.

What do I think about the scalability of the solution?

No.

What other advice do I have?

Spark gives the flexibility for developing custom applications.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user746673 - PeerSpot reviewer
Sr. Software Engineer at a tech vendor with 1-10 employees
Real User
Oct 1, 2017
Helped us reduce 3TB Google Ngrams in hours instead of days
Pros and Cons
  • "The most valuable feature is the Fault Tolerance and easy binding with other processes like Machine Learning, graph analytics."
  • "After using Spark, we were able to accomplish this task within hours."
  • "More ML based algorithms should be added to it, to make it algorithmic-rich for developers."

What is most valuable?

The most valuable feature is the Fault Tolerance and easy binding with other processes like Machine Learning, graph analytics. The community is growing and hence executing ML in a distributed fashion is quite good.

How has it helped my organization?

Previously we were using Hadoop MapReduce to reduce the Google Ngrams (3TB), which took us approximately five days on our cluster. After using Spark, we were able to accomplish this task within hours.

What needs improvement?

This product is already improving as the community is developing it rapidly. More ML based algorithms should be added to it, to make it algorithmic-rich for developers.

For how long have I used the solution?

Two and a half years.

What do I think about the stability of the solution?

No, I did not encounter any problems with the stability. It is also quite backwards compatible.

What do I think about the scalability of the solution?

No I did not as of now, it is quite scalable. Using simple scripts you can add as many workers as you want.

What other advice do I have?

This is a very good product for the big data analytics and integrates well with other parts like Machine Learning and graph analytics.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user326142 - PeerSpot reviewer
Architect at a healthcare company with 51-200 employees
Real User
Sep 27, 2017
Having everything in the same framework has helped us out a lot
Pros and Cons
  • "ETL and streaming capabilities."
  • "Having everything in the same framework has helped us out a lot."
  • "Stability in terms of API (things were difficult, when transitioning from RDD to DataFrames, then to DataSet)."

What is most valuable?

ETL and streaming capabilities.

How has it helped my organization?

Made Big Data processing more convenient and a uniform framework adds to efficiency of usage since the same framework can be used for batch and stream processing.

What needs improvement?

Stability in terms of API (things were difficult, when transitioning from RDD to DataFrames, then to DataSet).

For how long have I used the solution?

I have used Spark since its inception in March 2015, from Spark 1.1 onwards.

Currently, I use 2.2 extensively.

What do I think about the stability of the solution?

Yes, occasionally with different APIs.

What do I think about the scalability of the solution?

No.

How are customer service and technical support?

Since we were using the Open Source version of Apache Spark, without the Databricks support, we never used technical support form Databricks.

Which solution did I use previously and why did I switch?

Yes we used Hive, Pig, and Storm. Having everything in the same framework has helped us out a lot.

Which other solutions did I evaluate?

Yes, we considered other big data products in the Big Data Ecosystem.

What other advice do I have?

Go for it.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user372393 - PeerSpot reviewer
Big Data Consultant at a tech services company with 501-1,000 employees
Consultant
Aug 25, 2017
We are able to solve problems, e.g., reporting on big data, that we were not able to tackle in the past.
Pros and Cons
  • "The good performance. The nice graphical management console. The long list of ML algorithms."
  • "We are able to solve problems, e.g., reporting on big data, that we were not able to tackle in the past."
  • "Apache Spark provides very good performance The tuning phase is still tricky."
  • "The initial set-up is quite complex because you have to set-up many different configuration parameters that are deployment-specific."

What is most valuable?

The good performance. The nice graphical management console. The long list of ML algorithms.

How has it helped my organization?

We are able to solve problems, e.g., reporting on big data, that we were not able to tackle in the past.

What needs improvement?

Apache Spark provides very good performance The tuning phase is still tricky.

For how long have I used the solution?

I've used it for 2 years.

What was my experience with deployment of the solution?

We didn't have an issue with the deployment.

What do I think about the stability of the solution?

In the past we deployed Spark 1.3 to use Spark SQL but unfortunately one of our queries failed because of a bug fixed in following releases. Then we moved to Spark 1.6 but still some queries were failing when run against huge datasets. Now we are using version 2.1: it is more stable, it ensures better performances and the SQL/ML parts are reacher than before.

What do I think about the scalability of the solution?

I've had no issues with the scalability.

How is customer service and technical support?

Customer Service:

I've never had to use customer service.

Technical Support:

I've never had to use technical support.

How was the initial setup?

The initial set-up is quite complex because you have to set-up many different configuration parameters that are deployment-specific. It is not trivial to set-up the correct configuration with so many variables involved.

What about the implementation team?

In-house team. The setup itself is not a problem when you have just to test the system. The challenging part is discovering the optimal configuration needed to obtain a production system proving good performance.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
it_user371832 - PeerSpot reviewer
Chief System Architect at a marketing services firm with 501-1,000 employees
Vendor
Mar 30, 2016
Spark gives us the ability to run queries on MySQL database without pressurising our database
Pros and Cons
  • "With Spark SQL we've now the capabilities to analyse very large quantities of data located in S3 on Amazon at very low cost comparing other solution we checked."
  • "Spark Streaming is difficult to stabilize as you're always dependant to your stream flow."

What is most valuable?

With spark SQL we've now the capabilities to analyse very large quantities of data located in S3 on Amazon at very low cost comparing other solution we checked. 

We also use our own Spark cluster to aggregate data on near real time and save the result on MySQL database. 

We've started new projects using the machine learning library ML.

How has it helped my organization?

Until Spark we didn't have the ability to analyse this quantity of data we're talking about two TB/hour. So we're now able to produce a lot of reports, and are also able to develop machine learning based analysis to optimize our business. 

We've central access to every piece of data in the company including finance, business, debug etc. and the ability to join all this data together.

What needs improvement?

Spark is actually very good for batch analysis much more good than Hadoop, it's much simple, much more quicker etc., but it actually lacks the ability to perform real-time querying like Vertica or Redshift.  

Also, it is more difficult for an end user to work with Spark than normal database. even comparing with analytic database like Vertica or Redshift.

For how long have I used the solution?

We're now using Spark-Streaming and Spark-SQL for almost 2 years. 

What was my experience with deployment of the solution?

We're working on AWS so we need to have a managed environment. We've choose to go with a solution based on Chef to deploy and configure the spark clusters. Tip : if you don't have any devops you can use the ec2 script (provided by spark distro) to deploy cluster on amazon. We've tested it and work perfectly.  

What do I think about the stability of the solution?

Spark Streaming is difficult to stabilize as you're always dependant to your stream flow. If you start to be late on the consumer you've a serious problem. We've encountered a lot of stability issue to configure it as expected

What do I think about the scalability of the solution?

It's linked to stability in our case it's takes time to evaluate what is the correct size of the cluster you need. It's very important to always add to you jobs monitoring to be able to understand what's the problem. We use datadog as monitoring platform

Which solution did I use previously and why did I switch?

Yes to make this job we've used a MySQL database. We switch because MySQL is not a scalable solution and we've reach it's limits.

How was the initial setup?

Setup a spark cluster can be difficult. it's related to your clustering strategy. There is 4 solution at least. 

ec2 script : work only on Amazon AWS

Standalone : manually configuration (hard)

Yarn : to leverage your already existing Hadoop environment.

Mesos : to use with your other Mesos ready application

What about the implementation team?

We use Databricks as online DB ad hoc query. It's work on AWS as managed service, it manage for you the cluster creation, configuration and monitoring.

Give a notebook oriented user interface to query any data source using Spark: DB, Parquet, CSV, Avro etc...

Which other solutions did I evaluate?

Yes we've started to evaluate analytics databases : vertica, exasol, and other for all the them the price was an issue regarding the quantity of data we want to manipulate.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Download our free Apache Spark Report and get advice and tips from experienced pros sharing their opinions.
Updated: August 2026
Buyer's Guide
Download our free Apache Spark Report and get advice and tips from experienced pros sharing their opinions.