Try our new research platform with insights from 80,000+ expert users

Collibra Catalog vs StreamSets comparison

 

Comparison Buyer's Guide

Executive Summary

Review summaries and opinions

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
 

Categories and Ranking

Collibra Catalog
Average Rating
8.0
Reviews Sentiment
7.3
Number of Reviews
11
Ranking in other categories
Metadata Management (3rd)
StreamSets
Average Rating
8.4
Reviews Sentiment
7.0
Number of Reviews
21
Ranking in other categories
Data Integration (23rd)
 

Mindshare comparison

Collibra Catalog and StreamSets aren’t in the same category and serve different purposes. Collibra Catalog is designed for Metadata Management and holds a mindshare of 11.8%, up 10.2% compared to last year.
StreamSets, on the other hand, focuses on Data Integration, holds 1.6% mindshare, up 1.4% since last year.
Metadata Management
Data Integration
 

Featured Reviews

Tejbir Singh - PeerSpot reviewer
Facilitates data quality monitoring and AI governance with a complete suite of tools
When I initially started with Collibra, it was just a data cataloging platform with governance workflows around it. Now they have acquired a lot of other tools, or they have merged or acquired different platforms. It is a complete suite of tools for managing data. We can monitor data quality and take actions on the profiling results obtained by running data quality checks. Collibra helps catalog data assets, monitor the health of data assets, and take necessary actions. If we find data quality issues, it also provides a medium to capture those issues and how to remediate them. The workflows allow the creation of custom workflows based on needs. The newest addition in their tool suite is AI governance, which allows cataloging all AI models currently deployed or even in the pre-production stage. It helps document model meanings and the risks involved, thus managing all risks related to AI deployments.
Ved Prakash Yadav - PeerSpot reviewer
Useful for data transformation and helps with column encryption
We use various tools and alerting systems to notify us of pipeline errors or failures. StreamSets supports data governance and compliance by allowing us to encrypt incoming data based on specified rules. We can easily encrypt columns by providing the column name and hash key. If you're considering using StreamSets for the first time, I would advise first understanding why you want to use it and how it will benefit you. If you're dealing with change tracking or handling large amounts of data, it could be cost-effective compared to services like Amazon. It's easy to schedule and manage tasks with the tool, and you can enhance your skills as an ETL developer. You can easily migrate traditional pipelines built on platforms like Informatica or Talend to StreamSets. I rate the overall solution an eight out of ten.

Quotes from Members

We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
 

Pros

"Except for data quality, everything is perfect."
"We have had no complaints about the stability."
"The workflows allow the creation of custom workflows based on needs."
"Collibra Catalog allows us to automate metadata management, significantly saving time, effort, and finances."
"Collibra Catalog is simple to use and user-friendly for those who are not technically inclined since it is easy to find while also easy to see data lineage diagrams."
"The data lineage capability is valuable as it shows how different sources are connected and how data flows, which is crucial for projects like migrations. Moreover, data lineage visualization in Collibra Catalog aids our data governance initiatives."
"Gartner identifies Collibra Catalog as the leader, which aligns with our observations."
"Collibra Catalog's best feature is the data quality checker."
"The ETL capabilities are very useful for us. We extract and transform data from multiple data sources, into a single, consistent data store, and then we put it in our systems. We typically use it to connect our Apache Kafka with data lakes. That process is smooth and saves us a lot of time in our production systems."
"The ability to have a good bifurcation rate and fewer mistakes is valuable."
"In StreamSets, everything is in one place."
"The Ease of configuration for pipes is amazing. It has a lot of connectors. Mainly, we can do everything with the data in the pipe. I really like the graphical interface too"
"I really appreciate the numerous ready connectors available on both the source and target sides, the support for various media file formats, and the ease of configuring and managing pipelines centrally."
"The best feature that I really like is the integration."
"The entire user interface is very simple and the simplicity of creating pipelines is something that I like very much about it. The design experience is very smooth."
"StreamSets’ data drift resilience has reduced the time it takes us to fix data drift breakages. For example, in our previous Hadoop scenario, when we were creating the Sqoop-based processes to move data from source to destinations, we were getting the job done. That took approximately an hour to an hour and a half when we did it with Hadoop. However, with the StreamSets, since it works on a data collector-based mechanism, it completes the same process in 15 minutes of time. Therefore, it has saved us around 45 minutes per data pipeline or table that we migrate. Thus, it reduced the data transfer, including the drift part, by 45 minutes."
 

Cons

"The tool's overall functionalities need to improve since, nowadays, many tools, from a business perspective, are easy to use."
"I'd like to see more integration with other reporting sources."
"There is an issue with Collibra Catalog's pricing model, especially for organizations with many databases, as the initial package comes with a limited number of connectors."
"Collibra Catalog could improve its automation to increase the efficiency of the software."
"More automation and artificial intelligence involvement are necessary. Reducing required employee involvement and enhancing ease of use are vital."
"A key area for improvement in Collibra Catalog lies in its integration capabilities, particularly with a broader range of sources."
"If it can become more user-intuitive and work on integrating with communication platforms like Slack or Teams, it would significantly help business users."
"If the price is a bit reduced, that would be better."
"StreamSets should provide a mechanism to be able to perform data quality assessment when the data is being moved from one source to the target."
"They need to improve their customer care services. Sometimes it has taken more than 48 hours to resolve an issue. That should be reduced. They are aware of small or generic issues, but not the more technical or deep issues. For those, they require some time, generally 48 to 72 hours to respond. That should be improved."
"The logging mechanism could be improved. If I am working on a pipeline, then create a job out of it and it is running, it will generate constant logs. So, the logging mechanism could be simplified. Now, it is a bit difficult to understand and filter the logs. It takes some time."
"One thing that I would like to add is the ability to manually enter data. The way the solution currently works is we don't have the option to manually change the data at any point in time. Being able to do that will allow us to do everything that we want to do with our data. Sometimes, we need to manually manipulate the data to make it more accurate in case our prior bifurcation filters are not good. If we have the option to manually enter the data or make the exact iterations on the data set, that would be a good thing."
"We create pipelines or jobs in StreamSets Control Hub. It is a great feature, but if there is a way to have a folder structure or organize the pipelines and jobs in Control Hub, it would be great. I submitted a ticket for this some time back."
"One area for improvement could be the cloud storage server speed, as we have faced some latency issues here and there."
"If you use JDBC Lookup, for example, it generally takes a long time to process data."
"Sometimes, it is not clear at first how to set up nodes. A site with an explanation of how each node works would be very helpful."
 

Pricing and Cost Advice

"Collibra offers a per-user licensing model."
"I think they can bring a few more features and align better with other quality products."
"Collibra Catalog is fairly priced - I would rate their pricing seven out of ten."
"The product is highly priced compared to other vendors."
"There are different versions of the product. One is the corporate license version, and the other one is the open-source or free version. I have been using the corporate license version, but they have recently launched a new open-source version so that anybody can create an account and use it. The licensing cost varies from customer to customer. I don't have a lot of input on that. It is taken care of by PMO, and they seem fine with its pricing model. It is being used enterprise-wide. They seem to have got a good deal for StreamSets."
"It has a CPU core-based licensing, which works for us and is quite good."
"It's not so favorable for small companies."
"Its pricing is pretty much up to the mark. For smaller enterprises, it could be a big price to pay at the initial stage of operations, but the moment you have the Seed B or Seed C funding and you want to scale up your operations and aren't much worried about the funds, at that point in time, you would need a solution that could be scaled."
"The pricing is too fixed. It should be based on how much data you need to process. Some businesses are not so big that they process a lot of data."
"StreamSets is an expensive solution."
"I believe the pricing is not equitable."
"We are running the community version right now, which can be used free of charge."
report
Use our free recommendation engine to learn which Metadata Management solutions are best for your needs.
863,429 professionals have used our research since 2012.
 

Top Industries

By visitors reading reviews
Financial Services Firm
30%
Computer Software Company
8%
Manufacturing Company
7%
Government
6%
Computer Software Company
11%
Financial Services Firm
10%
Manufacturing Company
10%
Insurance Company
8%
 

Company Size

By reviewers
Large Enterprise
Midsize Enterprise
Small Business
 

Questions from the Community

What do you like most about Collibra Catalog?
The data lineage capability is valuable as it shows how different sources are connected and how data flows, which is crucial for projects like migrations. Moreover, data lineage visualization in C...
What is your experience regarding pricing and costs for Collibra Catalog?
Pricing is not under my purview as I am an architect. The platform team handles the licensing aspects.
What needs improvement with Collibra Catalog?
I have utilized the sophisticated search capability in Collibra Catalog, and it can be improved by implementing more natural language search capabilities. Currently, we need to enter the asset name...
What do you like most about StreamSets?
The best thing about StreamSets is its plugins, which are very useful and work well with almost every data source. It's also easy to use, especially if you're comfortable with SQL. You can customiz...
What needs improvement with StreamSets?
One issue I observed with StreamSets is that the memory runs out quickly when processing large volumes of data. Because of this memory issue, we have to upgrade our EC2 boxes in the Amazon AWS infr...
What is your primary use case for StreamSets?
We are using StreamSets for batch loading.
 

Overview

 

Sample Customers

AXA XL, DNB, Adobe, PMI, Holland America Line, UC Davis Health, Cox Automotive
Availity, BT Group, Humana, Deluxe, GSK, RingCentral, IBM, Shell, SamTrans, State of Ohio, TalentFulfilled, TechBridge
Find out what your peers are saying about Informatica, Alation, Collibra and others in Metadata Management. Updated: July 2025.
863,429 professionals have used our research since 2012.