

IBM InfoSphere DataStage and AWS Glue compete in the data integration and ETL tools category. AWS Glue appears to have the advantage due to its cloud-based architecture, cost-efficiency, and ease of integration with other AWS services.
Features: IBM InfoSphere DataStage offers robust scalability, extensive transformation capabilities, and strong metadata management, allowing users to handle large data sets and high customization for different data latencies. AWS Glue provides seamless integration with other AWS services, automation, and serverless architecture, enhancing its scalability and cost-efficiency. Its data catalog and support for Python scripting further improve its usability for data management tasks.
Room for Improvement: IBM InfoSphere DataStage needs improvement in cost-effectiveness, user interface complexity, and support for modern data sources such as cloud integrations. Simplified administration tools are also needed. AWS Glue could improve in cost predictability, expand support beyond AWS environments, and refine ETL capabilities. Enhanced documentation and no-code features for complex transformations are also desirable improvements.
Ease of Deployment and Customer Service: IBM InfoSphere DataStage is mainly used on-premises, making cloud integration deployments complex. Customer service experiences vary, with mixed reviews regarding response speed. AWS Glue offers flexibility in public, private, and hybrid cloud deployments, benefiting from AWS ecosystem integration. Customer service is generally good but can improve in handling complex inquiries.
Pricing and ROI: IBM InfoSphere DataStage is costly, particularly for small and medium enterprises, with significant licensing and maintenance expenses. AWS Glue operates on a pay-as-you-go model, offering cost-effective solutions for varying usage levels, although scaling up may incur considerable costs. Customers appreciate AWS Glue's serverless billing flexibility, although some find its overall cost high. Both solutions yield a positive ROI when managed effectively, with IBM InfoSphere DataStage noted for optimization improvements and AWS Glue praised for efficiency.
I advocate using Glue in such cases.
Upgrades occur every four months, and new developments coincide with version updates.
I would rate AWS support eight out of ten because they are technically strong and helpful in debugging complex Glue and cloud issues, and they are very responsive.
We also have the flexibility to submit a feature request to be included as part of the wishlist, potentially becoming a product feature in subsequent releases.
I rate their support as nine on a scale from one to ten.
IBM tech support has allocated dedicated resources, making it satisfactory.
It can easily handle data from one terabyte to 100 terabytes or more, scaling nicely with larger datasets.
For jobs requiring multiple RAM usage, we increase the number of workers accordingly.
If the job provided suggestions about running this kind of parallel processing and how many virtual nodes are required, it would help.
As a managed service, it reduces management burdens.
Learning the latest functionalities is crucial, and while challenging, it is a vital part of staying current and ensuring an efficient ETL process.
With AWS, I gather data from multiple sources, clean it up, normalize it, de-duplicate it, and make it presentable.
A more user-friendly and simpler process would help speed up the deployment process.
If the job itself gave some guidance, such as running this parallel processing with this many nodes, it would help; I think that is missing.
I wonder if it supports other areas, such as cloud environments with open source support, or EdgeShift.
The solution needs improvement in connectivity with big data technologies such as Spark.
AWS charges based on runtime, which can be quite pricey.
The smallest cost for a project is around €700, while the largest can reach up to €7,000 based on the scale of the usage.
Costing depends on resource usage, and cost optimization may involve redesigning jobs for flexibility.
Pricing for IBM InfoSphere DataStage is moderate and not much expensive.
AWS Glue is very efficient and integrates well with the AWS ecosystem.
AWS Glue also enhances job scheduling and orchestration capabilities, integrating with AWS Glue Studio for comprehensive data workflow management.
For ETL, I feel the performance is excellent. If I create jobs in a standard way, the performance is great, and maintenance is also seamless.
It is straightforward from a design and development perspective, and also for deployment.
IBM InfoSphere DataStage is very scalable, allowing us to extend it according to our processing needs.
I have leveraged IBM InfoSphere DataStage's integration with IBM's Information Server suite, and it is indeed beneficial.
| Company Size | Count |
|---|---|
| Small Business | 11 |
| Midsize Enterprise | 6 |
| Large Enterprise | 32 |
| Company Size | Count |
|---|---|
| Small Business | 23 |
| Midsize Enterprise | 4 |
| Large Enterprise | 26 |
AWS Glue is a serverless cloud data integration tool that facilitates the discovery, preparation, movement, and integration of data from multiple sources for machine learning (ML), analytics, and application development. The solution includes additional productivity and data ops tooling for running jobs, implementing business workflows, and authoring.
AWS Glue allows users to connect to more than 70 diverse data sources and manage data in a centralized data catalog. The solution facilitates visual creation, running, and monitoring of extract, transform, and load (ETL) pipelines to load data into users' data lakes. This Amazon product seamlessly integrates with other native applications of the brand and allows users to search and query cataloged data using Amazon EMR, Amazon Athena, and Amazon Redshift Spectrum.
The solution also utilizes application programming interface (API) operations to transform users' data, create runtime logs, store job logic, and create notifications for monitoring job runs. The console of AWS Glue connects all of these services into a managed application, facilitating the monitoring and operational processes. The solution also performs provisioning and management of the resources required to run users' workloads in order to minimize manual work time for organizations.
AWS Glue Features
AWS Glue groups its features into four categories - discover, prepare, integrate, and transform. Within those groups are the following features:
AWS Glue Benefits
AWS Glue offers a wide range of benefits for its users. These benefits include:
Reviews from Real Users
Mustapha A., a cloud data engineer at Jems Groupe, likes AWS Glue because it is a product that is great for serverless data transformations.
Liana I., CEO at Quark Technologies SRL, describes AWS Glue as a highly scalable, reliable, and beneficial pay-as-you-go pricing model.
IBM InfoSphere DataStage is a high-quality data integration tool that aims to design, develop, and run jobs that move and transform data for organizations of different sizes. The product works by integrating data across multiple systems through a high-performance parallel framework. It supports extended metadata management, enterprise connectivity, and integration of all types of data.
The solution is the data integration component of IBM InfoSphere Information Server, providing a graphical framework for moving data from source systems to target systems. IBM InfoSphere DataStage can deliver data to data warehouses, data marts, operational data sources, and other enterprise applications. The tool works with various types of patterns - extract, transform and load (ETL), and extract, load, and transform (ELT). The scalability of the platform is achieved by using parallel processing and enterprise connectivity.
The solution has various versions, catering to different types of companies, which include the Server Edition, the Enterprise Edition, and the MVS Edition. Depending on which version a company has bought, different goals can be achieved. They include the following:
IBM InfoSphere DataStage can be deployed in various ways, including:
IBM InfoSphere DataStage Features
The tool has various features through which users can integrate and utilize their data effectively. The components of IBM InfoSphere DataStage include:
IBM InfoSphere DataStage Benefits
This solution offers many benefits for the companies that utilize it for data integration. Some of these benefits include:
Reviews from Real Users
A data/solution architect at a computer software company says the product is robust, easy to use, has a simple error logging mechanism, and works very well for huge volumes of data.
Tirthankar Roy Chowdhury, team leader at Tata Consultancy Services, feels the tool is user-friendly with a lot of functionalities, and doesn't require much coding because of its drag-and-drop features.
We monitor all Cloud Data Integration reviews to prevent fraudulent reviews and keep review quality high. We do not post reviews by company employees or direct competitors. We validate each review for authenticity via cross-reference with LinkedIn, and personal follow-up with the reviewer when necessary.