We use this solution for finding anomalies and applying the rules to the streaming data.
There are around 50 people using this solution in my organization, including data scientists.
We use this solution for finding anomalies and applying the rules to the streaming data.
There are around 50 people using this solution in my organization, including data scientists.
The ability to stream data and the windowing feature are valuable. There are a number of targeted integration points, so that is a difference between Stream Analytics and Databricks. The integrations input or output are better in Databricks. It's accessible to use any of the Python or even Java. I can use the third party, deploy it, and use it.
Support for Microsoft technology and the compatibility with the .NET framework is somewhat missing. There should be reliability between these two. Databricks is based on open sources. If it's more synchronous between the Microsoft technology and the programming languages, it'll be better. Python has better languages, but compatibility would be a great help.
I would like to have better support for Microsoft technology and better language components.
With Azure or Cosmo DB, I can store other data links or time series data tables. That would be a great help for analytics in real time.
I have been using Databricks for eight months.
The scalability is fine. We had thousands of devices and were sending data infrequently, so that worked for us. If the amount increases, the windowing function and job schedule may not perform as expected.
I would rate technical support 4 out of 5. We had some issues with setup, and they were finally solved but it was after following up a few times.
Azure Stream Analytics is easy to use and easy to deploy. It's a little bit better. Databricks is still having some stability issues. Azure Stream Analytics has a few input and output sources, and it's scalable to all types of third party or interfaces.
Setup was complex. There were some issues with setting up a database and installing the third party component on top of services. I would rate the setup 3 out of 5.
Implementation was done in-house.
The cost is around $600,000 for 50 users.
I would rate the price 2 out of 5.
I would rate this solution 8 out of 10.
I use Databricks to manage the setting up of data lakes for SaaS.
The biggest problem associated with the product is that it is quite pricey. We cannot find a better solution than Databricks in the market currently.
I have been using Databricks for a year.
It is an expensive tool. The licensing model is a pay-as-you-go one.
The tool helps with data processing and analytics with large-scale data or big data since it is associated with managing data at a large scale.
For my general use cases, I would say that I am not a technical person, so I cannot explain to you how the tool helps with the area of data engineering tasks.
There is another team in my company that is involved in the use of machine learning and AI features in Databricks. My team is mostly into operations. The tool is used in a multi-country project.
For example, in my company, they make some shopping decisions related to solutions based on what is the product chosen by the whole company.
I rate the tool an eight out of ten.
We use Databricks for data science work in projects that create data pipelines, pre-processing, data wrangling, big data cluster management and ML, machine learning and deep learning tasks.
Databricks collaborates very well with the Azure platform, Dataiku, and enterprise AI tool. Databricks is a new connection to pull the data or connect to the Spark cluster. It is helpful for us to instance it or distribute the load through the Spark cluster, and it is very user-friendly.
The most valuable feature is the Spark cluster which is very fast for heavy loads, big data processing and Pi Spark.
Databricks as a solution is integrated with Azure, but Google Cloud has some restrictions. I'm not sure about AWS Cloud, but it would be great if Databricks could integrate all the cloud platforms. Regarding additional features, we would like to see them mostly on the data engineering side, where we have a Spark cluster and some inbuilt ML. In addition, pre-processing steps will be useful.
We have been using this solution for two years and are using the latest update.
It is a stable solution as long as the Microsoft Azure Platform is stable too.
It is a scalable solution, both vertically and horizontally, which is good. My organization is big, and we have a lot of users. In my department, we have about 15 people using Databricks.
We have not escalated any issues to technical support, but we initially struggled with configuration and the settings of Hive metastore, but we resolved it. I rate the technical support a nine out of ten.
Positive
We were using the looped EMR elastic MapReduce from AWS before using Databricks. We switched to Databricks because the whole platform changed from AWS to Azure platform, and Databricks comes as a package.
The initial setup was easy to complete and not complex. It may initially be challenging for a new user, but it improves over time. The CICD pipeline works well with the Microsoft Azure platform because the continuous integration, development and deployment come with the Git integration. It makes it easier for Databricks and the CICD. The deployment should be improved from the perspective of auto ML functionality, so it doesn't have intensive automation learning capability.
We don't use Databricks directly because we work on a data science project. It requires an auto ML and inbuilt machine learning capability. We found capabilities like the large language model using NLP and other deep learning models that are not that intensive. It is meant for data engineering purposes rather than data science purposes. It'll be great if Databricks could be intensive for data science.
We used a third-party, Dataiku platform for the deployment, where we connected to Databricks and completed the ML ops. We required about three people for deployment, and it is easy to maintain the solution.
We have seen an ROI but cannot differentiate because it also comes with the Azure platform.
I do not have details about the pricing.
I rate this solution a nine out of ten. Regarding advice, Databricks is a very good platform, popular and easy to use daily for data engineers and data scientists who rely on a large dataset to do advanced analytics reporting. It's a very good tool.
We are using Databricks for machine learning workloads specifically.
Databricks aligns well with our skillset and overall approach. We sought out their solution specifically for a big data application we are currently working on, as we needed a platform capable of handling large amounts of data and building models. Additionally, the fact that they use open-source software and can integrate data warehouse and data lake systems was particularly appealing, as we have encountered such issues in the past. We determined that Databricks would be an effective solution for our needs.
The most valuable feature of Databricks is the integration of the data warehouse and data lake, and the development of the lake house. Additionally, it integrates well with Spark for processing data in production.
The solution could be improved by adding a feature that would make it more user-friendly for our team. The feature is simple, but it would be useful. Currently, our team is more familiar with the language R, but Databricks requires the use of Jupyter Notebooks which primarily supports Python. We have tried using RStudio, but it is not a fully integrated solution. To fully utilize Databricks, we have to use the Jupyter interface. One feature that would make it easier for our team to adopt the Jupyter interface would be the ability to select a specific variable or line of code and execute it within a cell. This feature is available in other Jupyter Notebooks outside of Databricks and in our own IDE, but it is not currently available within Databricks. If this feature were added, it would make the transition to using Databricks much smoother for our team.
The most important feature other than the Jupyter interface would be to have the RStudio interface inside Databricks. This would be perfect.
We have been using Databricks for approximately one year.
The stability of Databricks is good.
I rate the stability of Databricks a nine out of ten.
Databricks is scalable.
I rate the scalability of Databricks a nine out of ten.
I have been receiving responsive answers from Databricks's support. I have been pleased with the support.
I rate the support from Databricks a ten out of ten.
Positive
The initial setup of Databricks is simple. I did not experience any challenges. The time it takes for the deployment is approximately four hours.
I rate the initial setup of Databricks.
We did the deployment of the solution in-house. There were three people involved in the deployment. A data engineer, data analyst, and machine learning engineer.
We have only incurred the cost of our AWS cloud services. This is because during this period, Databricks provided us with an extended evaluation period, and we have not spent much money yet. We are just starting to incur costs this month, I will know more later on the full cost perspective.
We only pay standard fees for the solution.
We use a data engineer, data analyst, and machine learning engineer for the maintenance of the solution.
I rate Databricks a nine out of ten.
Our team is currently utilizing machine learning for various applications, and a few members are also exploring Databrick's use for ML operations.
In the manufacturing industry, Databricks can be beneficial to use because of machine learning. It is useful for tasks, such as product analysis or predictive maintenance.
I have been using Databricks for approximately six months
The stability of the clusters or the instances of Databricks would be better if it was a much more stable environment. We've had issues with crashes.
The scalability of Databricks is good as long as you have a data lake, and it's easy to scale.
We have approximately 50 users using this solution in my company.
We have a different team who handles the support. I do not have contact with Databricks support.
I have not used a similar solution to Databricks.
I have seen an ROI using Databricks.
I rate the price of Databricks as eight out of ten.
Having a good understanding of physical security in relation to cybersecurity in an OT (Operational Technology) environment would be beneficial, and utilizing an existing data lake prior to implementing a Databricks initiative would greatly aid in its success.
I rate Databricks an eight out of ten.
I am using Databricks in my company.
The most valuable feature of Databricks is the notebook, data factory, and ease of use.
I have been using Databricks for approximately nine months.
The performance and stability of Databricks are good. It is quick and I have not had problems.
Databricks is highly scalable.
We have 200 people using the solution in my organization.
When I used the support, I had communication problems because of the language barrier with the agent. The accent was difficult to understand.
I have not worked with another solution prior to Databricks.
The price of Databricks is reasonable compared to other solutions.
I rate Databricks an eight out of ten.
I am using Databricks for creating business intelligence solutions.
The most valuable feature of Databricks is the integration with Microsoft Azure.
Databricks can improve by making the documentation better.
I have been using Databricks for approximately one year.
Databricks is stable.
The scalability of Databricks is good.
We have approximately 500 users using this solution in my organization.
I have not used the support from Databricks.
We previously used Microsoft stacks. We chose Databricks because the processing power was better and it was a better fit for our use case.
The initial setup of Databricks was not straightforward. We had to do trial and error and we learned as we went along.
I rate the initial setup of Databricks a four out of five.
We did the implementation of Databricks in-house. The solution requires ongoing maintenance.
I would recommend this solution to others.
My advice to others is for them to first do a small proof of concept and then see how it works out and then take it from there.
I rate Databricks an eight out of ten.
This solution offers a lake house data concept that we have found exciting. We are able to have a large amount of data in a data lake and can manage all relational activities. All asset complaints properties are available and this is very useful to ensure the quality of all data.
The connectivity with various BI tools could be improved, specifically the performance and real time integration. There is also some improvement required in the semantic layers to manage the data match as well as the data warehouse features.
In a future release, we would like to have features to better manage all ML development activities.
I have been using this solution for three years.
This is a stable solution, especially compared to other technology on the market.
It is a scalable solution but this depends on the platform that is being used. If you use a cloud platform such as Azure, it offers scalability. However, some platforms will not support scalability using Databricks.
We have around 20 users in our development team using Databricks.
The customer service and support for this solution is good.
Positive
The initial setup is pretty simple and requires minimal configuration compared to other technology.
I would rate the pricing for this solution a four out of five. This does depend on the environment or the infrastructure that one is using. There is a difference in pricing between using Azure or being on-premises.
Azure Synapse is a competitor that we evaluated but it is not mature enough to provide better performance than Databricks. We choose Databricks due to the ability to have a lot of data in Data Lakes and the Data Warehouse. We are also able to run data science activities using ML flow.
If you are looking for custom model development and a lot of data management in a cloud agnostic manner, then Databricks is a good solution.
I would rate this solution an eight out of ten.
Our use case is confidential, but I can say we use it for a deep learning model for machine learning.
The solution is built from Spark and has integration with MLflow, which is important for our use case.
Databricks is also user-friendly, providing customizable codes and models that allow people to experiment quickly.
Integration of Delta Lake is another useful feature.
Writing pandas-profiling reports could be easier.
The ability to customize our own pipelines would enhance the product, similar to what's possible using ML files in Microsoft Azure DevOps.
I have been using this product for one and a half years.
For now the solution seems stable.
The solution is easy to scale horizontally and it has a useful auto-scaling feature. For vertical scaling, you need to bring the system down and make some adjustments.
On my current project I have a team of 30 members under me, including data engineers and data science people. Our data science, engineering, and MLOps projects are expanding, so we are planning to do some vertical scaling to increase the team size to over 100 members. In our company, we are trying to certify more and more people in Databricks because it's cloud-agnostic.
We have never needed to contact customer support, online resources have been sufficient to solve our problems.
The initial setup of the solution is straightforward, once you understand the UI it is easy to implement. I would rate Databricks a four out of five for ease of setup.
One migration project took two to three months, including writing all the code and implementing end-to-end pipelines.
We are planning to deploy the solution in stages over the next 15 months to completely implement MLOps for our organization.
I'm not involved in the financing, but I can say that the solution seemed reasonably priced compared to the competitors. Similar products are usually in the same price range. With five being affordable and one being expensive, I would rate Databricks a four out of five.
I find that deployed systems work out cheaper than having to operate manually, which appeals to our customers.
I would rate this solution an eight out of ten.
There is an issue where clusters are automatically deleted after termination or after 100 days of non-usage. This could be more user-friendly, and they could include an enabler to pin the clusters you want to keep, instead of having to go and research why clusters got deleted after implementing the product. That documentation needs to be right in front of the user to avoid issues.
I definitely recommend this product to other users.
We use Databricks to define tool data and have many use cases to analyze and distribute the data.
Data is open to everyone; they can access it through many channels, including notebooks or SQL. That on its own democratizes the data.
I like cloud scalability and data access for any type of user.
It would be better if it were faster. It can be slow, and it can be super fast for big data. But for small data, sometimes there is a sub-second response, which can be considered slow.
In the next release, I would like to have automatic creation of APIs because they don't have it at the moment, and I spend a lot of time building them.
I have been using Databricks for roughly one and a half years.
Stability is excellent.
Databricks is scalable. You can use the power of the cloud to scale your cluster size, either CPU or memory. The data doesn't work like a standard database, so you don't have it based on files, and you don't copy the data. It's super scalable. It's only the computing that you have to scale with the data.
We probably have 40 users with roles like developers, business analysts, and data scientists. We have big plans to increase the usage and have more departments using it.
Technical support has helped us.
On a scale from one to ten, I would give technical support a five.
Positive
We used Cloudera before switching to Databricks.
The initial setup was fairly okay. It takes about two minutes to deploy this solution. It's all code, so we click a button, and then it's done.
On a scale from one to five, I would give the initial setup a four.
We set up and deployed this solution.
On a scale from one to five, I would give our ROI a three.
We only pay for the Azure compute behind the solution. If you want to compute, you have to have a database layer and Azure below.
On a scale from one to five, I would give their pricing a two.
We looked at other options such as Snowflake and Cloudera on the cloud,
I would tell potential users that they need proper cloud engineers and a
cloud infrastructure team to use this solution.
On a scale from one to ten, I would give Databricks a nine.
