I am a developer, data engineer, and data analyst. I use Databricks by Carahsoft Technology Corp [Private Offer Only] to build a project for data integration for the client, starting from collecting data and making data available to be used for business intelligence purposes. This is what I do exactly. With this, I work with many different architectures, but the most one that we use is Medallion Architecture.
What is our primary use case?
How has it helped my organization?
We are now moving from a traditional database to Databricks by Carahsoft Technology Corp [Private Offer Only]. We are handling all the different kinds of data, structured and semi-structured data. With Databricks by Carahsoft Technology Corp [Private Offer Only], we solved any problem related to optimization and the time of handling data on other systems.
It has improved the data analytics. The most time that we use Spark is with Python to make the ingestion of the data. After this, we use SQL Spark. Most of the time that we use is SQL Spark for SQL warehouse. We use this one to analyze the data. It is more for analytics.
What is most valuable?
Databricks by Carahsoft Technology Corp [Private Offer Only] is a new tool that we used with Spark. We are using it to make different treatments on the data. The most important thing is that we can use Unity Catalog as a tool that can give us more privilege to access the data and share it with clients and also make different transformations. This is very simple. We work in Databricks by Carahsoft Technology Corp [Private Offer Only] with DBT and also we are using Auto Loader to get data into Databricks by Carahsoft Technology Corp [Private Offer Only] from different file shares. With Databricks by Carahsoft Technology Corp [Private Offer Only], there is the volume that we use to get external data from different data sources. We are able to get data and we can process it.
We are now using Auto Loader, which is a module that is available in Databricks by Carahsoft Technology Corp [Private Offer Only], in which we handle data that are coming to the storage. We can say it is semi-real-time data. When the data is available, the jobs are starting their compute. When the data is coming, it keeps handling the data. This was a very interesting process that we handle in Databricks by Carahsoft Technology Corp [Private Offer Only].
What needs improvement?
Last year, Databricks by Carahsoft Technology Corp [Private Offer Only] added a new module for AI in this product. It was a very interesting AI that is integrated with the data in Databricks by Carahsoft Technology Corp [Private Offer Only]. I think it is a good one. For the moment, there is some limitation, specifically when we are handling data in Databricks by Carahsoft Technology Corp [Private Offer Only] and we want to send it to another system, like Azure SQL. We need to build a different pipeline to send the data. Basically, we can create a catalog that is pointing on SQL Azure, but this one is in one direction. It is only for read. Some functionality we can add to write data from Databricks by Carahsoft Technology Corp [Private Offer Only] to another system can be a good option to add to Databricks by Carahsoft Technology Corp [Private Offer Only].
For how long have I used the solution?
I have been using Databricks by Carahsoft Technology Corp [Private Offer Only] since 2022.
What do I think about the stability of the solution?
I see sometimes some updates on some modules that can cause some crash, but it is very rare. We have some jobs that we use to make treatments of data. They are using some library that is pre-installed in Databricks by Carahsoft Technology Corp [Private Offer Only]. Sometimes there is a new version that they are updated, and in which we need to change the version of this packaging so that they can continue working. This is a kind of crash that we got with Databricks by Carahsoft Technology Corp [Private Offer Only], but it is resolved when we get the last version of the packages.
What do I think about the scalability of the solution?
Databricks by Carahsoft Technology Corp [Private Offer Only] is installed on the Azure cloud. In Databricks by Carahsoft Technology Corp [Private Offer Only], there are some CPU, some VMs that are allocated to Databricks by Carahsoft Technology Corp [Private Offer Only]. We use this resource from Databricks by Carahsoft Technology Corp [Private Offer Only] with serverless. With this one, Databricks by Carahsoft Technology Corp [Private Offer Only] can handle the time of computing between different tasks. It was very interesting to use a schedule. We let Databricks by Carahsoft Technology Corp [Private Offer Only] handle different tasks on different jobs and schedules.
How are customer service and support?
With the project that we are working on, there were some people that were allocated to the project and we got the contact with them. They are from Databricks by Carahsoft Technology Corp [Private Offer Only]. They gave us more support to resolve some issues related to data processing and optimization with non-structured data. I remember they were also providing some use cases of using limits on sharing data with other people outside of the company, how we can share it, and also using some kind of filters with RLS. This is what I remember with the point that I was checking with the Databricks by Carahsoft Technology Corp [Private Offer Only] team.
I think they are, with different projects, around eight out of ten.
Which solution did I use previously and why did I switch?
I used Palantir.
How was the initial setup?
Databricks by Carahsoft Technology Corp [Private Offer Only] is a user interface. But actually, we use CI/CD. We use a pipeline of deployment to deploy to Databricks by Carahsoft Technology Corp [Private Offer Only]. For me, it is not very hard. There are some new options to know, new technology. We go to this technology, and we learn about it and we start using it.
What's my experience with pricing, setup cost, and licensing?
Basically, the price in Databricks by Carahsoft Technology Corp [Private Offer Only] is based on the compute, the different type of compute that we are using. There are some units, unit computes. The type that we are using to handle and to process the data depends if it is job compute or it is personalized compute, customized compute. The price depends. Also, if we are using serverless as an option to handle the data. With different options that we can use for the treatment, sometimes it is very high to use serverless. It depends on the scenario, how much data we can process.
Which other solutions did I evaluate?
It depends on the client. The client shows the solution which is cheaper than the other. There are many people that can develop in the solution. For me, it is the same. There is no big difference between them. They are both a user interface that can use Spark with storage and with compute. If I work with Databricks by Carahsoft Technology Corp [Private Offer Only] or Palantir, it is the same thing.
What other advice do I have?
For data engineering and BI, I have been working for around twelve years.
It is required to use the CI/CD tools and use DBT with Databricks by Carahsoft Technology Corp [Private Offer Only] because with Databricks by Carahsoft Technology Corp [Private Offer Only] we can handle two kinds of treatment: using Spark or we can use SQL. What Databricks by Carahsoft Technology Corp [Private Offer Only] offers exactly is the compute. Compute using SQL warehouse. With this one, we can use DBT. Or we can use the servers, small compute servers, that we can use PySpark with the treatment.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Microsoft Azure
![Databricks by Carahsoft Technology Corp [Private Offer Only] Logo](https://images.peerspot.com/image/upload/c_scale,dpr_3.0,f_auto,q_100,w_100/3wz26y4eobbnb0ea9r1hctqbpq6x.png?_a=BACAGSGT)