We use Datadog for observability and system/application health, mainly for product support, triaging, debugging, and incident responses.
We use a lot of the logging and the Datadog agent to collect logs, metrics, and traces from our GKE workloads. We use APM and continuous profiling for latency and performance measurement. We use RUM to observe frontend user events, such as tracing on request and what actions they take before errors occur. We also use error tracking and source maps to debug production failures.
We are still relatively new to the product, and we are planning to use more of the notebook functionality and power packs to record run books and break knowledge silos. We also need to utilize dashboards and continuous profiling more for performance measurement and integrate Datadog alerts for incident response.
Customizable and helpful for isolating and filtering environments
Pros and Cons
- "We have way more observability than what we had before - on the application and the overall system."
- "Since we started using the product, we were able to create dashboards, and utilize APM, continuous profiling, RUM, and distributed tracing for production support and user trends."
- "Auto instrumentation on tracing has not been very easy to find in the documentation."
What is our primary use case?
How has it helped my organization?
We have way more observability than what we had before - on the application and the overall system. That includes the GKE cluster, nodes, and pods. It's helped with our cloud-run instances, databases, and data storage.
We also started observability in the CI pipeline to measure our CI performance, as it was a pain point for us. We are aiming to do incremental deployments and releases, and the bottleneck so far has been our CI performance. The visibility on which actions or functions take the most time allows us to pinpoint and focus on improving configurations on these.
What is most valuable?
We use structure logging a lot to triage production issues. The querying, attributes and tags manipulation, and customization have been very helpful in isolating and filtering environments. The integration with Winston logger has also been a breeze.
First and foremost, was that structured logging, tags, and attributes have not only allowed us to narrow down to a problem quickly in production, they have also let us create dashboards from these logs to understand more user behaviors, such as how many users stop and leave our application before an upload has completed. That helps us understand how important processing time is to a user.
We also intend to use distributed tracing more to understand where the error has occurred in a particular request.
What needs improvement?
Definitely, documentation could use improvement. As I navigated and try to find instrumentation and implementation details, I discovered inconsistency among SDKs based on languages.
There are also places where highlighting can be improved. I once created an issue on GitHub, and it was resolved right away by an engineer. He pointed out that it was actually in the documentation. I looked again and found it was not very obvious. We were stuck on the problem for days.
Auto instrumentation on tracing has not been very easy to find in the documentation. We ended up using OpenTelemetry, yet the conversion between tracing contexts has been difficult.
Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
914,394 professionals have used our research since 2012.
For how long have I used the solution?
We've used the solution between six months and a year.
How are customer service and support?
Customer service and support are generally very fast. I did experience one ticket, which involved changing the log index retention period, not being responded to. Any support tickets related to technical issues were resolved pretty fast.
Which solution did I use previously and why did I switch?
We used to use GCP Stackdriver for logging and monitoring since our infrastructure is all GCP based. It was lacking a lot, particularly on tracing and structured logging. We often had a lot of trouble triaging and diagnosing a production problem. Datadog's specialty is observability. Since we started using the product, we were able to create dashboards, and utilize APM, continuous profiling, RUM, and distributed tracing for production support and user trends.
Datadog also offers labs and workshops for its products, which is very helpful.
What about the implementation team?
We implemented the product ourselves.
What was our ROI?
I'm not sure what our ROI would be.
What's my experience with pricing, setup cost, and licensing?
We started with on-demand pricing as we were re-writing our product, and we weren't sure about the total usage. After we went into production and released the product, we experienced a price surge. Fortunately, our Datadog account manager reached out to us and suggested a monthly subscription, which is what we'll be switching to.
I'd advise keeping an eye on the usage and possibly setting up some monitoring on price. We didn't have much of a setup cost; we started with a free trial and continued with on-demand after the trial ended.
Which other solutions did I evaluate?
We didn't evaluate many of the other options. However, we do also use OpenTelemetry, which is vendor agnostic and integrates with Datadog.
What other advice do I have?
We always keep the Datadog agent to the latest version.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Google
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Senior Manager at a manufacturing company with 10,001+ employees
Great network monitoring, testing, and integration tools
Pros and Cons
- "The visibility into our network has allowed for quick diagnosis of failures, identification of underutilized or over-utilized resources, and allowed for cloud cost optimization opportunities."
- "I would love to see more metrics or analytics in IoT devices."
What is our primary use case?
This solution is for physical device monitoring across breweries, including PLCs, HMI Cameras, RFID panels, scales, etc. We want to gain visibility into these devices to influence predictive maintenance and unscheduled downtime. We want to monitor physical devices across the zone from a control tower perspective for end users and support teams alike. Understanding more about the performance of the devices and mechanical components will allow us to schedule downtime to fix imminent catastrophic failures and prevent unplanned downtime and lost revenue.
How has it helped my organization?
Previously, we had no visibility into the architectural layout of our infrastructure. The UI of Datadog has allowed for increased visibility and access to broken or underperforming resources or critical pieces of infrastructure. Beyond this, it has allowed us to identify areas where we can optimize cost in our cloud infrastructure.
What is most valuable?
The most valuable features I have found are network monitoring, testing, and integration tools. The visibility into our network has allowed for quick diagnosis of failures, identification of underutilized or over-utilized resources, and allowed for cloud cost optimization opportunities. The ability to correlate metrics has proven useful in determining downstream or upstream issues influencing the device, machine, or database having issues.
What needs improvement?
I would love to see more metrics or analytics in IoT devices.
For how long have I used the solution?
I've been using the solution for approximately two years.
What do I think about the stability of the solution?
I have never experienced an issue or outage.
What do I think about the scalability of the solution?
The solution is very scalable and developed in a fashion that provides the ability to scale easily.
How are customer service and support?
Customer service has been outstanding. They have been timely and knowledgeable with all of my questions.
How would you rate customer service and support?
Positive
Which solution did I use previously and why did I switch?
We used a different product for the total stack solution.
How was the initial setup?
The initial setup was straightforward.
What about the implementation team?
We handled the setup process in-house.
What was our ROI?
I'm unsure as to if we've seen an ROI.
Which other solutions did I evaluate?
We did evaluate SolarWinds.
Which deployment model are you using for this solution?
Hybrid Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Microsoft Azure
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
914,394 professionals have used our research since 2012.
Software Engineer at Spring Health
Great dashboards and custom metrics with the ability to parse logs
Pros and Cons
- "The dashboards are great."
- "It is exceptionally helpful for making our engineering more data-driven."
- "We need more advanced querying against logs."
What is our primary use case?
We share dashboards, set up alerts, and monitor everything that happens in our system. We use it in staging, features, production, and our load test environment. It is exceptionally helpful for making our engineering more data-driven.
I came from a company that believes we should focus on being telemetry driven. Instilling this in a smaller, less mature engineering organization has been challenging. However, it is much easier while using Datadog.
What is most valuable?
The dashboards are great. They are an easy way to give visibility into what we need to watch with others who are not SMEs.
I enjoy the custom metrics. With this, we can take things that were once logs and then retain them longer.
We are able to parse logs. To be honest, this was only useful due to the fact that we had not yet set up the Datadog agent properly in PHP. Once we did this, the Datadog log parsing was no longer needed.
The ability to pin to a date and time is very helpful. This allows us to pinpoint exactly what was happening.
What needs improvement?
We need more advanced querying against logs. While most issues I have had here can be alleviated by way of sending better-formatted logs, it would be cool to do SQL-type queries against our data.
We need a way to see dashboard metadata. We launched a huge customer, and we saw more people using Datadog than ever across the entire organization, yet had no way to tell.
It would be ideal if we had some way to compare arbitrary date times more easily. We would love to use the Diff Graph command against some hard-coded value, for instance, against some known event.
For how long have I used the solution?
I've used the solution for eight months.
What do I think about the scalability of the solution?
The scalability is great!
Which solution did I use previously and why did I switch?
We previously used New Relic. I was not part of the decision-making team that made the switch.
What was our ROI?
The ROI is the speed at which we can debug live sites. It has been excellent. It's amazing how many incidents we can capture before customers notice.
Which other solutions did I evaluate?
We looked into New Relic and a home-brewed solution as potential other options.
Which deployment model are you using for this solution?
Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Software Engineer at Enable Medicine
Good technical documentation and overall education with improved visibility
Pros and Cons
- "We've found it most useful for managing Rstudio Workbench, which has its own logs that would not be picked up via Cloudwatch."
- "Datadog allows for much better visibility across our entire fleet and has saved us countless eng hours as a result."
- "We primarily use the log management functionality, and the only feedback I have there is better fuzzy text searching in logs (the kind that Kibana has)."
What is our primary use case?
We primarily use the solution for log monitoring across our entire cloud infra (EB, EC2, Batch, and Lambda).
This is in addition to Rstudio Workbench, which has its own logs that would not be picked up via Cloudwatch(https://docs.rstudio.com/ide/server-pro/server_management/logging.html#default-log-file-locations).
We own several dozen of these servers, and we used to manage instance logs by tailing logs when incidents occurred. Datadog allows for much better visibility across our entire fleet and has saved us countless hours.
How has it helped my organization?
It is now way easier to search in one place rather than across all of Cloudwatch (and needing to know log groups, etc.).
Primarily, we run several separate deployments of Rstudio Workbench, which has its own logs that would not be picked up via Cloudwatch.
We own several dozen of these servers. We used to manage instance logs manually.
Datadog allows for much better visibility.
What is most valuable?
We've found it most useful for managing Rstudio Workbench, which has its own logs that would not be picked up via Cloudwatch.
Datadog allows for much better visibility across our entire fleet and has saved us countless eng hours as a result.
We plan on trying out offerings such as APM moving forward too.
Some things that Datadog does very well:
- Technical documentation (the docs are clear, concise, and include realistic code samples)
- Overall education efforts (e.g. the codelabs/workshops)
What needs improvement?
We primarily use the log management functionality, and the only feedback I have there is better fuzzy text searching in logs (the kind that Kibana has).
I've learned about a ton of other offerings, like APM, NPM, etc., over the course of workshops. Once I try those out, I'm sure I will have additional feedback.
For how long have I used the solution?
I've used the solution for one year.
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Senior Engineer at a educational organization with 5,001-10,000 employees
I like the amount of tooling and the number of solutions they sold with their monitoring.
Pros and Cons
- "I like the amount of tooling and the number of solutions they sold with their monitoring. Datadog was highly intuitive to use."
- "Datadog is a complete solution with easy-to-use templates and excellent scalability."
- "Datadog needs more local Asia-Pacific support, and if they don't have a SaaS solution in Asia-Pacific, they should offer an on-prem version. I'm told that's not possible."
- "I'd rate Datadog support four out of 10. It was primarily an issue with support in the Asia-Pacific region."
What is our primary use case?
Datadog is a SaaS solution we tried for URL and synthetic monitoring. You record a transaction going into a website and replay that transaction from various locations. Datadog is mainly used by the admin, but three or four other guys had access to the reports and notifications, so it's five altogether.
We probably tried no more than 8 percent of what Datadog can do. There are so many other bits and modules. I've only gone into about half of what APM can do in the Datadog stack.
How has it helped my organization?
We could detect outages on particular websites or problems in specific locations. If I had paid for the full solution, I'm sure I could get a lot of value out of Datadog.
What is most valuable?
I like the amount of tooling and the number of solutions they sold with their monitoring. Datadog was highly intuitive to use.
What needs improvement?
Datadog needs more local Asia-Pacific support, and if they don't have a SaaS solution in Asia-Pacific, they should offer an on-prem version. I'm told that's not possible.
For how long have I used the solution?
I have used Datadog for about two or three years.
What do I think about the scalability of the solution?
I was only using Datadog to monitor on a small scale.
How are customer service and support?
I'd rate Datadog support four out of 10. It was primarily an issue with support in the Asia-Pacific region. I sent them several emails, and they responded around three weeks later.
They said it went around the houses. Nobody knew who to respond to. That's not good enough. They should have at least told me they'd received the email. I used to work in support.
How would you rate customer service and support?
Neutral
Which solution did I use previously and why did I switch?
We were just trying Datadog, and we've switched temporarily to Site24x7. We're looking for one of the bigger ones. They've all given us proposals, whereas Datadog hasn't come forward with a proposal for what they could do.
I used Datadog because I already had a relationship with them at a previous company. However, that guy's moved on now, and I wanted to see how good they were.
How was the initial setup?
Setting up Datadog is pretty straightforward. I have a lot of experience doing that sort of thing. It took maybe a day and a half to deploy because I was picking externally facing websites.
I deployed it by myself. One person is enough for the small system we had. However, if we were moving forward, I'd recommend at least two or three people to manage it.
What's my experience with pricing, setup cost, and licensing?
Datadog would've cost around $850 a month based on the loads we were doing, and you could estimate roughly what you would be paying monthly. I liked their pricing model. It was flexible, so you only paid for what you used. I rate Datadog pricing eight out of 10.
Which other solutions did I evaluate?
We looked at several URL and APM monitoring solutions like Site24x7 and Pingdom. They weren't big players like Dynatrace or any of the those that had already provided us a request for information.
What other advice do I have?
Even with our negative experiences, I'd still give Datadog an eight out of 10. Datadog is a complete solution with easy-to-use templates and excellent scalability. People should know exactly what they're going to configure before they try it out. The trial is brief. Don't start a trial until you know exactly what you're going to do.
You must be certain that you can meet any internal security requirements. If you're in the Asia-Pacific region, you might not be able to run something that's running abroad.
Which deployment model are you using for this solution?
Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Director at CBRE
Flexible, excellent support, and reliable
Pros and Cons
- "The most valuable features of Datadog are the flexibility and additional features when compared to other solutions, such as AppDynamics and Dynatrace. Some of the features include AI and ML capabilities and cloud and analysis monitoring"
- "Datadog is far better than any other monitoring tool in introducing any of the new capabilities because they think before Amazon AWS and Microsoft Azure before they introduce the concepts."
- "Datadog could improve the flexibility with AI and ML concepts. This will allow customers to be more leveraged towards publishing."
What is most valuable?
The most valuable features of Datadog are the flexibility and additional features when compared to other solutions, such as AppDynamics and Dynatrace. Some of the features include AI and ML capabilities and cloud and analysis monitoring
What needs improvement?
Datadog could improve the flexibility with AI and ML concepts. This will allow customers to be more leveraged towards publishing.
For how long have I used the solution?
I have been using Datadog for approximately one year.
What do I think about the stability of the solution?
Datadog is stable. We did not have a single outage.
What do I think about the scalability of the solution?
I have found Datadog to be scalable.
We have approximately 2,000 users using the solution in my organization.
How are customer service and support?
The support from Datadog is excellent.
Which solution did I use previously and why did I switch?
I have previously used AppDynamics and Dynatrace.
How was the initial setup?
Datadog's initial setup is easy because they have helped us come up with the easiest way of instrumenting any of the features which need to be deployed. We worked on it with their engineers and we were able to happily do it. We have done approximately 60 application monitoring through Datadog since our deployment.
What about the implementation team?
We have a very tiny team of four members that do the maintenance of Datadog.
What's my experience with pricing, setup cost, and licensing?
The price of Datadog is reasonable. Other solutions are more expensive, such as AppDynamics.
What other advice do I have?
Datadog is far better than any other monitoring tool in introducing any of the new capabilities because they think before Amazon AWS and Microsoft Azure before they introduce the concepts. Datadog is a good tool to have for monitoring your own infrastructure.
I rate Datadog a ten out of ten.
Which deployment model are you using for this solution?
Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Good alerting and issue detection for many valuable features
Pros and Cons
- "Thanks to frequent concurrent deployments, the DataDog alerts monitors allow us quickly detect issues if anything occurs."
- "The monitors can be improved."
What is our primary use case?
Our company has a microservice architecture, with different teams in charge of different services. Also, it is a start, which means that we have to build fast and move very fast as well. So before we were properly using DD, we often had issues of things breaking, but without much information on where in our system the breaking happened. This was quite a big-time sync as teams were unfamiliar with other teams' codes, so they needed the help of other teams to debug. This slowed our building down a lot. So implementing dd traces fixed this
What is most valuable?
DataDog has many features, but the most valuable have become our primary uses.
Also, thanks to frequent concurrent deployments, the DataDog alerts monitors allow us quickly detect issues if anything occurs.
What needs improvement?
The monitors can be improved. The chart in the monitors only goes back a couple of hours, clunky. Also, it can provide more info, like traces within the monitors. We have many alerts connected to different notification systems, such as Slack and Opsgenie.
When the on-caller receives notifications fired by the alerts, we are taken to the monitors. Yet often, we have to open up many different tabs to see logs, traces and info that is not accessible on the monitors. I think it would make all of the on callers' lives easier if the monitor had more data
For how long have I used the solution?
We've used the solution for three years.
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Project senior at Moka Cloud factory
An expensive solution with easy deployment
Pros and Cons
- "The tool's deployment is easy."
- "Datadog is expensive."
What needs improvement?
Datadog is expensive.
How was the initial setup?
The tool's deployment is easy.
What's my experience with pricing, setup cost, and licensing?
The solution's pricing depends on project volume.
What other advice do I have?
I rate Datadog a seven out of ten.
Disclosure: My company has a business relationship with this vendor other than being a customer. partner
Software Engineering Manager at a healthcare company with 501-1,000 employees
Great CI visibility, logging, and monitoring
Pros and Cons
- "Datadog helps us detect issues early on and helps in troubleshooting."
- "We would really like to see more from the Service Catalog."
What is our primary use case?
We mainly use the product to monitor our infrastructure and apps. It is the go-to tool when we want to check that things are running properly. We use Datadog synthetic monitors to ensure our app works across different locations in the United States.
We also have set up Datadog monitors to send alerts if things stop working as expected.
We use Continuous Integration Pipeline visibility to make sure our developers are not being blocked by infrastructure and other things that might be out of their control.
How has it helped my organization?
Datadog helps us detect issues early on and helps in troubleshooting. Creating Service Level Objectives and defining monitors is helping us to stay on top of potential issues that might affect our users.
We take advantage of Application Performance Monitoring to ensure our applications are working as expected, and our users can get the healthcare they need at a price they can afford.
Synthetic monitoring also helps us in testing our application in different browsers.
What is most valuable?
The most valuable aspects of the solution include:
CI visibility, which helps us in making sure our CI systems are running efficiently and are not blocking our developers from releasing new software and fixing bugs.
Logs, which help us in debugging issues where we can search for logs and can make sure they are relevant to the issues we are looking at.
APM, which can help us to stay on top of our applications by giving us the confidence that our apps are running.
Monitoring. We use monitoring a lot to ensure we know about potential issues and fix them before they affect our customers.
What needs improvement?
Overall, we really like the quality and relevance of all of the Datadog products that are currently being used.
The documentation is very well organized and is the go-to place for us to find answers to our questions.
We would really like to see more from the Service Catalog. It is something that we are interested in. However, some might think it lacks some key features at this time. We will definitely keep our eye out for this and adopt it when all the features are implemented.
We're really looking forward to all the great things DD will do.
For how long have I used the solution?
I've used the solution for three years.
What do I think about the stability of the solution?
The stability is great.
What do I think about the scalability of the solution?
The scalability is great.
How are customer service and support?
Technical support is great.
What about the implementation team?
We handled the initial setup in-house.
What's my experience with pricing, setup cost, and licensing?
I don't have any insights into pricing.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Amazon Web Services (AWS)
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Product Manager, Delivery Engineering at a media company with 1,001-5,000 employees
Intuitive to set up with great dashboards and dashboards and APM
Pros and Cons
- "The tools are powerful and intuitive to set up."
- "Billing should be more transparent."
What is our primary use case?
The main use case is observability and reliability as part of a platform/delivery engineering solution. We use the product to assist tenants and clients within the company to get more ramped up on SRE/DevOps.
How has it helped my organization?
The solution has provided us with a lot more insight into service-level metrics, which is especially useful with APM/tracing. It gives us all-up dashboards and alerts to assist with incident management.
What is most valuable?
The most useful aspects of the product are the dashboards and APM/tracing. The tools are powerful and intuitive to set up as well.
What needs improvement?
Custom-level metrics could be improved.
Billing should be more transparent.
For how long have I used the solution?
I've used the solution for three to four years at multiple companies.
What do I think about the stability of the solution?
The solution is very stable.
What do I think about the scalability of the solution?
The scalability is fantastic so far.
How are customer service and support?
Technical support has been great so far.
How was the initial setup?
The initial setup is straightforward. The documentation we use is very clean and concise.
What about the implementation team?
We handled the initial setup in-house.
What's my experience with pricing, setup cost, and licensing?
I would advise others to be cautious around custom metrics and be picky when setting them up.
Which deployment model are you using for this solution?
Hybrid Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Amazon Web Services (AWS)
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros
sharing their opinions.
Updated: September 2026
Product Categories
Cloud Monitoring Software Application Performance Monitoring (APM) and Observability Network Monitoring Software IT Infrastructure Monitoring Log Management Container Monitoring AIOps Cloud Security Posture Management (CSPM) AI ObservabilityPopular Comparisons
Cloudflare
Splunk Enterprise Security
Snyk
Qualys TotalCloud
Zabbix
SentinelOne Singularity Cloud Security
Dynatrace
Wazuh
Microsoft Defender for Cloud
Darktrace
IBM Security QRadar
Prisma Cloud by Palo Alto Networks
Check Point Cloud Firewall (formerly CloudGuard Network Security)
Splunk AppDynamics
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros
sharing their opinions.
Quick Links
Learn More: Questions:
- Datadog vs ELK: which one is good in terms of performance, cost and efficiency?
- Any advice about APM solutions?
- Which would you choose - Datadog or Dynatrace?
- What is the biggest difference between Datadog and New Relic APM?
- Which monitoring solution is better - New Relic or Datadog?
- Do you recommend Datadog? Why or why not?
- How is Datadog's pricing? Is it worth the price?
- Anyone switching from SolarWinds NPM? What is a good alternative and why?
- Datadog vs ELK: which one is good in terms of performance, cost and efficiency?
- What cloud monitoring software did you choose and why?


















