No more typing reviews! Try our Samantha, our new voice AI agent.
Sid Nigam - PeerSpot reviewer
Works at RAPDEV LLC
User
Sep 23, 2024
Unified platform with customizable dashboards and AI-driven insights
Pros and Cons
  • "The infrastructure monitoring capabilities, especially for our Kubernetes clusters, have helped us optimize resource allocation and reduce costs."
  • "We'd like to see more advanced incident management capabilities integrated directly into the platform."

What is our primary use case?

Our primary use case for this solution is comprehensive cloud monitoring across our entire infrastructure and application stack. 

We operate in a multi-cloud environment, utilizing services from AWS, Azure, and Google Cloud Platform. 

Our applications are predominantly containerized and run on Kubernetes clusters. We have a microservices architecture with dozens of services communicating via REST APIs and message queues. 

The solution helps us monitor the performance, availability, and resource utilization of our cloud resources, databases, application servers, and front-end applications. 

It's essential for maintaining high availability, optimizing costs, and ensuring a smooth user experience for our global customer base. We particularly rely on it for real-time monitoring, alerting, and troubleshooting of production issues.

How has it helped my organization?

Datadog has significantly improved our organization by providing us with great visibility across the entire application stack. This enhanced observability has allowed us to detect and resolve issues faster, often before they impact our end-users. 

The unified platform has streamlined our monitoring processes, replacing several disparate tools we previously used. This consolidation has improved team collaboration and reduced context-switching for our DevOps engineers. 

The customizable dashboards have made it easier to share relevant metrics with different stakeholders, from developers to C-level executives. We've seen a marked decrease in our mean time to resolution (MTTR) for incidents, and the historical data has been invaluable for capacity planning and performance optimization. 

Additionally, the AI-driven insights have helped us proactively identify potential issues and optimize our infrastructure costs.

What is most valuable?

We've found the Application Performance Monitoring (APM) feature to be the most valuable, as it provides great visibility on trace-level data. This granular insight allows us to pinpoint performance bottlenecks and optimize our code more effectively. 

The distributed tracing capability has been particularly useful in our microservices environment, helping us understand the flow of requests across different services and identify latency issues. 

Additionally, the log management and analytics features have greatly improved our ability to troubleshoot issues by correlating logs with metrics and traces. 

The infrastructure monitoring capabilities, especially for our Kubernetes clusters, have helped us optimize resource allocation and reduce costs.

What needs improvement?

While Datadog is an excellent monitoring solution, it could be improved by building more features to replace alerting apps like OpsGenie and PagerDuty. Specifically, we'd like to see more advanced incident management capabilities integrated directly into the platform. This could include features like sophisticated on-call scheduling, escalation policies, and incident response workflows. 

Additionally, we'd appreciate more customizable machine learning-driven anomaly detection to help us identify unusual patterns more accurately. Improved support for serverless architectures, particularly for monitoring and tracing AWS Lambda functions, would be beneficial. 

Enhanced security monitoring and threat detection capabilities would also be valuable, potentially reducing our reliance on separate security information and event management (SIEM) tools.

Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
914,351 professionals have used our research since 2012.

For how long have I used the solution?

I've used the solution for two years.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Michael Johnston1 - PeerSpot reviewer
Senior Software Engineer at angel Studios
Real User
Sep 20, 2024
A great tool with an easy setup and helpful error logs
Pros and Cons
  • "The setup cost was minimal."
  • "We did have an issue where a synthetic test was set up before the holiday break, and we were quickly charged a great amount. Our team worked with Datadog, and they were able to help us out since it was inadvertent on our end and was a user error."

What is our primary use case?

We currently have an error monitor to monitor errors on our prod environment.  Once we hit a certain threshold, we get an alert on Slack. This helps address issues the moment they happen before our users notice. 

We also utilize synthetic tests on many pages on our site. They're easy to set up and are great for pinpointing when a bug is shipped, but they may take down a less visited page that we aren't immediately aware of. It's a great extra check to make sure the code we ship is free of bugs.

How has it helped my organization?

The synthetic tests have been invaluable. We use them to check various pages and ensure functionality across multiple areas. Furthermore, our error monitoring alerts have been crucial in letting us know of problems the moment they pop up.  

Datadog has been a great tool, and all of our teams utilize many of its features.  We have regular mob sessions where we look at our Datadog error logs and see what we can address as a team. It's been great at providing more insight into our users and logging errors that can be fixed.

What is most valuable?

The error logs have been super helpful in breaking down issues affecting our users. Our monitors let us know once we hit a certain threshold as well, which is good for momentary blips and issues with third-party providers or rollouts that we have in the works. Just last week, we had a roll-out where various features were broken due to a change in our backend API. Our Datadog logs instantly notified us of the issues, and we could troubleshoot everything much more easily than just testing blind. This was crucial to a successful rollout.

What needs improvement?

I honestly can't think of anything that can be improved. We've started using more and more features from our Datadog account and are really grateful for all of the different ways we can track and monitor our site. 

We did have an issue where a synthetic test was set up before the holiday break, and we were quickly charged a great amount. Our team worked with Datadog, and they were able to help us out since it was inadvertent on our end and was a user error. That was greatly appreciated and something that helped start our relationship with the Datadog team.

For how long have I used the solution?

We've been using Datadog for several months. We started with the synthetic tests and now use It for error handling and in many other ways.

What do I think about the stability of the solution?

Stability has been great. We've had no issues so far.

What do I think about the scalability of the solution?

The solution is very easy to scale. We've used it on multiple clients.

How are customer service and support?

We had a dev who had set up a synthetic test that was running every five minutes in every single region over the holiday break last year. The Datadog team was great and very understanding and we were able to work this out with them.

How would you rate customer service and support?

Positive

Which solution did I use previously and why did I switch?

We didn't have any previous solution. At a previous company, I've used Sentry. However, I also find Datadog to be much easier, plus the inclusion of synthetic tests is awesome.

How was the initial setup?

The documentation was great and our setup was easy.

What about the implementation team?

We implemented the solution in-house.

What was our ROI?

This has had a great ROI as we've been able to address critical bugs that have been found via our Datadog tools.

What's my experience with pricing, setup cost, and licensing?

The setup cost was minimal. The documentation is great and the product is very easy to set up.

Which other solutions did I evaluate?

We also looked at other providers and settled on Datadog. It's been great to use across all our clients.

Which deployment model are you using for this solution?

Private Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
914,351 professionals have used our research since 2012.
Tejaswini A - PeerSpot reviewer
Application Engineer at Discover Financial Services
Real User
Jul 3, 2024
Consolidates all our logs into a single place, making it easier to find errors
Pros and Cons
  • "The best way it has helped us is by consolidating all our logs into a single place and making it easier to find errors."
  • "Another issue that I have is with the search syntax, it could be simpler and it feels like there are too many ways to do the same things."

What is our primary use case?

We have a tech stack including all backend services written in TS/Node (mostly) and as a full stack engineer, it is crucial to keep track of new and existing errors. Our logs have been consolidated in Datadog and are accessible for search and review, so the service has become a daily tool for my work. 

More recently, session replay has been adopted at my company, but I do not like it so much because the UI elements are not in their place, so it is very hard to see what the users on the web app are actually clicking on.

How has it helped my organization?

The best way it has helped us is by consolidating all our logs into a single place and making it easier to find errors. Previously using AWS Cloudwatch was cumbersome and time-consuming. One issue I do have with logs is the length of time they are on the platform. Some issues happen sporadically, so it would be good to have logs for longer than one month by default or make it a configuration. 

Another issue that I have is with the search syntax, it could be simpler and it feels like there are too many ways to do the same things.

What is most valuable?

Logs search is the most valuable feature because it has consolidated all of our backend services logs into one place. Now we can see the relationship between them as requests are going from one service to other dependencies. 

What needs improvement?

One issue I do have with logs is the length of time they are on the platform. Some issues happen sporadically, so it would be good to have logs for longer than one month by default or make it a configuration. I have yet to try rehydrating logs, so this might be an option I need to try. Another issue I have is with the search syntax, it could be simpler. The syntax is a bit cumbersome and there is not an intuitive to save them to look for similar searches in the future. 

Finally, while my company replaced a different tool for session replay with DataDog's version, I find it clunky and in need of further improvements. For example, when troubleshooting a web portal issue, it is super important to know what the user clicked, but the elements are not where they should be in the replay.

It is also hard to find details about the sessions, and metadata such as user email, account, etc. that exist on other services with replay features.

For how long have I used the solution?

I have been using Datadof for approximately five years.

What do I think about the stability of the solution?

So far we haven't had any issues with uptime and Datadog has been available when needed.

What do I think about the scalability of the solution?

It seems to scale well as we continue to add services that need monitoring.

How are customer service and support?

I haven't had to contact support.

Which solution did I use previously and why did I switch?

Cloudwatch was not a great tool for what we need to do to troubleshoot issues.

What about the implementation team?

We deployed it in-house with intermediate expertise.

What was our ROI?

I am not sure how much we are paying, but I use the app often enough to feel like we are getting a good ROI.

Which other solutions did I evaluate?

I was not involved in the choosing process as a software engineer

Which deployment model are you using for this solution?

Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
reviewer2561892 - PeerSpot reviewer
Principal. Performance Engineering at Invitation Homes
User
Top 20
Oct 2, 2024
A go-to tool for analyzing, understanding, and investigating application performance
Pros and Cons
  • "Log analytics give us a powerful mechanism for error tracking, research, and analysis."
  • "Network device and performance monitoring could be improved, as we've faced some limitations in this area."

What is our primary use case?

The soluton is used for full stack enterprise performance monitoring for our primarily cloud-based stack on AWS. We have implemented monitoring coverage using RUM for critical apps and websites and utilize APM (integrated with RUM) for full stack traceability.  

We use Datadog as our primary log repository for all apps and platforms, and the advanced log analytics enable accurate log-based monitoring/alerting and investigations. 

Additionally, we some advanced RUM capabilities and metrics to track and optimize client-side user experience. We track SLO's for our critical apps and platforms using Datadog.

How has it helped my organization?

We now have full-stack observability, which allows us to better understand application behavior, quickly alert users about issues, and proactively manage application performance.  

We've seen value by implementing observability coordinated across multiple applications, allowing us to track things like customer shopping and orders across multiple applications and services.  

For critical application launches, we've built dashboards that can track user activity and confirm users are able to successfully utilize new features, tracking user activities in real-time in a war-room situation.  

Datadog is our go-to tool for analyzing, understanding, and investigating application performance and behavior.

What is most valuable?

APM accurately tracks our service performance across our ecosystem. RUM gives us client-side performance and user experience visibility, and the rate of new features implemented in the Digital Experience area recently has been high. Log analytics give us a powerful mechanism for error tracking, research, and analysis.  

Custom metrics that we've created allow us to track KPIs in real-time on dashboards. All of these have proven valuable in our organization.  Additionally, Datadog product support teams are responsive and have provided timely support when needed.

What needs improvement?

Agent remote configuration should be provided/improved and streamlined, allowing for config changes/upgrades to be performed via the portal instead of at the host.   

Cost tracking via the admin portal is a bit lacking, even though it has gotten better.  I'm looking for usage trends (that drive cost) across time and better visibility or notifications about on-demand charges.  

Network device and performance monitoring could be improved, as we've faced some limitations in this area.  

The Datadog usage-based cost model, while giving us better transparency, is difficult to follow at times and is constantly evolving.  

For how long have I used the solution?

I've used the solution for three years.

How are customer service and support?

Support has been responsive and helpful.  

How would you rate customer service and support?

Positive

What's my experience with pricing, setup cost, and licensing?

Pricing is straightforward. That said, it's sometimes difficult to estimate usage volumes.

Which other solutions did I evaluate?

We evaluated Datadog and New Relic in detail and chose Datadog due to their straightforward and competitive pricing model, and their full coverage of monitoring features that we desired, and an easy-to-use UI.  

Which deployment model are you using for this solution?

Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
SecOps Engineer at Ava Labs
User
Top 20
Sep 30, 2024
Helpful support, with centralized pipeline tracking and error logging
Pros and Cons
  • "Real user monitoring gives us invaluable insights into actual user experiences, helping us prioritize improvements where they matter most."
  • "While the documentation is very good, there are areas that need a lot of focus to pick up on the key details."

What is our primary use case?

Our primary use case is custom and vendor-supplied web application log aggregation, performance tracing and alerting. 

How has it helped my organization?

Through the use of Datadog across all of our apps, we were able to consolidate a number of alerting and error-tracking apps, and Datadog ties them all together in cohesive dashboards. 

What is most valuable?

The centralized pipeline tracking and error logging provide a comprehensive view of our development and deployment processes, making it much easier to identify and resolve issues quickly. 

Synthetic testing is great, allowing us to catch potential problems before they impact real users. Real user monitoring gives us invaluable insights into actual user experiences, helping us prioritize improvements where they matter most. And the ability to create custom dashboards has been incredibly useful, allowing us to visualize key metrics and KPIs in a way that makes sense for different teams and stakeholders. 

What needs improvement?

While the documentation is very good, there are areas that need a lot of focus to pick up on the key details. In some cases the screenshots don't match the text when updates are made. 

I spent longer than I should trying to figure out how to correlate logs to traces, mostly related to environmental variables.

For how long have I used the solution?

I've used the solution for about three years.

What do I think about the stability of the solution?

We have been impressed with the uptime.

What do I think about the scalability of the solution?

It's scalable and customizable. 

How are customer service and support?

Support is helpful. They help us tune our committed costs and alert us when we start spending out of the on-demand budget.

Which solution did I use previously and why did I switch?

We used a mix of SolarWinds, UptimeRobot, and GitHub actions. We switched to find one platform that could give deep app visibility.

How was the initial setup?

Setup is generally simple. .NET Profiling of IIS and aligning logs to traces and profiles was a challenge.

What about the implementation team?

We implemented the solution in-house.

What was our ROI?

There has been significant time saved by the development team in terms of assessing bugs and performance issues.

What's my experience with pricing, setup cost, and licensing?

I'd advise others to set up live trials to asses cost scaling. Small decisions around how monitors are used can have big impacts on cost scaling. 

Which other solutions did I evaluate?

NewRelic was considered. LogicMonitor was chosen over Datadog for our network and campus server management use cases.

What other advice do I have?

We are excited to dig further into the new offerings around LLM and continue to grow our footprint in Datadog. 

Which deployment model are you using for this solution?

Hybrid Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Operations Manager at TodayTix
User
Top 20
Sep 30, 2024
Good dashboards, easy troubleshooting, and integrations
Pros and Cons
  • "The dashboards are super convenient to us for a more zoomed out view of what is going on with each integration that we utilize."
  • "There could be more easily identifiable documentation on how to find different things on the platform."

What is our primary use case?

We utilize Datadog mainly to monitor our API integrations and all of the inventory that comes in from our API partners. Each event has its own ID, so we can trace all activity related to each event and troubleshoot where needed.

How has it helped my organization?

Datadog gives non-dev teams insights as to what all is happening with a particular event as well as flags any errors so that we can troubleshoot more efficiently.

What is most valuable?

The dashboards are super convenient to us for a more zoomed out view of what is going on with each integration that we utilize.

What needs improvement?

There could be more easily identifiable documentation on how to find different things on the platform. It can be overwhelming at first glance, and it's hard to find appropriate documentation on the site to lead you to where you need to be. 

For how long have I used the solution?

I've used the solution for about 1.5 years.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
ZJ - PeerSpot reviewer
Software Engineer at a computer software company with 201-500 employees
User
Sep 23, 2024
Very good custom metrics, dashboards, and alerts
Pros and Cons
  • "The dashboards provide a comprehensive and visually intuitive way to monitor all our key data points in real-time, making it easier to spot trends and potential issues."
  • "One key improvement we would like to see in a future Datadog release is the inclusion of certain metrics that are currently unavailable. Specifically, the ability to monitor CPU and memory utilization of AWS-managed Airflow workers, schedulers, and web servers would be highly beneficial for our organization."

What is our primary use case?

Our primary use case for Datadog involves utilizing its dashboards, monitors, and alerts to monitor several key components of our infrastructure. 

We track the performance of AWS-managed Airflow pipelines, focusing on metrics like data freshness, data volume, pipeline success rates, and overall performance. 

In addition, we monitor Looker dashboard performance to ensure data is processed efficiently. Database performance is also closely tracked, allowing us to address any potential issues proactively. This setup provides comprehensive observability and ensures that our systems operate smoothly.

How has it helped my organization?

Datadog has significantly improved our organization by providing a centralized platform to monitor all our key metrics across various systems. This unified observability has streamlined our ability to oversee infrastructure, applications, and databases from a single location. 

Furthermore, the ability to set custom alerts has been invaluable, allowing us to receive real-time notifications when any system degradation occurs. This proactive monitoring has enhanced our ability to respond swiftly to issues, reducing downtime and improving overall system reliability. As a result, Datadog has contributed to increased operational efficiency and minimized potential risks to our services.

What is most valuable?

The most valuable features we’ve found in Datadog are its custom metrics, dashboards, and alerts. The ability to create custom metrics allows us to track specific performance indicators that are critical to our operations, giving us greater control and insights into system behavior. 

The dashboards provide a comprehensive and visually intuitive way to monitor all our key data points in real-time, making it easier to spot trends and potential issues. Additionally, the alerting system ensures we are promptly notified of any system anomalies or degradations, enabling us to take immediate action to prevent downtime. 

Beyond the product features, Datadog’s customer support has been incredibly timely and helpful, resolving any issues quickly and ensuring minimal disruption to our workflow. This combination of features and support has made Datadog an essential tool in our environment.

What needs improvement?

One key improvement we would like to see in a future Datadog release is the inclusion of certain metrics that are currently unavailable. Specifically, the ability to monitor CPU and memory utilization of AWS-managed Airflow workers, schedulers, and web servers would be highly beneficial for our organization. These metrics are critical for understanding the performance and resource usage of our Airflow infrastructure, and having them directly in Datadog would provide a more comprehensive view of our system’s health. This would enable us to diagnose issues faster, optimize resource allocation, and improve overall system performance. Including these metrics in Datadog would greatly enhance its utility for teams working with AWS-managed Airflow.

For how long have I used the solution?

I've used the solution for four months.

What do I think about the stability of the solution?

The stability of Datadog has been excellent. We have not encountered any significant issues so far. 

The platform performs reliably, and we have experienced minimal disruptions or downtime. This stability has been crucial for maintaining consistent monitoring and ensuring that our observability needs are met without interruption.

What do I think about the scalability of the solution?

Datadog is generally scalable, allowing us to handle and display thousands of custom metrics efficiently. However, we’ve encountered some limitations in the table visualization view, particularly when working with around 10,000 data points. In those cases, the search functionality doesn’t always return all valid results, which can hinder detailed analysis.

How are customer service and support?

Datadog's customer support plays a crucial role in easing the initial setup process. Their team is proactive in assisting with metric configuration, providing valuable examples, and helping us navigate the setup challenges effectively. This support significantly mitigates the complexity of the initial setup.

Which solution did I use previously and why did I switch?

We used New Relic before.

How was the initial setup?

The initial setup of Datadog can be somewhat complex, primarily due to the learning curve associated with configuring each metric field correctly for optimal data visualization. It often requires careful attention to detail and a good understanding of each option to achieve the desired graphs and insights

What about the implementation team?

We implemented the solution in-house.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Senior Software Engineer at Clearstory.build
User
Sep 20, 2024
Capable of pinpointing warnings and errors in logs and provide detailed context
Pros and Cons
  • "Being able to filter requests by latency is invaluable, as it provides immediate insight into which endpoints require further analysis and optimization."
  • "The query performance could be improved, particularly when handling large datasets, as slower response times can hinder efficiency."

What is our primary use case?

Our primary use case for Datadog is to monitor, analyze, and optimize the performance and health of our applications and infrastructure. 

We leverage its logging, metrics, and tracing capabilities to pinpoint issues, track system performance, and improve overall reliability. 

Datadog’s ability to provide real-time insights and alerting on key metrics helps us quickly address issues, ensuring smooth operations. It’s integral for visibility across our microservices architecture and cloud environments.

How has it helped my organization?

Datadog has been incredibly valuable to our organization. Its ability to pinpoint warnings and errors in logs and provide detailed context is essential for troubleshooting. 

The platform's request tracing feature offers comprehensive insights into user flows, allowing us to quickly identify issues and optimize performance. 

Additionally, Datadog's real-time monitoring and alerting capabilities help us proactively manage system health, ensuring operational efficiency across our applications and infrastructure.

What is most valuable?

Being able to filter requests by latency is invaluable, as it provides immediate insight into which endpoints require further analysis and optimization. This feature helps us quickly identify performance bottlenecks and prioritize improvements. 

Additionally, the ability to filter requests by user email is extremely useful for tracking down user-specific issues faster. It streamlines the troubleshooting process and enables us to provide more targeted support to individual users, improving overall customer satisfaction.

What needs improvement?

The query performance could be improved, particularly when handling large datasets, as slower response times can hinder efficiency. 

Additionally, the interface can sometimes feel overwhelming, with so much happening at once, which may discourage users from exploring new features. 

Simplifying the layout or providing clearer guidance could enhance user experience. Any improvements related to query optimization would be highly beneficial, as it would further streamline workflows and boost productivity.

For how long have I used the solution?

I've used the solution for five years.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Mason Parry - PeerSpot reviewer
Data Engineer at Nursa
User
Sep 20, 2024
Customizable alerts, good dashboards, and improves reliability
Pros and Cons
  • "I like how we can customize alerts, and when alerts have become too noisy, we turn their threshold down fairly easily."
  • "It's not that straightforward when creating an alert. The syntax is a little confusing."

What is our primary use case?

We have several teams and several different projects, all working in tandem, so there are a lot of logs and monitoring that need to be done. We use Datadog mostly for alerting when things go down. 

We also have several dashboards to keep track of critical operations and to make sure things are running without issues. The Slack messaging is essential in our workflow in letting us know when an alert is triggered. I also appreciate all the graphs you can make, as it gives our team a good overview of how our services are doing.

How has it helped my organization?

It has improved our reliability and our time to get back up from an outage. By creating an alert and then messaging a Slack channel, we know when something goes down fairly fast. This, in turn, improves our response time to swarm on an issue without it affecting customers. The graphs have also been useful to demonstrate to higher-ups how our services are performing, allowing them to make more informed decisions when it comes to the team. 

What is most valuable?

The alerts are the most valuable. Having alerts have saved us countless times in the past and is essentially what we use data dog for. 

I like how we can customize alerts, and when alerts have become too noisy, we turn their threshold down fairly easily. This is also the case when alerts should be notifying us more often. 

I also like the graphs and how customizable they are. It allows us to create a nice-looking dashboard with all sorts of information relating to our project. This gives us a quick overview of how things are going.

What needs improvement?

It's not that straightforward when creating an alert. The syntax is a little confusing. I guess that the trade-off is customizability. But it would be nice to have a click-and-drag kind of way when creating an alert. So, if someone who isn't so familiar with Datadog or tech in general wanted to create an alert, they wouldn't need to know the syntax. 

It would also be great if AI could be used to generate alerts and graphs. I could write a short prompt, and then the AI could auto-generate alerts and graphs for me.

For how long have I used the solution?

I've used the solution for more than two years.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Felix Flores - PeerSpot reviewer
Staff Engineer at a tech services company with 1,001-5,000 employees
Real User
Oct 30, 2022
Great distributed tracing and flame graphs for debugging with a relatively painless setup
Pros and Cons
  • "We like the distributed tracing and flame graphs for debugging. This has been invaluable for us during periods of high traffic or red alert conditions."
  • "Both the engineering team and the product team are seeing tremendous value from this solution."
  • "Once Datadog has gained wide adoption, it can often be overwhelming to both know and understand where to go to find answers to questions."

What is our primary use case?

We are using a mixture of on-prem and cloud solutions to bridge the gap with healthcare entities in the service of providing patients with the medication they need to live healthy lives.

Since we're a heavily regulated company, a lot of our solutions grew from on-premises monoliths. However, as we scaled out, it became harder and harder to move forward with that architecture. Today, we're investing heavily in transforming our systems from monoliths into distributed systems.

With this change in mind, the ability for us to connect the dots using Datadog has been invaluable.

How has it helped my organization?

We have an API that serves as a critical aspect of our system for generating new requests for us to process in service of a patient. This service has many tentacles, and it was always hard to track down how issues from this API are affecting things downstream. Since we've added more instrumentation in this API, Datadog has changed our status from a reactive posture to a proactive one.

It has also served as a prime example to other applications on what the benefit of a well-instrumented system is for that application and other applications around it. Due to this, more and more people are using Datadog.

What is most valuable?

We like the distributed tracing and flame graphs for debugging. This has been invaluable for us during periods of high traffic or red alert conditions. It has also informed our developers on how our various systems are interconnected and the downstream effects of the problems we might encounter for certain services.

We're still working on getting widespread adoption of these products. Still, we're already seeing a shift in the developer's perspective from application-specific and starting to look at things from a more holistic systems perspective.

While this is not part of the question, this is relevant: Now that I've learned more about RUM, this will be something that we will heavily leverage moving forward to give us a whole complete view of our system from the front and back end perspective.

What needs improvement?

Once Datadog has gained wide adoption, it can often be overwhelming to both know and understand where to go to find answers to questions. Currently, we use a combination of documentation and COPs to ensure that folks know how to leverage what we have in Datadog properly.

While the guides for Datadog go a long way, a way to customize the user experience from "advanced" to "novice" mode would go a long way.

For how long have I used the solution?

I've been using the solution for two years.

What do I think about the stability of the solution?

It has never failed us and therefore I consider it to be very stable.

What do I think about the scalability of the solution?

It's magic. For the most part, we just installed the product and a lot of it just worked out of the box.

How are customer service and support?

Technical support is excellent.

How would you rate customer service and support?

Positive

Which solution did I use previously and why did I switch?

We have used Splunk, Sentry, and a suite of hand-made solutions. We switched since the Datadog solution was both comprehensive and cohesive. It was also easier to onboard people since the solution was well-documented and standardized.

How was the initial setup?

For the most part, it was really painless to set up.

What about the implementation team?

We implemented the solution in-house.

What was our ROI?

We're still early on in our transformation process. That said, we are gaining a lot of steam in terms of adoption. Both the engineering team and the product team are seeing tremendous value from this solution.

What's my experience with pricing, setup cost, and licensing?


Which other solutions did I evaluate?


What other advice do I have?

Adding more tooltips and links to documentation or how-tos within the application would really go a long way for those trying to get their feet wet with Datadog.

Which deployment model are you using for this solution?

Hybrid Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros sharing their opinions.
Updated: September 2026
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros sharing their opinions.