No more typing reviews! Try our Samantha, our new voice AI agent.
Binaya Moharana - PeerSpot reviewer
System Administrator And Application Support at ICE
MSP
Top 20
Apr 1, 2026
Automated incident workflows have reduced downtime and improve real-time on-call response
Pros and Cons
  • "PagerDuty Operations Cloud is the best tool available, and I can confidently say it is the best tool for all aspects, not just incident management or escalation, but for all analytical functions as well."
  • "I have observed that MTTR is very slow, and wrong escalation sometimes routes alerts to the wrong team rather than the proper team."

What is our primary use case?

PagerDuty Operations Cloud is used for production incident management to automate incidents when alerts come from our tools. When we have a critical issue and receive that alert and notification, we configure it accordingly. On-call scheduling, escalation policies, and tracking MTTR improvement are areas where we interact and interrogate with the tools. This approach reduces downtime, enables faster incident response, provides clear accountability, and improves reliability.

We support different clients including JPMorgan, Wells Fargo, PL, LPL, and Bank of America, though we are not a customer, partner, or reseller. More than 300 clients use the solution. In my company, we have 21 specialists working with PagerDuty Operations Cloud.

What is most valuable?

The best features of PagerDuty Operations Cloud are intelligent alerting, the code fixture, on-call scheduling, escalation capabilities, automation, runbooks, and integration. PagerDuty Operations Cloud has improved downtime mostly by 30 to 50%.

What needs improvement?

We are not using the autonomous AI agent in PagerDuty Operations Cloud, and we have not integrated with AI Ops, which uses machine learning to group similar alerts automatically and suggest root cause analysis from past incidents. However, we are planning to implement this functionality.

Improvement in PagerDuty Operations Cloud should focus on areas where we need to reduce alert noise by filtering unnecessary alerts. PagerDuty should send actionable alerts, and grouping and suppression should be managed from PagerDuty's side.

Alert noise and grouping do not work together seamlessly in PagerDuty Operations Cloud, but they should be consolidated. Alerts should be directed to the right person with real-time notification to ensure no critical issue is missed. Currently, when we receive alerts through call, SMS, or email, some users do not receive them, and the end client providing support to their clients may miss something important. This is the most critical feature that PagerDuty should improve.

Defining on-call scheduling and escalation by groups is also necessary. The duty roster should clearly indicate which alerts go to L1, L2, or L3 level support. Automatic escalation is not happening if nobody has responded to an alert, which I have observed.

While I did not deploy PagerDuty Operations Cloud, I performed the migration and reconfigured PagerDuty from scratch, then migrated to ServiceNow where I handled redeployment. I have observed that MTTR is very slow, and wrong escalation sometimes routes alerts to the wrong team rather than the proper team. On-call schedules should map the different teams we have, such as the application team, infrastructure team, or database team, who will take ownership. They need to align to the same team only for that incident. Proper service configuration for each application is required.

For how long have I used the solution?

I have almost three or more years of experience with PagerDuty Operations Cloud.

Buyer's Guide
PagerDuty Operations Cloud
July 2026
Learn what your peers think about PagerDuty Operations Cloud. Get advice and tips from experienced pros sharing their opinions. Updated: July 2026.
904,899 professionals have used our research since 2012.

What do I think about the stability of the solution?

The stability of PagerDuty Operations Cloud is good. I have worked with Opsgenie, and PagerDuty is better than Opsgenie and ServiceNow as well. I can give a rating of a minimum of eight.

What do I think about the scalability of the solution?

The scalability of PagerDuty Operations Cloud is also good.

How are customer service and support?

I can give PagerDuty Operations Cloud a rating of a minimum of 7.5 to eight for technical support.

Which solution did I use previously and why did I switch?

Previously, we were using PagerDuty Operations Cloud on-premises, and now we are planning to implement a hybrid approach with both cloud and on-premises solutions. I have worked with Opsgenie, and PagerDuty is better than Opsgenie and ServiceNow as well.

How was the initial setup?

Alert reduction in PagerDuty Operations Cloud does not help prevent costly incidents.

What other advice do I have?

PagerDuty Operations Cloud is the best tool available. I have never worked with other solutions, but whenever I have had the chance to work with PagerDuty Operations Cloud, there is nothing on my mind to say negatively. I can confidently say it is the best tool for all aspects, not just incident management or escalation, but for all analytical functions as well. This is the best tool in the market.

Regarding maintenance, I have not observed that part of PagerDuty Operations Cloud directly. The on-site USA team is located in Jacksonville, and some team members might be handling maintenance from their end, but I am not certain about their specific involvement.

I would recommend PagerDuty Operations Cloud to other users because it provides the best tool for instant alerting, ensuring the right person responds, automatic escalation, and reducing downtime and alert noise. It integrates with different monitoring tools including DataDog, New Relic, Prometheus, and Grafana. The main strengths are real-time alerting, proper escalation, and faster incident response, which help reduce downtime and improve MTTR.

I have experience with MTTR in PagerDuty Operations Cloud. Faster detection and alerting reduce MTTR significantly. When PagerDuty is integrated with other monitoring systems, alerts are real-time and actionable.

The end-to-end flow of PagerDuty Operations Cloud, including on-call schedules, escalation policies, and services, is excellent. It is highly reliable, offers easy escalation setup, and has strong interaction capabilities that reduce manual effort. However, the cost might be high.

Disclosure: My company has a business relationship with this vendor other than being a customer. partner
Last updated: Apr 1, 2026
Flag as inappropriate
PeerSpot user
Yarasi Harshavardhan Reddy - PeerSpot reviewer
Operations Lead at a tech vendor with 10,001+ employees
Real User
Top 10
Jun 13, 2026
Intelligent alerts have protected revenue and now drive faster incident triage with AI guidance
Pros and Cons
  • "Since PagerDuty Operations Cloud has all the data and provides forward-looking resolution steps and information about which team was involved, PagerDuty AI helps us tremendously."
  • "While PagerDuty has comment functionality, a chat option would be a potential addition."

What is our primary use case?

I have been using PagerDuty for the last nine years, but PagerDuty Operations Cloud for over one and a half years.

We work directly with merchants and need to trigger immediate alerts whenever there are 5xx errors or business errors like 4xx issues, as well as payment failures. We have configured every alert on a data log in some other monitoring tools that are integrated with PagerDuty. We receive alerts very immediately and trigger calls and Slack notifications. We integrate everything with PagerDuty and get notifications instantly, after which we start our triage process.

One use case I can mention is when we have an auth rate dip. Whenever there is an auth rate dip, we run into revenue losses with the merchants or partners that PayPal currently works with. Since everything is integrated, PagerDuty Operations Cloud catches when there is an auth rate dip for particular merchants and immediately triggers a notification for us. We then immediately dive into what the problem is and figure out how to fix the issue with the help of engineering teams.

What is most valuable?

PagerDuty Operations Cloud is one of the best tools we have seen because it is already integrated with AI. We use it as a barrier tool, meaning it is the top tool that we consider and we get notified when there is an issue.

The best features include integrating with any tool and analyzing all previous alerts that have been stored. When an alert occurred on a particular day, we can immediately be notified on Slack with historical data and, since it is integrated with AI, we receive suggestions on how it can be resolved, how it was resolved earlier, and who resolved it. These are the very best features we have seen on PagerDuty Operations Cloud.

Since we have historical data showing when an alert has triggered on a particular day, we can turn it into a problem incident and work with the relevant teams to get it fixed completely so it does not reoccur. We are recording these kinds of repetitive issues using that feature.

It is very helpful that we can integrate with numerous monitoring tools such as Datadog, Splunk, and Kibana. Since we have integrated many other tools, I feel this is one of the features that PagerDuty Operations Cloud offers that makes it great.

What needs improvement?

Since PagerDuty Operations Cloud is already equipped with the latest technologies, I do not feel that anything more needs to be added, including summarizing content, as it is already available. Since it is already connected with AI, I do not feel that any other features could be added, so I do not have a concrete answer right now since we already have a number of features available and this is already a highly improved state.

While PagerDuty has comment functionality, a chat option would be a potential addition.

For how long have I used the solution?

I have been using PagerDuty for the last nine years, but PagerDuty Operations Cloud for over one and a half years.

What do I think about the stability of the solution?

PagerDuty Operations Cloud is highly accurate and there are no issues with the accuracy. It is highly reliable in terms of alert triggering and we do not get any false alarms, with only very minimal ones based on our internal signals. We do not have any complaints about PagerDuty Operations Cloud.

What do I think about the scalability of the solution?

PagerDuty Operations Cloud definitely increases efficiency for us. Since we do not have much manual work with workflows and everything is automated, it definitely helps.

Which solution did I use previously and why did I switch?

We are using only PagerDuty and do not have any other tool in use. There is no other tool that can match PagerDuty Operations Cloud.

What was our ROI?

We definitely have an ROI in terms of earlier requiring multiple employees. Since we are now using AI, we have reduced our staffing needs and can save a lot of time and money as well.

Which other solutions did I evaluate?

There is no other tool that can match PagerDuty Operations Cloud.

What other advice do I have?

Earlier, PagerDuty Operations Cloud was just notifying incidents, but now it is showing historical data and we can see how it was resolved earlier and quickly get notes from that to resolve issues with the historical data and suggestions.

Earlier, when there was an auth rate dip or different signals that we received through Datadog or different platforms, we used to have some false alarms. Now, everything we are using is AI-based with agents that were configured with those signals. We have very accurately configured the AI using factors such as holiday seasons that will have high traffic, and everything was configured with historical data. We are getting very solid results and signals.

Since PagerDuty Operations Cloud has all the data and provides forward-looking resolution steps and information about which team was involved, PagerDuty AI helps us tremendously.

We definitely do not have any revenue loss since we are getting accurate signals and alerts and have a solution for all configured alerts.

Since it has all advanced features integrated with AI, I am really impressed with the ability to integrate with numerous monitoring tools very easily and the ease of onboarding any member to PagerDuty Operations Cloud. Setting up the alerts and everything is very easy with a number of monitoring tools. That is why I rated this product a nine out of ten. There is no other tool that can match PagerDuty Operations Cloud right now.

We have a number of layers in terms of governance and security since we are a payment gateway. PagerDuty Operations Cloud has its own governance and security at a great level, so we do not need to think about any security concerns from PagerDuty Operations Cloud governance.

Since it already has AI features, I am going to recommend others to use PagerDuty Operations Cloud. I rate this solution a nine out of ten.

Which deployment model are you using for this solution?

Private Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Amazon Web Services (AWS)
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: Jun 13, 2026
Flag as inappropriate
PeerSpot user
Buyer's Guide
PagerDuty Operations Cloud
July 2026
Learn what your peers think about PagerDuty Operations Cloud. Get advice and tips from experienced pros sharing their opinions. Updated: July 2026.
904,899 professionals have used our research since 2012.
Patel Dhulva - PeerSpot reviewer
Software Firmware Engineer at Kohler Co.
Real User
Top 5
May 30, 2026
AI-driven incident management has reduced downtime and improves focus on strategic work
Pros and Cons
  • "PagerDuty Operations Cloud has positively impacted my organization by enabling faster issue response, which helped reduce downtime, saved revenue by avoiding long outages, improved team accountability during incidents, reduced manual effort in handling alerts, and helped maintain a better customer experience."
  • "PagerDuty Operations Cloud needs improvements because sometimes integrations are not very seamless and misbehave."

What is our primary use case?

PagerDuty Operations Cloud is a multifunctional digital operations platform that meets my organization's needs.

I am impressed by this digital operations solution because it is the most appropriate tool for incident detection and alerting.

PagerDuty Operations Cloud is a very user-friendly tool, highly accurate, and an easy-to-customize digital operations management system that suits my organization's needs.

It has intelligent noise reduction capabilities that play a significant role in minimizing alert floods.

What is most valuable?

PagerDuty Operations Cloud offers top-tier features that enable real-time alerting and accelerate incident response.

The solution is reliable and effective when it comes to automating routine diagnostic tasks.

Regarding how the real-time alerting and automation features have helped my team, problem-solving became automatic, and incident management becomes less complex to manage.

PagerDuty Operations Cloud has positively impacted my organization by enabling faster issue response, which helped reduce downtime, saved revenue by avoiding long outages, improved team accountability during incidents, reduced manual effort in handling alerts, and helped maintain a better customer experience.

The solution's alert reduction feature has had a major impact on preventing costly incidents in my organization. By grouping related alerts and de-duplicating noise, my team was able to spot real issues faster instead of getting buried in alerts, helping us prevent two to three potential outages because engineers responded to the root alert instead of missing it in noise.

What needs improvement?

The user interface should be easier to customize and use.

The pricing could be less expensive, especially for smaller organizations.

The user interface could be made easier to customize and navigate so that users who are new to this platform find the learning curve smoother.

PagerDuty Operations Cloud needs improvements because sometimes integrations are not very seamless and misbehave.

For how long have I used the solution?

I have been using PagerDuty Operations Cloud for about one year and a few months.

What other advice do I have?

PagerDuty Operations Cloud is a great operational efficiency tool, not just for paging.

It is very cost-effective, especially for organizations that are not limited by budgets.

PagerDuty Operations Cloud solves a lot of problems.

For example, if any issue arises during our online exam with our client, then PagerDuty Operations Cloud alerts the right team and the right people, and tasks are assigned so those problems can be resolved at the correct time and our real task does not get disrupted.

PagerDuty Operations Cloud's AI functionality has improved my team's ability to focus on core tasks rather than routine issues by removing routine alert triage.

The AI groups and de-duplicates alerts automatically, so our engineers are not manually sorting through twenty duplicate notifications for one root issue, allowing them to save a lot of time and focus on other strategic tasks, which improves productivity in my organization.

We are using PagerDuty Operations Cloud's autonomous AI agents for low-severity incidents, which automatically triage, correlate, and resolve known issues without human intervention, such as restarting services or acknowledging flapping alerts.

This has contributed to efficiency by cutting manual workload by thirty-five percent and also reducing MTTR for routine incidents.

The effectiveness of PagerDuty Operations Cloud's generative AI in providing insights for decision-making is effective during incidents.

The AI provides clear insights through incident summaries and what-changed analysis, helping us decide where to start troubleshooting instead of guessing, enabling us to make data-driven decisions easily, and providing actionable insights that improve response decisions.

The influence of PagerDuty Operations Cloud's embedded AI on revenue protection in terms of reducing alert fatigue and incident costs has a positive impact by reducing downtime risks and operational costs per incident.

I would rate this review nine out of ten.

Which deployment model are you using for this solution?

Public Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: May 30, 2026
Flag as inappropriate
PeerSpot user
Navin Samuel - PeerSpot reviewer
Delivery Manager at Cognizant
Real User
Top 20
May 25, 2026
Modern alert automation has boosted noc efficiency and now streamlines on-call workflows
Pros and Cons
  • "Overall, it was a very effective product and really helped the productivity of the teams."
  • "Initially it was a nightmare. There were a lot of bugs and we were getting a lot of alerts, and it was a very messy period."

What is our primary use case?

My network operation center team monitors network alerts and digital alerts and pages engineers on-call. My team utilized this tool for supporting the customer Nike, and I led the NOC operations at Cognizant. We had various sources like Splunk, New Relic, VeloCloud, and SolarWinds triggering alerts for device failures or high CPU utilization. PagerDuty Operations Cloud correlates all events from multiple sources and triggers alerts to the respective teams based on the orchestration set for the particular event. My team handles alerts through automated processes or through manual intervention and resolves incidents. They also use it to trigger on-call engineers through SMS or email.

When we receive alerts that do not need manual intervention for 15 to 20 minutes, we have automation in place to monitor that alert and resolve it if the issue is remediated automatically. We also had automated functions to check the reachability of devices where we have on-demand ping commands in PagerDuty Operations Cloud. Event correlation helps to  action bulk alerts in a single incident. For example, if one location is triggering alerts for multiple devices, all these alerts are correlated in a single incident, and the engineer will action and resolve the alerts.

We were getting a lot of alerts on a daily basis that do not require immediate intervention. We automated those alerts to be in a different queue where PagerDuty Operations Cloud itself monitors and resolves the alert. We also implemented agentic AI use cases for device interface alerts where the alerts are triggered and not immediately actioned by humans. PagerDuty Operations Cloud monitors and performs certain actions using agentic AI. If agentic AI is not able to clear the alerts, it creates an incident and ticket to the respective engineering team. If it is able to clear the issue, it automatically resolves the alert. This automation helped in reducing a lot of noisy alerts without manual intervention, and without it, a lot of manual intervention would have been required for engineers to perform basic triage.

What is most valuable?

PagerDuty Operations Cloud definitely helped us. In terms of governance, it had very good reports regarding the alerts we were receiving and calculating the times and effort. Overall, it was a very effective product and really helped the productivity of the teams.

PagerDuty Operations Cloud has correlation features which help us when we get a lot of alerts from one particular location. Based on the correlation values set for the alerts, it automatically correlates all the incidents. It also attaches the KB articles with troubleshooting steps

What needs improvement?

PagerDuty Operations Cloud was at a premature stage and we were not utilizing the GenAI features. I was using it for nearly two years, and we migrated from a different tool to PagerDuty Operations Cloud. Many of these features were in the development stage, so I cannot comment on its capability when it comes to GenAI.

We were not fully utilizing PagerDuty Operations Cloud capabilities as our environment was different. We were just migrating to PagerDuty Operations Cloud and it had been nearly two years. The first year went fully into migrating the product to the environment and fixing bugs, and the next year went into development and stabilizing the product in the customer environment. By the time we were planning to use AI and all the advanced features, I left the project for a different engagement and we do not use PagerDuty Operations Cloud in this current environment.

Initially it was a nightmare. There were a lot of bugs and we were getting a lot of alerts, and it was a very messy period. We had to manually keep track of all the alerts and share the feedback to get it sorted out. The user interface that was provided in the initial stages was not friendly, and it was very difficult to manage the alerts. In later stages, based on feedback, they improved the customization options and user interface, so initially it was not good.

For how long have I used the solution?

I was using it for nearly two years.

What do I think about the stability of the solution?

PagerDuty Operations Cloud is reliable. We have faced only one or two outages in a span of two years and have not had a lot of outages that disrupt operations. I will not say it is 100% perfect, as we do run into bugs and disruptions. Overall, PagerDuty Operations Cloud is a reliable product.

What do I think about the scalability of the solution?

PagerDuty Operations Cloud is highly scalable as it can be integrated with multiple platforms and multiple ticketing tools, and with automation and AI features available, it is highly scalable.

How are customer service and support?

PagerDuty Operations Cloud support was good compared to BigPanda. We had a better experience with PagerDuty Operations Cloud. There were some features that we requested and which they delayed in delivering, but they had their reasoning. Overall, the support by PagerDuty Operations Cloud was good. They had a monthly cadence call and a bi-weekly call with the tools and support team to address concerns on a weekly basis. PagerDuty Operations Cloud is a good product for the organization and the support team is highly effective and responsive.

Which solution did I use previously and why did I switch?

We used a tool called BigPanda. BigPanda was very buggy and their support teams were not responding on time to some critical issues. It had issues in integrating with other tools like ServiceNow and we had a lot of issues with reporting. Reporting was not accurate and there were a lot of inaccurate data in the reporting, so we had difficulties in governing the team's performance. The customer decided that it was better to migrate to a different tool. We had PagerDuty Operations Cloud even during that period, just to page on-call engineers, but not as an event correlation platform tool. They provided a testing period for us and we decided to switch over to PagerDuty Operations Cloud rather than holding onto BigPanda.

What other advice do I have?

PagerDuty Operations Cloud is already up to date with the requirements in terms of cloud automation and AI enhancement. It is a very modern solution for NOC operations and for paging on-call engineers. It has all the orchestration and automation functionalities required to perform triages for all types of incidents, and it can correlate using intelligence automatically. I do not see a lot of room for improvements, but there were some bugs that we were working on with the vendor. Overall as a product, PagerDuty Operations Cloud is a very modern solution for NOC environments.

The user interface was very easy in PagerDuty Operations Cloud. There were some operation-related bugs that were not due to some configurations and challenges with integrating the tool with the customer environment. Overall as a product by default, it had a very user-friendly user interface.

I would recommend PagerDuty Operations Cloud as a good product for the organization. I gave this product a rating of 8.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: May 25, 2026
Flag as inappropriate
PeerSpot user
Akhil Viswam - PeerSpot reviewer
Senior Consultant at a consultancy with 10,001+ employees
Real User
Top 5Leaderboard
Dec 7, 2025
Runbook automation has reduced incident response time and now improves uptime and collaboration
Pros and Cons
  • "My advice to others looking into using PagerDuty Operations Cloud is that it is one of the best tools in the market for production support and SRE engineers."
  • "One suggestion for improving PagerDuty Operations Cloud is to provide more insights about incidents, such as root cause analysis or additional information, which could assist SRE teams in reducing remediation time and incident detection before jumping on a call."

What is our primary use case?

Our main use case for PagerDuty Operations Cloud is for alerting purposes whenever any kind of downtime or downstream incident happens with our application which causes any downtime, and PagerDuty Operations Cloud will alert us through calls and SMS so we can get notified and quickly remediate the issue.

A unique aspect of our main use case with PagerDuty Operations Cloud is using the Runbook flow. Whenever we experience a specific kind of incident, the Runbook will trigger automation to either remediate the issues or perform root cause analysis, thus enhancing our workflow automations.

What is most valuable?

PagerDuty Operations Cloud helps our team respond by increasing our response time. Whenever there is any incident, we will get notified and through PagerDuty Operations Cloud, we receive calls 24/7, allowing us to instantly get into a call or investigation and remediate the issue as early as possible. This way, PagerDuty Operations Cloud helps us reduce the MTTR and ensures our application is more reliable and resilient.

We have been using the Runbook automation feature for building automated flows that help us add extra monitoring for specific alerts or incidents and perform remediation tasks autonomously using this Runbook flow.

One feature I particularly appreciate about PagerDuty Operations Cloud is that it offers multiple notification options. I receive alerts via call as well as SMS, which is beneficial. If I miss the call, I may still receive the SMS and vice versa.

Through PagerDuty Operations Cloud, our MTTR has been reduced by at least 30% over the last year due to its instant notification features like SMS and calls, which help us jump on calls quickly to remediate issues. This reduction has impacted our application downtime, ensuring an uptime of approximately 99% throughout the year.

What needs improvement?

One suggestion for improving PagerDuty Operations Cloud is to provide more insights about incidents, such as root cause analysis or additional information, which could assist SRE teams in reducing remediation time and incident detection before jumping on a call.

From an integration point of view, everything is functioning well. However, we primarily use the desktop interface as our main tool, and adding more details on incidents directly from PagerDuty Operations Cloud's analysis would enhance the user experience.

For how long have I used the solution?

I have been using PagerDuty Operations Cloud for the last three years.

What do I think about the stability of the solution?

PagerDuty Operations Cloud is absolutely stable. We have never experienced any downtime or latency issues from PagerDuty Operations Cloud.

What do I think about the scalability of the solution?

We don't have much insight on scalability, as a separate enterprise PagerDuty Operations Cloud team is responsible for handling all scaling activities.

How are customer service and support?

We have internal enterprise support within the application, which is very interactive. They escalate issues to the external PagerDuty Operations Cloud team when necessary, and they are very supportive.

How would you rate customer service and support?

Which solution did I use previously and why did I switch?

We have not previously used a different solution. PagerDuty Operations Cloud is the first alerting tool I have been using since the beginning.

How was the initial setup?

PagerDuty Operations Cloud onboarding is pretty straightforward in our organization, as new candidates simply need to be part of specific Windows AD groups to complete the onboarding process and gain access.

What about the implementation team?

There are automations in our organization that connect PagerDuty Operations Cloud to other ticketing tools such as Jira and ServiceNow. Whenever an incident occurs, automation that uses the Runbook flow triggers to extract data from the PagerDuty Operations Cloud alert to create incidents and Jira tickets for the development team.

What was our ROI?

In terms of return on investment, we have reduced our MTTR by 30% in the last year, indirectly improving our application's uptime to nearly 99%, which enhances client experience and boosts our business.

What's my experience with pricing, setup cost, and licensing?

I have no personal experience with pricing, setup costs, or licensing, as a separate enterprise PagerDuty Operations Cloud team manages those processes.

What other advice do I have?

The escalation policies within PagerDuty Operations Cloud are user-friendly and customizable, allowing us to set up multi-level escalations from SRE engineers to SRE leads and then to management.

PagerDuty Operations Cloud helps our team collaborate during incidents by automatically updating incident status based on progress. We have alerting integrated with Slack for this, where incidents show as red when active, yellow when acknowledged, and green when resolved.

Regarding performance metrics, there is a dedicated enterprise PagerDuty Operations Cloud team that handles monitoring, so as an SRE, I don't need to manage these performance aspects myself.

My advice to others looking into using PagerDuty Operations Cloud is that it is one of the best tools in the market for production support and SRE engineers. It is essential for our operations, functioning as our bread and butter.

We have covered almost everything regarding PagerDuty Operations Cloud. It has been a great tool for SRE and production support teams, and we look forward to more features, especially with trending technologies like AI. I would rate this product an 8 out of 10.

Which deployment model are you using for this solution?

On-premises

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Other
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: Dec 7, 2025
Flag as inappropriate
PeerSpot user
Gulam Gauss - PeerSpot reviewer
Cloud Support Engineer at ATOS
MSP
Top 20
May 28, 2026
Real-time monitoring has ensured proactive issue resolution across critical cloud environments
Pros and Cons
  • "Using PagerDuty Operations Cloud is a very crucial part for us; if we do not use it, we don't know what is happening in the customer's environment."
  • "If there is an outage on PagerDuty's side, we sit idle waiting for it to be resolved."

What is our primary use case?

We use PagerDuty Operations Cloud for monitoring purposes, as we have our own agent installed in the customer's environment across AWS and Azure, virtual machines or instances. Our monitoring agent continuously monitors the customer's environment for high CPU utilization, high memory utilization, disk utilization, and if there is any problem within the customer's environment, then we get an alert on PagerDuty. After receiving the alert, we have to acknowledge it or resolve it. If the issue is still there, then we have to work on that alert and resolve the issue, so that is part of PagerDuty.

Using PagerDuty Operations Cloud is a very crucial part for us; if we do not use it, we don't know what is happening in the customer's environment. It plays a crucial role in our organization because we have no other option for receiving any kind of information from the customer's environment. If there is an outage on PagerDuty's side, we sit idle waiting for it to be resolved. We reach out to PagerDuty's team to sort out the issue quickly, as some of our customers are very important and we have to monitor their environments 24/7 because they do not want any problems.

What is most valuable?

I appreciate the notifications in PagerDuty Operations Cloud. We get real-time updates in PagerDuty, and that is the best part. The second feature is that we get each and every detail in PagerDuty as well, including the main problem, which account is affected, and in which account, which instance or which VM is affected. We also get graphs in PagerDuty Operations Cloud, making it so that 50% of our work is done by PagerDuty itself, with the rest of the 50% being the work we have to do.

What needs improvement?

The main part I would see improved or enhanced in PagerDuty Operations Cloud is the absence of a PagerDuty notifier in Chrome. Someone created a notifier in the Chrome extension, but it was removed after an update to the Chrome version. My request is for PagerDuty's team to create a notifier so we get pop-up alerts whenever we receive any kind of alert. We currently use the notifier in Microsoft Edge, which works fine, but we need it in Chrome as well, and I suggest it should have options for acknowledging, holding, and resolving alerts, as that is very important.

Enhancements are not needed for me except for having PagerDuty Operations Cloud's notifier from the official PagerDuty team.

For how long have I used the solution?

I have been working with PagerDuty Operations Cloud for more than five years.

What do I think about the stability of the solution?

Regarding the stability aspect of PagerDuty Operations Cloud, I am aware of problems we faced during the CrowdStrike incident last year. We encountered multiple issues, but after degrading our customer's systems, the problems were resolved once CrowdStrike released a new patch, normalizing the process.

What do I think about the scalability of the solution?

I think PagerDuty Operations Cloud is scalable, and I have not encountered any limitations. We can configure it as needed, and I receive detailed information through PagerDuty. As I mentioned earlier, 50% of our work is already handled, and the remaining part requires us to log in to the customer's environment to check what we can do.

How are customer service and support?

I have only interacted with PagerDuty's technical support team during the CrowdStrike incident. They worked very efficiently and responded on time. We also have their contact numbers, emails, and technical assistance email IDs, so whenever we need to modify anything like increasing or degrading resources, we reach out to them without any problem.

On a scale of one to ten, I would rate PagerDuty's technical support as nine out of ten.

Which solution did I use previously and why did I switch?

We have been using PagerDuty Operations Cloud from the initial phase; there was no previous solution.

How was the initial setup?

The initial setup process for PagerDuty Operations Cloud involves creating a user first with our official email ID. After that, we select which team we are working in, and as we're operating 24/7, we add on-call availability, with three engineers assigned for different shifts, which means splitting 24 hours into three 8-hour shifts.

After adding the user, we have to specify our name and the shift timings during which we are working. We also select our teams, whether L1, L2, or L3, based on that, we receive alerts from the customer's environment. That is the basic process we use.

What about the implementation team?

We handle the deployment in-house for PagerDuty Operations Cloud.

What was our ROI?

In terms of measurable benefits from PagerDuty Operations Cloud, it is cost-effective and saves resources, as we receive alerts with real-time data. If the alert is resolved in the customer's environment, then PagerDuty Operations Cloud alerts also get resolved in real time. So far, I have not encountered any errors in PagerDuty Operations Cloud system; this is a very good thing for us.

What's my experience with pricing, setup cost, and licensing?

I don't know anything about the pricing, setup costs, or licensing part as we have a different team for that.

Which other solutions did I evaluate?

I have no idea about evaluating other options available in the market as our company decided to use PagerDuty Operations Cloud based on recommendations from contacts who suggested we use it.

What other advice do I have?

My advice to other organizations considering PagerDuty Operations Cloud is to focus on real-time updates, as we have configurable options. Their support is very timely, so we receive prompt answers to any queries. My only concern is that PagerDuty needs to develop their notification system to provide similar notifications on Windows or macOS as we get with their app on Android or iOS. I would rate this product a ten out of ten overall.

Which deployment model are you using for this solution?

Public Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Disclosure: PeerSpot contacted the reviewer to collect the review and to validate authenticity. The reviewer was referred by the vendor, but the review is not subject to editing or approval by the vendor.
Last updated: May 28, 2026
Flag as inappropriate
PeerSpot user
DeepakReddy - PeerSpot reviewer
Sr.Devops engineer at Scaler
Real User
Top 5Leaderboard
Apr 8, 2026
Centralized incident response has reduced downtime and now needs more predictable costs
Pros and Cons
  • "PagerDuty Operations Cloud has positively impacted our organization by accelerating incident response and reducing MTTR by up to 27%, and we are also integrating AI and ML into PagerDuty Operations Cloud."
  • "One area for improvement in PagerDuty Operations Cloud is the unpredictable costs that can cause issues in our organization and project complexity, along with the occasional perception of an outdated user interface by non-tech personnel."

What is our primary use case?

My main use case for PagerDuty Operations Cloud is for cloud-based operations, including incident management and resolving incident responses to reduce downtime and improve reliability.

PagerDuty Operations Cloud provides a central command center that collects data signals from various IT systems, which helps us detect high-priority incidents and reduce the noise.

In addition to my main use case, we are able to perform on-call scheduling and routing with the help of PagerDuty Operations Cloud very easily.

What is most valuable?

The best features PagerDuty Operations Cloud offers include automated on-call scheduling, AI-driven alert grouping, and an impressive number of integrations, as it has more than 700 integrations, which really help us.

The user interface is user-friendly for non-technical persons, and it has been maintained well over the years, making it easier for our project managers and product managers to navigate PagerDuty Operations Cloud dashboards.

PagerDuty Operations Cloud's embedded AI has helped reduce alert fatigue and lower costs from incidents, which has contributed to retaining revenue by minimizing financial losses.

The AI-driven alert grouping, particularly for incident management, helps us significantly as it streamlines our processes.

PagerDuty Operations Cloud has positively impacted our organization by accelerating incident response and reducing MTTR by up to 27%, and we are also integrating AI and ML into PagerDuty Operations Cloud.

What needs improvement?

One area for improvement in PagerDuty Operations Cloud is the unpredictable costs that can cause issues in our organization and project complexity, along with the occasional perception of an outdated user interface by non-tech personnel.

For how long have I used the solution?

I have been using PagerDuty Operations Cloud for around three years.

What do I think about the stability of the solution?

PagerDuty Operations Cloud is stable and scalable, capable of handling enterprise environments and multi-cloud setups efficiently.

How are customer service and support?

I would rate customer support a seven out of 10.

Which solution did I use previously and why did I switch?

Previously, we were using incident.io, AlertOps, and DataDog, but we switched to PagerDuty Operations Cloud due to its all-in-one solution capabilities.

What about the implementation team?

I have implemented AI and automation through PagerDuty Operations Cloud for incident response, which has significantly changed our operational efficiency, allowing us to accomplish more with less manual input.

I have experimented with PagerDuty Operations Cloud's autonomous AI agents, striving to automate repetitive tasks, which improves operational efficiency.

What was our ROI?

While I cannot provide exact return on investment metrics, I estimate that we save around 20% of costs and approximately 10 to 15 hours a week with the efficient use of PagerDuty Operations Cloud.

The 27% reduction in MTTR has had a direct business impact as it enhances business growth and supports impactful business decisions.

The alert reduction feature has greatly impacted our ability to prevent costly incidents, as we can accurately respond to alerts with the help of autonomous AI agents, which reduces erroneous notifications.

What's my experience with pricing, setup cost, and licensing?

My experience with pricing, setup cost, and licensing has been positive compared to other tools, as PagerDuty Operations Cloud simplifies many management tasks that would otherwise be burdensome.

Which other solutions did I evaluate?

Before choosing PagerDuty Operations Cloud, I evaluated previous tools such as incident.io and DataDog, but we selected PagerDuty Operations Cloud for its comprehensive features and strong alerting.

What other advice do I have?

My advice for others considering PagerDuty Operations Cloud is to first understand their organizational needs before selecting any tool to prevent mismatches.

I believe all aspects of my experience with PagerDuty Operations Cloud have been covered.

I would rate my overall experience with PagerDuty Operations Cloud a seven out of 10.

Which deployment model are you using for this solution?

Hybrid Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Amazon Web Services (AWS)
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: Apr 8, 2026
Flag as inappropriate
PeerSpot user
Saurab Gnagurde - PeerSpot reviewer
IT Analyst | Aws Cloud Ops | Dev Ops | Fin Ops at Tata Consultancy
Real User
Top 5Leaderboard
Jun 9, 2026
On-call automation has reduced critical incident impact and ensures faster production responses
Pros and Cons
  • "PagerDuty Operations Cloud helps us manage production incidents beyond service outages, mostly high CPU utilization where we set alerts, application failures, pod issues in Kubernetes, and infrastructure-related alerts."
  • "One area where I believe improvement can be made is reporting and dashboard customization to make it more user-friendly."

What is our primary use case?

As a cloud operation team, I was a user who set the alerts, and whatever important incidents or anomalies were detected that needed to be immediately taken care of were bifurcated through our APM tools that we integrated with PagerDuty Operations Cloud. As a cloud operation team, we supported the platform for rotational shifts. My roles involved setting the person in the shift according to the shift roster, so whenever any incidents triggered, they would get the call. The primary use was supporting production operations and cloud activities.

Our multi-environment consists of AWS infrastructure, Linux servers, Kubernetes clusters, and customer-facing applications. PagerDuty Operations Cloud was mainly used for incident management and alerting. We integrated it with AppDynamics, Instana, and CloudWatch, where it would monitor the patterns and platform, and then PagerDuty Operations Cloud would generate the critical alerts that the appropriate support team who was working in that present shift would get notified of immediately. This platform really helped us manage production incidents beyond service outages, mostly high CPU utilization where we set alerts, application failures, pod issues in Kubernetes, and infrastructure-related alerts. We configured all kinds of alerts, which ensured that alerts were routed to the correct on-call person, helping us reduce response time in critical situations.

What is most valuable?

One of the best features I would mention about PagerDuty Operations Cloud is its on-call rotational scheduling support and escalation management practices. If an engineer did not acknowledge the alert within a defined time frame, the incident was automatically escalated to the next person, support team, or manager of that specific team. Another useful feature was its integration capability. We were able to integrate PagerDuty Operations Cloud with monitoring and observability tools that allow alerts to generate automatically whenever issues were detected in the environment within a fraction of time. We also had the mobile application that was very helpful because the engineer could receive calls, notifications, and acknowledge the incident and track the updates even when they were away from their laptop.

I also valued the centralized incident management dashboard that provides visibility into active incidents, response status, escalation history, and overall operational health. I used to get all the data accumulated there through the dashboard.

PagerDuty Operations Cloud helps us manage production incidents beyond service outages, mostly high CPU utilization where we set alerts, application failures, pod issues in Kubernetes, and infrastructure-related alerts.

What needs improvement?

My experience with PagerDuty Operations Cloud has been positive overall. One area where I believe improvement can be made is reporting and dashboard customization to make it more user-friendly. The operations team often requires different views compared to the management team. Having more flexibility in generating custom reports would be helpful. Another improvement could be providing more advanced AI-driven collaboration capabilities to reduce unnecessary noise alerts and help the team focus on the most critical issues. Apart from these areas, the platform is very reliable and effective for managing production incidents and on-call operations.

For how long have I used the solution?

I have been using PagerDuty Operations Cloud for almost five to six years.

What do I think about the stability of the solution?

PagerDuty Operations Cloud has been stable and performing well wherever our incident management or alerting was configured for production support. Timely notifications and incident responses were critical. PagerDuty Operations Cloud delivers alerts immediately through multiple channels which we configured, including mobile on-call notifications, email, SMS, and phone calls. Since PagerDuty Operations Cloud was integrated with our monitoring and observability tools, it helped ensure that critical incidents were captured and routed to the appropriate on-call team. During my usage, I did not encounter any significant outages or stability issues that impacted our operations due to PagerDuty Operations Cloud.

What do I think about the scalability of the solution?

PagerDuty Operations Cloud is highly scalable and works well with small and large environments. The project I worked on was integrated with multiple application servers and cloud resources for monitoring. PagerDuty Operations Cloud handles all the alerts from different resources and routes them to the appropriate teams. As the infrastructure grows, new services get implemented, escalation policies get defined, and schedules and teams are easily available without requiring major changes in our existing setup. This makes it suitable for an organization to manage large cloud infrastructure and multiple team supports.

Which solution did I use previously and why did I switch?

When I joined this project, they had already implemented PagerDuty Operations Cloud. When I joined, the SOPs and testing were already in process. After a few days, when I was actually onboarded, many of the alerts were configured in PagerDuty Operations Cloud. I did not get the chance to work on different tools besides PagerDuty Operations Cloud.

How was the initial setup?

During the initial setup of PagerDuty Operations Cloud, when I joined the project, I got a Jira ticket listing a few of the servers where I needed to install PagerDuty agents so it could trigger any alerts or integrate with the server. I was mostly involved in the configuration part.

The setup was straightforward. PagerDuty Operations Cloud also helped us in this process. It was not directly integrated on the individual servers, but we integrated our monitoring tools and observability with PagerDuty Operations Cloud. The servers and applications were monitored through application monitoring tools such as Instana, Zabbix, and Splunk. Whenever critical alerts were generated, they would automatically forward to PagerDuty Operations Cloud through the configured integrations we set up with the application. PagerDuty Operations Cloud would notify the on-call engineers and follow different escalation policies if the alerts were not acknowledged within a specific time. Our flow was that we had EC2 instances, AWS servers, and CloudWatch alarms, and if any alert triggered, it would send through SNS, AWS Simple Notification Service, and then to PagerDuty Operations Cloud and the on-call engineer.

What about the implementation team?

We followed the documentation provided by PagerDuty Operations Cloud for the configuration part.

The documentation is full-fledged with proper details on how to configure it depending on the integration with any application monitoring tool. They specify what steps need to be followed. If integrating with servers, they mention which type of server, whether it is Windows or Linux, and accordingly, they have provided all the documents. The documentation is comprehensive and easy to understand, such that even a layperson can do the configuration part with the way they have provided the documentation.

What other advice do I have?

We are not mostly focused on utilizing PagerDuty's autonomous AI agents because we are working on cloud infrastructure where we do the deployments. We have not implemented AI in our cloud to that extent. Going forward, if our infrastructure is AI-based, then we will definitely explore where PagerDuty Operations Cloud can help in that.

As of now, we do not use generative AI capabilities of PagerDuty Operations Cloud. Our infrastructure is huge, and there is a dedicated developer team working on AI-related things. They are still in two POCs, and the POC is being evaluated. If it looks good, then only we can roll this out into production because my application is customer-facing, and we do not want anything to go wrong or if the alert triggers unnecessarily due to some AI alert that did not notify us. That would ultimately cause us to lose our SLAs and SLOs, and all the other escalation matrices would come into the picture. That is why we are still in POCs as it is critical.

That part is taken care of by a different team or mostly the clients themselves. My main role is to keep the environment always up and running, and all alerts should be properly centralized and customized accordingly.

PagerDuty Operations Cloud is basically where we get the alert, and we can integrate through Slack and on-call rotational shifts on cell phones. Prior to this, we were mostly relying on application monitoring tools only and emails and Slack notifications. If an on-call shift person is not at their desk and if any alert has been triggered and no one is there to acknowledge it or look into it and take necessary action, then ultimately there will be customer impact. That is why we implemented PagerDuty Operations Cloud. Even if the on-call person is not near their laptop, they will get the call and can immediately acknowledge and report to the team that we have received a P1 call for this specific environment or that the alert is regarding a production issue. Another team member will immediately take action, so there will not be any miss.

I did not encounter any issues that required contacting support for PagerDuty Operations Cloud. This review represents an overall rating of 9 out of 10.

Which deployment model are you using for this solution?

Public Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Amazon Web Services (AWS)
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
Last updated: Jun 9, 2026
Flag as inappropriate
PeerSpot user
Lead Data Ops Engineer at Wipro Limited
Real User
Top 20
May 31, 2026
Incident workflows have transformed and now reduce downtime for critical gaming services
Pros and Cons
  • "We have seen a positive return on investment from PagerDuty Operations Cloud through improved operational efficiencies, faster incident response, and reduced downtime."

    What is our primary use case?

    My name is Dinesh Singh Negi and I currently work as a Lead DataOps Engineer in the online gaming industry. My primary responsibility is ensuring the reliability, availability, and performance of our data platform and complete production system. I work extensively with AWS services, Prometheus, Grafana, and PagerDuty for monitoring, alerting, and incident management. My team supports critical gaming workloads and data pipelines that require high uptime and quick incident response. A significant part of my role involves setting up monitoring strategies, managing on-call operations, handling production incidents, and performing root cause analysis. We drive operational improvements, and we use PagerDuty Operations Cloud as our central incident management platform to ensure alerts are routed to the right team and escalated appropriately. I have been working in operations and reliability for nine to ten years and have hands-on experience managing large-scale customer-facing environments where managing, minimizing downtime, and reducing meantime to resolution are key priorities. We use PagerDuty Operations Cloud to understand the maximum time of acknowledgment and maximum time of resolution to derive meaningful analysis from the incidents that have been triggered to different teams.

    I have been working for nine to ten years in operation, production support, reliability engineering, and mixed roles during this time. I have worked extensively on monitoring, incident management, system reliability, and operational excellence while particularly supporting large-scale online platforms and data operations. For five to six years, my focus has been on ensuring high availability, managing production incidents, optimizing monitoring and alerting strategies, and improving operational processes. Throughout these years, I have gained hands-on experience with AWS Cloud, Prometheus, Grafana, and PagerDuty Operations Cloud, which are the core tools we use for monitoring, alerting, and incident responses.

    What is most valuable?

    The best features are those we have been using for incident management. We have been using PagerDuty Operations Cloud for on-call scheduling, escalation policies, and integration capabilities. Incident management is extremely valuable because it ensures critical alerts are delivered to the right people immediately. On-call scheduling and escalation policies are very helpful because we can define clear ownership for the services and automatically escalate incidents if they are not acknowledged within a specific timeframe. Another key strength is the integration ecosystem. We can integrate it with our monitoring stack including Prometheus, Grafana, and AWS services, which helps us automate alerts ingestion and incident creation without manual intervention. The most valuable features are automating alerts, escalations, on-call management, integrations, and incident analytics.

    One example that stands out was a production incident where we experienced a sudden spike in database latency during peak gaming hours. This started impacting player transactions and causing delays in some backend services. Our Prometheus and Grafana monitoring detected this abnormal latency and error rate increase, which went beyond a threshold, and the alert was automatically routed to PagerDuty Operations Cloud. PagerDuty Operations Cloud immediately notified the on-call engineer of our team and triggered the escalation workflow based on the incident severity. Since the issue occurred during peak traffic, quick response was critical, which was maintained. PagerDuty Operations Cloud helped us coordinate multiple teams, including DataOps, application, and other infrastructure teams. The platform helped ensure everyone was engaged quickly and that no critical notifications were missed. While we were under investigation, we identified a resource bottleneck in the database layer caused by an unexpected traffic surge. With the help of the database team, we scaled the required AWS resource and optimized a few long-running queries. This restored normal performance.

    What needs improvement?

    A significant positive impact is improving incident response efficiency and overall service reliability. Before we had a mature incident management process, coordinating responses during critical issues often required manual communication and follow-ups. PagerDuty Operations Cloud automated all of those things, including alert ownership, escalation, ensuring that incidents are routed to the right team members immediately. One of the most measurable benefits is the reduction in meantime to acknowledge and meantime to resolve. Faster detection and response help minimize service disruptions and maintain a stable experience for our users, which is especially important in the online gaming industry where availability and performance directly affect customer satisfaction. The platform has helped us mature our operational practices by analyzing incident trends, alert volumes, and escalation patterns. We have been able to refine our monitoring, reduce alert fatigue, and proactively address recurring issues before they become major bottlenecks in production.

    For how long have I used the solution?

    I have been using PagerDuty Operations Cloud for approximately more than five years.

    What do I think about the stability of the solution?

    PagerDuty Operations Cloud is stable.

    What do I think about the scalability of the solution?

    When you are using a tool for incident response, you need to trust that notifications and escalations work when a critical event occurs. PagerDuty Operations Cloud has been very dependable in that regard. Another aspect we have found valuable is the flexibility to support different teams and services as our environment grows. We have added new applications, data pipelines, and AWS service resources. We are able to extend our PagerDuty Operations Cloud configuration without major challenges or changes to our overall operational model.

    Which solution did I use previously and why did I switch?

    I have not used any solution previously. Since the beginning of 2021, I have been using PagerDuty Operations Cloud.

    How was the initial setup?

    The setup and customization process was relatively straightforward. The integrations were one of the easiest parts. PagerDuty Operations Cloud provides well-documented integrations for monitoring tools and cloud platforms. Connecting it with our Prometheus, Grafana, and AWS monitoring stack did not require significant development efforts. The initial setup involved configuring alert routing, defining service ownership, and mapping severity levels to appropriate escalation policies. Customizing on-call schedules and escalation workflows was also quite flexible. We were able to create different schedules for various teams, define escalation paths based on incident severity, and establish notification rules that match our operational requirements. As our team and environment grew, we refined the configuration further by tuning alert thresholds and reducing noise to avoid alert fatigue. It is important to ensure engineers receive only actionable alerts rather than excessive notifications.

    What about the implementation team?

    PagerDuty Operations Cloud's AI and automation capabilities are primarily used for alert correlation, event intelligence, noise reduction, incident prioritization, and providing operational context to responders. These capabilities help engineers identify and respond to issues more quickly while keeping humans in control of critical decisions. We see value in the direction of autonomous operations. If AI agents continue to improve in areas such as incident triage, root cause analysis, and automated remediation for well-understood scenarios, they could further reduce response times and operational overhead.

    What was our ROI?

    We have seen a positive return on investment from PagerDuty Operations Cloud through improved operational efficiencies, faster incident response, and reduced downtime. I cannot share financial figures, but I can speak to operational outcomes we have observed. Since implementing PagerDuty Operations Cloud and integrating it with AWS, Prometheus, and Grafana monitoring stack, we have seen measurable improvements in incident processes such as MTTA and MTTR, or reduced alert fatigue by using event correlation and alert deduplication. These improvements have helped us a great deal.

    Which other solutions did I evaluate?

    I did not get a chance to evaluate any other applications. When I was in the company, they were using PagerDuty Operations Cloud only, so I started with that.

    What other advice do I have?

    My advice would be to start with a clear incident management strategy rather than focusing only on the tool itself. PagerDuty Operations Cloud delivers the most value when you have well-defined service ownership, escalation policies, severity levels, and monitoring practices in place. The platform is very powerful, but its effectiveness depends on the quality of the alerts and operational processes behind it. I would also recommend investing time in alert tuning early on and integrating PagerDuty Operations Cloud with your monitoring stack, whether it is AWS, Prometheus, Grafana, or any other observability tool. Make sure the alerts being sent are actionable. Reducing noise from the beginning will help prevent alert fatigue and improve adoption among engineering teams. I would rate this product an eight out of ten.

    Which deployment model are you using for this solution?

    Public Cloud

    If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

    Amazon Web Services (AWS)
    Disclosure: My company does not have a business relationship with this vendor other than being a customer.
    Last updated: May 31, 2026
    Flag as inappropriate
    PeerSpot user
    Divyajyoti Ghosh - PeerSpot reviewer
    CSO C Apac at Autodesk, Inc.
    Real User
    Top 10
    Mar 18, 2026
    Reliable on-call workflows have supported incident response and integrations across teams
    Pros and Cons
    • "PagerDuty Operations Cloud has been reliable as it has never gone down in my experience; I have never seen it fail."
    • "Our developers noted that integrating incident bots or chatbots created with AI tools occasionally presents challenges with PagerDuty Operations Cloud."

    What is our primary use case?

    We took a subscription and started integrating many applications with it. We have integrated it with ServiceNow and also use our wiki repository by Spotify called Beacon. We onboard multiple services and integrate them with PagerDuty so that different engineering teams can update their line of escalation policy and primary, secondary, and tertiary users who are on call 24/7 or following a Follow the Sun model.

    Additionally, we have integrated PagerDuty for incident commanders who can acknowledge any incident page they receive and update their responses. If they need to hand it over to someone else, they can do that as well. This is how we are actually using PagerDuty. We primarily leverage it for service onboarding. For instance, if you have created a product and need to support it, which might run on any cloud such as EC2, AWS, or Azure, the service teams or triage engineering teams have their members added into PagerDuty along with different playbooks, runbooks, and SOPs integrated with any ticketing tool such as ServiceNow or Jira, whichever you are using. On these fronts, we effectively use PagerDuty.

    PagerDuty Operations Cloud has been reliable as it has never gone down in my experience; I have never seen it fail. I rely on PagerDuty Operations Cloud for on-call support for any high severity incidents or sev zero scenarios. This is a great feature, and I can update it using the mobile app because we also use PagerDuty Operations Cloud mobile app. Occasionally, I may be on call during weekends, and something might come up. For example, on October 20th, I was on leave but used PagerDuty Operations Cloud to stay in sync, even without my laptop. I could join Teams and Zoom calls and simultaneously update required documentation via Copilot and ChatGPT on PagerDuty Operations Cloud regarding different incidents. This capability was extremely helpful.

    What is most valuable?

    The integration with ServiceNow is one of the most valuable features. I rely on PagerDuty Operations Cloud for on-call support for high severity incidents or sev zero scenarios, which is a great feature. PagerDuty Operations Cloud has been reliable as it has never gone down, and I trust it for incident response. Additionally, I find the capability to update using the mobile app extremely helpful.

    Our developers noted that integrating incident bots or chatbots created with AI tools occasionally presents challenges with PagerDuty Operations Cloud. This is the only negative feedback I have encountered. However, I am very satisfied with PagerDuty Operations Cloud because my team is also pleased with it. We have onboarded multiple teams using this tool, and it functions well. From a security perspective, I believe there could be more layers. When scheduling on-call rotations for different team members, access should be restricted to specific users to prevent unauthorized changes to the on-call module. Despite this, the security features have been functioning well, and overall, I appreciate PagerDuty Operations Cloud.

    What needs improvement?

    In terms of integration, while I cannot speak for all developers, some have encountered anomalies, but I expect they will resolve over time. Our developers noted that integrating incident bots or chatbots created with AI tools occasionally presents challenges with PagerDuty Operations Cloud. This is the only negative feedback I have encountered.

    From a security perspective, I believe there could be more layers. When scheduling on-call rotations for different team members, access should be restricted to specific users to prevent unauthorized changes to the on-call module. Despite this, the security features have been functioning well.

    For how long have I used the solution?

    I have been using PagerDuty Operations Cloud since 2021.

    What do I think about the stability of the solution?

    I cannot recall ever having to contact support. I have been using PagerDuty Operations Cloud for more than seven years across two organizations without ever needing assistance. It never breaks down for us, and considering I have devoted 20 years of my career to IT infrastructure operations, where everything typically breaks down, including Jira and ServiceNow, it is impressive to say that PagerDuty Operations Cloud has not caused disruptions. I can also mention MongoDB Atlas as a vendor we subscribe to and similarly, I have never experienced disruptions in their services, aside from scheduled maintenance.

    PagerDuty Operations Cloud performs maintenance, which we are notified of in advance, usually two weeks prior, and this occurs during low-activity periods such as holidays. We adapt our workflows accordingly but typically, these notifications for maintenance are infrequent, occurring once or twice a year, making them manageable.

    What do I think about the scalability of the solution?

    PagerDuty Operations Cloud is scalable, but I emphasize I am a user. In our field, whether supporting applications or web technologies, it is very scalable. However, if developers assess it from an AI perspective, I cannot comment due to our newness to AI. Nevertheless, I find it scalable in all other aspects and have worked with some of the top tools available, apart from Remedy, Siebel, or Lotus Notes. Overall, it is highly scalable.

    How are customer service and support?

    I cannot recall ever having to contact support.

    How would you rate customer service and support?

    Negative

    Which solution did I use previously and why did I switch?

    I have never used any alternatives to PagerDuty Operations Cloud. I have been a bit old school, where we used to get paged on our phone numbers. Aside from PagerDuty Operations Cloud, I cannot recall using anything else.

    Which other solutions did I evaluate?

    Some alternatives to PagerDuty Operations Cloud could be Automation Anywhere, Tray.io, IBM RPA, and a few solutions from Hyland and SAP that also do automation.

    What other advice do I have?

    While I cannot provide specific pricing details, I can share my perspective as an operations professional. Though we use Jira and initially relied on ServiceNow, we have transitioned more towards Vulcan. We never considered moving away from PagerDuty Operations Cloud. I believe that whatever the cost is, it is beneficial because the IT infrastructure operations industry cannot function without PagerDuty Operations Cloud or a similar product. Furthermore, PagerDuty Operations Cloud has an excellent reputation.

    New users are onboarded to ServiceNow or Jira, and they immediately create a PagerDuty Operations Cloud account profile that goes through a verification and approval process by a hierarchy. Once approved, they can set up their numbers and build their profiles to reflect their department, area of expertise, and time zone, allowing them to track incidents outside their shifts while remaining informed about ongoing schedules.

    OpenScape is one product that I used before. I worked with Siemens in healthcare IT infrastructure operations, and during that period, we used BMC Remedy integrated with OpenScape, which was back around 2013 to 2016. Back then, our phone numbers were connected to it, but it was not particularly helpful. If you were not in front of your laptop, you received a call, and the automated IVR provided a brief description of the incident logged. You could only acknowledge or resolve the incident without having the option to assign it. OpenScape was what I had used before opting for PagerDuty Operations Cloud at Salesforce starting in 2021, and continuing with PagerDuty Operations Cloud became more widespread at Autodesk in 2022. As I mentioned earlier, some developers have indicated that integrating bots presents a challenge that has been somewhat resolved over time, but that is the only negativity I have heard about PagerDuty Operations Cloud. I rate this product overall a nine out of ten.

    Disclosure: My company does not have a business relationship with this vendor other than being a customer.
    Last updated: Mar 18, 2026
    Flag as inappropriate
    PeerSpot user
    Buyer's Guide
    Download our free PagerDuty Operations Cloud Report and get advice and tips from experienced pros sharing their opinions.
    Updated: July 2026
    Buyer's Guide
    Download our free PagerDuty Operations Cloud Report and get advice and tips from experienced pros sharing their opinions.