What is our primary use case?
I am currently working with Splunk Cloud, but I could review Splunk Enterprise as well because it is fundamentally the same product but with hosting being on-premises instead. Splunk Enterprise, Splunk Cloud, or Splunk ITSI (IT Service Intelligence) are all options we use.
We are a customer of Splunk On-Call service. We use Splunk On-Call as an acknowledgment tool. It is not our primary ticketing solution for tickets; however, it creates tickets in Splunk On-Call. When we have priority one tickets assigned to us in ServiceNow, that is when we get the notification in Splunk On-Call to acknowledge that we have a priority one ticket in ServiceNow. We use it for the on-call part and we use the roster to schedule who is going to be on call. We also schedule and allocate who is the primary responder during office hours and during nighttime.
We do use the on-call scheduling feature of Splunk On-Call. We are integrating Splunk On-Call directly with different back-end monitoring systems. First of all, you can initiate some of the alerts directly from monitoring tools, but our main primary integration is through Splunk IT Service Intelligence. That is where you have the backbone of the events being handled within IKEA. When you set the attribute that this event goes into the pipeline, you can say that this should be creating a ticket into ServiceNow or a Splunk On-Call acknowledgment as well. That integration has been developed, and it is very tight. Whenever you look at Splunk On-Call, which is the entry point on the acknowledgment, you will be able to see that it is coming from this monitoring system, and you have this associated ServiceNow ticket or associated Jira ticket. From there you can directly dive into runbooks to validate if this issue is still relevant or if it has been self-healing. We also have the bi-directional part that we are working on. Whenever a system is self-healing itself, Splunk On-Call acknowledgment will be closed by the system, and they will no longer be paging people based on something that has been solved.
We have the acknowledgment part that we have discussed with Splunk On-Call. We have the escalation policies. We have the possibility to broadcast during daytime to as many responders as possible, so we get wide coverage with that. But also during nighttime, if it is just a single responder, we are sure that someone picks it up. We also have some built-in capabilities to see the reports and the assessment afterwards, in the form of post-mortem reports. We can see what happened during this time, who was involved, exactly who was paged and why they did not pick it up. We can see measurements of the mean time to acknowledge. Since it is not our primary responding primary ticketing system, we are not measuring mean time to recover, but we are working with that with the bi-directional integration. The report part is excellent for doing the follow-up on that aspect. It helps the organization both to have transparency and visibility of the current situation, but also transparency and visibility afterwards, regarding what went wrong or what went well.
What is most valuable?
The most valuable feature of Splunk On-Call is that you have clear traceability on who is acknowledging the ticket and who has been paged, and who they are paging for the notification. Additionally, the advanced way that you can integrate with our existing event pipeline is valuable, where we can get the annotations in Splunk On-Call to see what the runbook is for this type of issue, what the associated tickets in ServiceNow are, and so on.
It has helped a lot because we now have reassurance that whenever there is a notification in Splunk On-Call, we are making sure that it has been acknowledged through different escalation policies and through different routing. If the primary responder is not taking care of the acknowledgment of the ticket, we have a different escalation so it goes to a second set of managers in this example. We are making sure that there is a fallback in the acknowledgment.
We are able to know for certain that the tickets are being acknowledged and being worked on. It does not really matter how you set up your own paging policy. As the initiator of the acknowledgment, I do not really care about the details. I basically say I have a priority one, and this is the team that should take care of that. How they do it behind the scenes, I do not care about, but there is full transparency on the follow-up and they can have very advanced or very simple acknowledgment principles, but I know it has been taken care of. It is not fire and forget. It is fire and make sure it lands.
For critical incidents, the escalation policies of Splunk On-Call is critical. If the primary responders are not picking up, I know that the escalation policy is falling back to a higher tier of escalation. In most of the cases that we have implemented, it goes to managers. They are making sure that within IKEA, we have said that for a priority one ticket, we need to have acknowledgment within thirty minutes. The escalation policy is making sure that we are keeping that promise to our consumers because the escalation policies make sure that if the primary responders are not picking up, the secondary responders are taking it.
What needs improvement?
I would say a little bit of the user experience and the user interface of Splunk On-Call can be improved. I have already started to gather metrics from the tool and started to analyze that. I would like to see more about that as a part of the solution. Overall, it is a very good tool. They have a lot of integrations to different back-ends, which I appreciate a lot. However, there is some slight development that could be improved in the tool.
There are a couple of small things where you do not have insights to the tool. For example, if you invite a consumer based on the email address and there are some issues during the invitation process when they are failing to sign up, there are some insights to the tool that you as a consumer do not have knowledge about. You need to reach out to Splunk support and ask why this user cannot sign up, for example, and then they need to resolve that. There is some lack of insights to the internal working, which makes me not rate it as a ten. However, those use cases are not common.
For how long have I used the solution?
We purchased Splunk On-Call one year ago and we started to implement it six months ago.
What do I think about the stability of the solution?
It has been very stable. It has sometimes been a little slow in the response. The first page takes some time to log on and to get started. However, I have not experienced any type of downtime on Splunk On-Call. There is no downtime at all measured on it.
What do I think about the scalability of the solution?
From what I have seen, you can scale Splunk On-Call, and it is not a problem. You just need to make sure that you are using an implementation structure that is scalable. For example, if you have assignment groups in ServiceNow, you should have the equal structure in Splunk On-Call to make sure that you have the cross-reference between the tools. Otherwise, if one team is called A in one tool and it is called B in another tool, you will get into an administrative nightmare. However, I think that is the same challenge for any type of solutions that is scalable, and you can make an administrative nightmare for yourself if you are not cautious about it.
How are customer service and support?
I use Splunk Support for technical support as that is the entry point for all our Splunk solutions to Splunk Support. Splunk Support itself has opinions about that. The first level of support is not always that great, but if you pass that and get into the second line support, they are really excellent. I would rate the technical support from Splunk On-Call at an eight from my perspective.
Which solution did I use previously and why did I switch?
We did use a home-built solution before Splunk On-Call that was much more expensive and not as flexible as this.
How was the initial setup?
I did participate in the initial setup of Splunk On-Call and the deployment.
What about the implementation team?
It was pretty straightforward. The sales engineers gave us a quick start. However, the biggest improvement that we had was when I took the extended education and really learned how to master the tool and to use all the components in it. I think that is the biggest suggestion I have for anyone and any tool. Whatever tool you select, you should master it. You should really learn it. There is very much that, if you get a tool that is just fire and forget, and one that is a five-click installation type of tool, they are not that advanced and flexible in the long run. Easy to start is not necessarily guaranteeing that you have a solution that is worth the money and affordable.
What was our ROI?
The big benefit of using Splunk On-Call is that I know I have a high-performance team, but now I can really measure and prove it. I can prove that we are acknowledging within seconds and minutes, rather than hours, that when we have an issue, we are on it.
What's my experience with pricing, setup cost, and licensing?
I would say it is very affordable. I have compared Splunk On-Call internally with our own homemade solution. I have also compared it with other vendors such as PagerDuty, and it is really affordable.
Which other solutions did I evaluate?
We did evaluate other options or other vendors before choosing Splunk On-Call. However, there were very few vendors that have this isolated acknowledgment and notification service. A lot of the times, they are built in as a part of the overall solution. For example, Jira and those types of services. You also have notification services from ServiceNow built in, but then we cannot use it incorporated and integrated to other solutions that we have internally.
What other advice do I have?
Splunk On-Call is only hosted on the cloud. I am not certain what cloud platform Splunk On-Call is hosted on. It is hosted by Splunk itself, but what cloud platform Splunk is using, I do not know. I know that they are currently hosting it in the US and they are looking at having it hosted in Europe as well. However, what the underlying cloud platform is, I do not know.
We are using primarily only the Splunk ecosystem in my department. We, as a company, are using Sentinel, but not my department.
I would rate this review as a nine overall.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Other