What is our primary use case?
My main use case for Splunk Observability Cloud at my current company is monitoring apps. We have a lot of microservices because we sell different SaaS solutions to businesses, so monitoring, tracing, and everything related to that is essential given our microservices architecture.
We have traces and a service map set up where currently we only have two of the apps onboarded. The goal is to get the other four or five apps onboarded so we can have a single pane of glass view of how all of our apps work, how they're all connected, and how the traffic flows through them.
What is most valuable?
The best features Splunk Observability Cloud offers include in-depth capabilities. From going over it at the boot camp these last two days, there is definitely a learning curve, which could be a positive or negative. It's a positive thing because there are so many features, and it definitely takes a little time to get ramped up, but it is feature-rich with any feature one could think of, whether it's monitoring backends, frontends, or RUM metrics.
Regarding the RUM metrics feature, I'm not entirely sure yet, but we definitely have users of our SaaS solutions, so I can see it being useful to see how our apps load for users, how they are liking them, and what their sessions are like so we can research what it's from our end users' point of view.
I'm still relatively new, but I consider pricing one of the positives of how Splunk Observability Cloud has impacted my organization. From what I hear, they were using DataDog, and the pricing was very obscure, so they made a pivot to Splunk for pricing reasons. That combined with giving us clearer insights into the apps has been beneficial.
Clear insights are important because we host in AWS using ECS Fargate. It's difficult to rely on CloudWatch alone to trace everything that's going on with the apps running in the containers.
What needs improvement?
There is a pretty steep learning curve with Splunk Observability Cloud. There is a lot going on, which has pros and cons. The pros include the abundance of offerings in Splunk Observability Cloud, but the cons are that there is so much to learn that you need full-day boot camps just to really understand it.
I feel there is a lot going on, and I think some things could be simplified for certain, particularly the menus and how to accomplish things. Additionally, my understanding is that it is expensive to get infrastructure data into it, though I'm not sure if that's AWS's fault or Splunk's. My opinions may change since I just started using the product.
For how long have I used the solution?
I've just started using Splunk Observability Cloud as I'm new at this company. I previously worked at Choice Hotels for practically my entire tech career. I started at this new company about a month and a half ago, and they use Splunk. I'm at the annual Splunk conference doing a three-day boot camp to learn Splunk Observability Cloud because we use it at this new company. I've been using it for approximately a month at this point.
What do I think about the stability of the solution?
Splunk Observability Cloud is stable.
What do I think about the scalability of the solution?
Splunk Observability Cloud's scalability appears to be infinite from my understanding since it's a SaaS solution, so Splunk will ingest as many logs as we can send. I have heard that there are some limits on things, but I imagine if we pay for it, we can scale as large as we want.
How are customer service and support?
I have not used customer support, but the people I've interacted with here today at the conference have been great, so I can imagine it's more of the same.
Which solution did I use previously and why did I switch?
At this company they used DataDog, but at my previous company, we used OpenSearch.
How was the initial setup?
I wasn't here when the setup, pricing, and licensing happened. However, I have heard that it was cheaper than DataDog, which is why we switched.
What was our ROI?
I have seen a return on investment in terms of time saved. I don't have hard numbers, but using the MCP server has saved me a ton of time. I'm able to multitask and do some work on the side while querying Claude to pull data and traces, which comes back to me rather than me having to manually do it and have that take my full attention.
What other advice do I have?
Splunk Observability Cloud has helped improve my operational performance. I'm quite new, but I would definitely say it has improved operationally. I already got Claude hooked up to Splunk Observability Cloud and the regular Splunk Cloud MCP servers, which has helped me tremendously to be able to track down issues and traces, which could have taken me much longer to do manually.
It's really important for my organization that Splunk Observability Cloud has end-to-end visibility into our cloud-native environment. Otherwise, we're flying blind, especially because some of these new products we're about to release in an early release program require us to know what we're doing and to hit it out of the park to make it a success.
To help our organization scale, we only have two apps piped into it so far. We have acquired a couple of other companies, so Splunk Observability Cloud has allowed us to consider piping in that data as well to get the single pane of glass situation as we continue to build and grow the company.
I haven't noticed any specific outcomes or improvements in lowering the cost of unplanned digital downtime using Splunk Observability Cloud yet. I haven't done a traditional on-call role yet because I'm new to this company. We are beginning our on-call rotation, and I will definitely be leaning on Splunk Observability Cloud more heavily as I am on call and handling outages and after-hours calls. The solution's out-of-the-box dashboards and detectors are great. I'm doing the boot camp at the conference, and creating dashboards and filtering was probably my favorite part. Learning the roll-ups and the functions and how all of that works was excellent. The detectors seem very easy to create, and I appreciate how it creates a border on a chart when that chart is in the alert state.
I haven't really looked into the AI capabilities too much. I have used the AI tool that can help me write SPL, but that's about it. I typically use the MCP server for it and plug it into Claude, which has been my main way to interact with it from an AI perspective.
Regarding the accuracy and reliability of output from Splunk Observability Cloud's AI capabilities, I haven't used it too much beyond creating SPL, and so far there are no issues to report.
Disclosure: My company does not have a business relationship with this vendor other than being a customer.