What is our primary use case?
The main use for
Apache SkyWalking is application monitoring and troubleshooting during integration.
What is most valuable?
The best features of
Apache SkyWalking are distributed tracing, service topology visualization, and trace-correlated logs and metrics. This makes it much easier to follow requests across multiple services, identify where an integration failure occurs, and troubleshoot performance issues. I also find the dashboards, alerting, and profiling capabilities useful for investigating issues during testing.
These specific features of Apache SkyWalking have made troubleshooting more collaborative and efficient. Instead of relying only on application logs, our team, including the QA and development teams, can follow the same trace, quickly identify the service causing an issue, and share concrete evidence. This helps us respond to integration failures faster and makes communication between QA, developers, and operations more focused.
The dashboards of Apache SkyWalking are particularly useful for getting a quick overview of service health, performance, and dependencies. The alerting capabilities also help the team detect unusual behavior and performance degradation earlier, rather than waiting for users or testers to report an issue.
Apache SkyWalking has improved our ability to troubleshoot distributed applications and resolve integration issues more quickly. It gives QA, developers, and operations better visibility into service interactions, which reduces the time spent identifying root causes. It has also improved collaboration because we can share traces and monitoring data as concrete evidence when investigating defects.
Regarding the metrics, we saw improvements mainly in troubleshooting and incident resolution. For example, Apache SkyWalking helped reduce the time needed to identify the affected service during integration failures from roughly 30 to 60 minutes of manual log investigation to around 10 to 20 minutes when a trace was available. We also had better visibility into response times, error rates, and service dependencies, which helped us identify performance regressions earlier.
What needs improvement?
Apache SkyWalking could be improved by making the initial setup and configuration more straightforward, especially for teams that are new to observability. More guided tutorials, ready-to-use dashboard templates, and perhaps some troubleshooting documentation would help reduce the learning curve.
I would also prefer to see even more out-of-the-box integrations with common CI/CD, cloud, logging, and incident management tools, so teams can connect Apache SkyWalking to their existing workflows with less configuration. The UI is already quite powerful, but it could be made more intuitive for new users, particularly when navigating between services, traces, logs, and dashboards.
For how long have I used the solution?
I have been using Apache SkyWalking for around one year.
What do I think about the stability of the solution?
Apache SkyWalking has been very stable in our experience.
What do I think about the scalability of the solution?
Apache SkyWalking is highly scalable, particularly for distributed and cloud-native environments.
How are customer service and support?
Regarding customer support for Apache SkyWalking, I would give it eight out of ten because it is good. It is more community-driven than traditional commercial support. There is strong official documentation,
GitHub discussions, and issue tracking, as well as Apache mailing lists and Slack channels where users can ask questions and troubleshoot problems. The main limitation is that there is not the same dedicated, guaranteed response support model you would typically get from a commercial observability vendor.
How was the initial setup?
The pricing experience was quite positive because Apache SkyWalking is open source, so there are no traditional software licensing fees. The main costs are associated with deployment infrastructure, storage, monitoring, and the time required for initial setup and configuration.
What was our ROI?
We have seen a positive return, mainly through time saved during troubleshooting rather than direct licensing savings. In our experience, Apache SkyWalking reduced the time spent investigating integration issues by roughly 30 to 40%. For instance, an issue that previously could take 30 to 60 minutes of manually checking logs across different services could often be narrowed down to the affected service within 10 to 20 minutes using distributed traces and service topology.
Which other solutions did I evaluate?
We evaluated several options such as
Grafana with Tempo and Elastic
APM. We chose Apache SkyWalking because it provides a more complete observability platform, combining distributed tracing, metrics, logs, service topology, profiling, and alerting in one solution.
What other advice do I have?
Apache SkyWalking's AI capabilities have a good approach to governance and security because the AI assistant is designed to work within existing permissions and is read-only, so it does not automatically make changes to the monitored environment. It also supports authentication, TLS, and access controls over the underlying observability data.
The AI output of Apache SkyWalking is very useful and generally reliable for investigation and troubleshooting, especially because it is grounded in live Apache SkyWalking data such as metrics, traces, logs, and topology, rather than relying only on the model's general knowledge. However, I would not treat it as completely correct. I would still need to validate things before moving to production.
Start with a focused use case rather than trying to monitor everything at once. I would give Apache SkyWalking an overall rating of nine out of ten.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?