I hope Private AI can achieve better performance because each time I input questions, I have to wait approximately ten seconds or sometimes twenty seconds to get the answer. I hope it can improve the speed and overall performance so that I can get answers in one or two seconds instead. The performance issue prevents Private AI from being a perfect ten because I have to wait for a long time before getting the answer. I believe the suggestion is to improve the performance before deployment so that we do not have to wait for a long time when using Private AI.
The provisioning of hardware is unique because we are a software company. We identify and provide to the client the hardware, such as the GPU for creating a Kubernetes cluster and running large language models. Provisioning the hardware is the most difficult task. For this, we use partners, such as Lenovo. The most important challenge is related to developing and providing generative AI solutions. For example, deploying a large language model of 70 billion parameters is not an easy task on an on-premise architecture because you have to provide many GPUs. For the deployment of big large language models, it is a challenge. For many clients, we try to deploy small language models specific for a particular task. Providing a fast way for creating an on-premise architecture for large language models should be an important challenge for the future.
Principal Software Developer at a consultancy with 11-50 employees
Real User
Top 20
Jul 2, 2026
Regarding how Private AI can be improved, currently serving open weights models is complicated. There are many parameters that you have to adjust to make it run well. The ecosystem is maturing and becoming easier. Tools like Ollama allow running models with no friction, but with suboptimal performance. The speed of development is important as the ecosystem is advancing rapidly, and the tools and overlays need to catch up to make it easier for normal users. I have several computers of different kinds, such as a Mac Studio, a box in my house with NVIDIA GeForce cards like 3090, and a rented server with H100s or H200s. I do not know which quantization I should choose because quantization is complicated. The performance of models is affected by quantization type, MTU, token predictions, KV cache type, and cache size. Recipes for running models better would be appreciated. Some providers like Ansloth and G-Lang provide them, but they cover so much hardware that it is really hard to make a choice.
Private AI can be improved by adding more features to the platform. For example, two people could use it at the same time, and there could be options for sharing libraries or folders. Adding more advanced and helpful features for AI students and for day-to-day life is needed.
Private AI can be improved for security. Encryption is a specific area that could be stronger. I don't think Private AI needs any other improvements outside of security.
Private AI safeguards sensitive data by offering advanced privacy protection tailored to today's digital landscape. It integrates seamlessly with existing systems while ensuring compliance with privacy regulations. Designed for organizations prioritizing data privacy, Private AI provides a dynamic approach to protecting user information. Using cutting-edge techniques, it identifies and redacts personal data in datasets, ensuring data remains confidential throughout processes. With a focus on...
I hope Private AI can achieve better performance because each time I input questions, I have to wait approximately ten seconds or sometimes twenty seconds to get the answer. I hope it can improve the speed and overall performance so that I can get answers in one or two seconds instead. The performance issue prevents Private AI from being a perfect ten because I have to wait for a long time before getting the answer. I believe the suggestion is to improve the performance before deployment so that we do not have to wait for a long time when using Private AI.
The provisioning of hardware is unique because we are a software company. We identify and provide to the client the hardware, such as the GPU for creating a Kubernetes cluster and running large language models. Provisioning the hardware is the most difficult task. For this, we use partners, such as Lenovo. The most important challenge is related to developing and providing generative AI solutions. For example, deploying a large language model of 70 billion parameters is not an easy task on an on-premise architecture because you have to provide many GPUs. For the deployment of big large language models, it is a challenge. For many clients, we try to deploy small language models specific for a particular task. Providing a fast way for creating an on-premise architecture for large language models should be an important challenge for the future.
Regarding how Private AI can be improved, currently serving open weights models is complicated. There are many parameters that you have to adjust to make it run well. The ecosystem is maturing and becoming easier. Tools like Ollama allow running models with no friction, but with suboptimal performance. The speed of development is important as the ecosystem is advancing rapidly, and the tools and overlays need to catch up to make it easier for normal users. I have several computers of different kinds, such as a Mac Studio, a box in my house with NVIDIA GeForce cards like 3090, and a rented server with H100s or H200s. I do not know which quantization I should choose because quantization is complicated. The performance of models is affected by quantization type, MTU, token predictions, KV cache type, and cache size. Recipes for running models better would be appreciated. Some providers like Ansloth and G-Lang provide them, but they cover so much hardware that it is really hard to make a choice.
Private AI can be improved by adding more features to the platform. For example, two people could use it at the same time, and there could be options for sharing libraries or folders. Adding more advanced and helpful features for AI students and for day-to-day life is needed.
Private AI can be improved for security. Encryption is a specific area that could be stronger. I don't think Private AI needs any other improvements outside of security.