66 points | by zxdc7896 1 day ago
4 comments
> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...
> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...