Repos

llama.cpp

August 30, 2026 1 min read SvaNews
llama.cpp

LLM inference in C/C++ by Georgi Gerganov — runs on CPU with no GPU required, on minimal hardware. Introduced the GGUF quantized model format used industry-wide.

Category: Model Inference & LLM Tooling  •  Stars: 123K

View on GitHub →

Discover more from Svanews

Subscribe now to keep reading and get access to the full archive.

Continue reading