High-performance serving framework for LLMs and multimodal models in Python. RadixAttention shares KV cache across common prefixes, boosting throughput on batch workloads.
Category: Model Inference & LLM Tooling • Stars: 32K
IT News, How To's & updates
High-performance serving framework for LLMs and multimodal models in Python. RadixAttention shares KV cache across common prefixes, boosting throughput on batch workloads.
Category: Model Inference & LLM Tooling • Stars: 32K