2026-08 Model Serving 分类最佳 AI 开源项目
Community maintained hardware plugin for vLLM on Ascend
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
LLM inference in C/C++
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
A Datacenter Scale Distributed Inference Serving Framework
FlashInfer: Kernel Library for LLM Serving