2026-07 Model Serving 分类最佳 AI 开源项目
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Community maintained hardware plugin for vLLM on Ascend
7.4 billion tokens per month. 34 free LLM providers. 635 free model endpoints. All behind one /v1 endpoint, plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experimentation only.
Distribute and run LLMs with a single file.
WasmEdge is a lightweight, high-performance, and extensible WebAssembly runtime for cloud native, edge, and decentralized applications. It powers serverless apps, embedded functions, microservices, smart contracts, and IoT devices.
ncnn is a high-performance neural network inference framework optimized for the mobile platform
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference