Model Serving

LLM 推理引擎和模型服务基础设施

共 22 个项目
排序:综合评分
LMCache / LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11.6k
活跃维护 Apache 2.0
ai-dynamo / dynamo

A Datacenter Scale Distributed Inference Serving Framework

Rust 7.9k
活跃维护 NOASSERTION
ggml-org / llama.cpp

LLM inference in C/C++

C++ 126.3k
正常维护 MIT
microsoft / onnxruntime

ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator

C++ 21.7k
正常维护 MIT
modelscope / FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Python 20.1k
活跃维护 MIT
vllm-project / vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 90.4k
维护停滞 Apache 2.0
vllm-project / vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

C++ 2.7k
维护停滞 Apache 2.0
iree-org / iree

A retargetable MLIR-based machine learning compiler and runtime toolkit.

C++ 3.9k
维护停滞 Apache 2.0
kvcache-ai / Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6.4k
活跃维护 Apache 2.0
Tencent / ncnn

ncnn is a high-performance neural network inference framework optimized for the mobile platform

C++ 23.8k
维护停滞 NOASSERTION
WasmEdge / WasmEdge

WasmEdge is a lightweight, high-performance, and extensible WebAssembly runtime for cloud native, edge, and decentralized applications. It powers serverless apps, embedded functions, microservices, smart contracts, and IoT devices.

C++ 10.8k
活跃维护 Apache 2.0
raullenchai / Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

Python 3.6k
活跃维护 NOASSERTION
EricLBuehler / mistral.rs

Fast, flexible LLM inference

Rust 7.6k
活跃维护 MIT
k2-fsa / sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++ 14.5k
正常维护 Apache 2.0
qualcomm / GenieX

Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code

Rust 8.3k
活跃维护 BSD-3
maximhq / bifrost

Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.

Go 7.6k
活跃维护 Apache 2.0
hao-ai-lab / FastVideo

A unified inference and post-training framework for accelerated video generation.

Python 4.2k
活跃维护 Apache 2.0
beclab / Olares

Open-Source Personal Cloud OS for Always-On Agents

Go 5.2k
活跃维护 AGPL-3
kvcache-ai / ktransformers

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19.3k
活跃维护 Apache 2.0
mozilla-ai / llamafile

Distribute and run LLMs with a single file.

C++ 25.7k
活跃维护 NOASSERTION
altic-dev / FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp

Swift 11.1k
活跃维护 GPL-3
handy-computer / transcribe.cpp

ggml speech-to-text inference for 16+ model families

C++ 1.9k
活跃维护 MIT