Model Serving
LLM 推理引擎和模型服务基础设施
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A Datacenter Scale Distributed Inference Serving Framework
LLM inference in C/C++
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A high-throughput and memory-efficient inference and serving engine for LLMs
Community maintained hardware plugin for vLLM on Ascend
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
ncnn is a high-performance neural network inference framework optimized for the mobile platform
WasmEdge is a lightweight, high-performance, and extensible WebAssembly runtime for cloud native, edge, and decentralized applications. It powers serverless apps, embedded functions, microservices, smart contracts, and IoT devices.
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
Fast, flexible LLM inference
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
A unified inference and post-training framework for accelerated video generation.
Open-Source Personal Cloud OS for Always-On Agents
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Distribute and run LLMs with a single file.
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp
ggml speech-to-text inference for 16+ model families