2026-08 Multimodal 分类最佳 AI 开源项目
Port of OpenAI's Whisper model in C/C++
The open-source AI voice studio. Clone, dictate, create.
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
Cross-platform, customizable ML solutions for live and streaming media.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Deepfakes Software For All
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
基于图像识别的自动化黑盒测试框架 | An automation black-box testing framework based on image recognition
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
Visualize, query, and stream to train on multimodal robotics data.