👁️

Multimodal

视觉、音频和多模态模型工具

共 26 个项目
排序:综合评分
unslothai / unsloth

Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.

Python 75.2k
正常维护 Apache 2.0
Comfy-Org / ComfyUI

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.

Python 130.7k
维护停滞 GPL-3
MaaAssistantArknights / MaaAssistantArknights

《明日方舟》小助手,全日常一键长草!| A one-click tool for the daily tasks of Arknights, supporting all clients.

C++ 22.9k
正常维护 AGPL-3
huggingface / transformers

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 164.6k
正常维护 Apache 2.0
openvinotoolkit / openvino

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

C++ 10.8k
活跃维护 Apache 2.0
calesthio / OpenMontage

World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.

Python 54.2k
活跃维护 AGPL-3
sgl-project / sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 32.8k
正常维护 Apache 2.0
lance-format / lance

Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..

Rust 7.0k
正常维护 Apache 2.0
invoke-ai / InvokeAI

Invoke is a leading creative engine for Stable Diffusion models, empowering professionals, artists, and enthusiasts to generate and create visual media using the latest AI-driven technologies. The solution offers an industry leading WebUI, and serves as the foundation for multiple commercial products.

Python 28.0k
活跃维护 Apache 2.0
leejet / stable-diffusion.cpp

Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++

C++ 6.9k
活跃维护 MIT
rerun-io / rerun

Visualize, query, and stream to train on multimodal robotics data.

Rust 11.4k
维护停滞 Apache 2.0
Blaizzy / mlx-audio

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

Python 7.8k
活跃维护 MIT
roboflow / supervision

We write your reusable computer vision tools. 💜

Python 49.8k
活跃维护 MIT
ggml-org / whisper.cpp

Port of OpenAI's Whisper model in C/C++

C++ 53.3k
维护停滞 MIT
MaaXYZ / MaaFramework

基于图像识别的自动化黑盒测试框架 | An automation black-box testing framework based on image recognition

C++ 4.7k
活跃维护 LGPL-3
NVIDIA-NeMo / Speech

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

Python 18.4k
活跃维护 Apache 2.0
sgl-project / sglang-omni

SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

Python 976
活跃维护 Apache 2.0
google-ai-edge / mediapipe

Cross-platform, customizable ML solutions for live and streaming media.

C++ 36.8k
活跃维护 Apache 2.0
opencv / opencv

Open Source Computer Vision Library

C++ 90.6k
维护停滞 Apache 2.0
mayocream / koharu

ML-powered manga translator, written in Rust.

Rust 5.4k
活跃维护 Apache 2.0
microsoft / data-formulator

🪄 Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data.

Python 17.0k
活跃维护 MIT
Anil-matcha / Open-Generative-AI

Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT licensed.

JavaScript 27.4k
活跃维护 MIT
debpalash / OmniVoice-Studio

Local voice clone, video dubbing, dictation and audiobook maker. The open-source ElevenLabs alternative.

Python 9.6k
正常维护 AGPL-3
liustack / modlens

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

TypeScript 3.8k
活跃维护 MIT
index-tts / index-tts

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 23.6k
活跃维护 NOASSERTION
abus-aikorea / voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

Python 12.7k
正常维护 GPL-3