OpenDCAI / DataFlow
An all-in-one LLM-powered data preparation framework that streamlines data cleaning, generation, and processing with vLLM and SGLang backends.
活跃维护 Apache 2.0 Python Tracked
7.8k 1.1k 11 天前
CIDockerPyPI
data data-cleaning data-pipelines data-processing data-science data-synthesis llms operators data-agent sglang-bankend vllm-backend gradio-interface quick-data-processing
星标趋势
数据积累中,暂无足够数据生成趋势图
AI 分析
项目摘要
DataFlow is an all-in-one data preparation framework that leverages LLMs for data cleaning, generation, and processing. It provides operators and pipelines with support for vLLM and SGLang backends, featuring a Gradio interface and comprehensive tooling for LLM data workflows.
为什么值得关注
With nearly 7k stars and 52 contributors in under 2 years, DataFlow has rapidly become a go-to solution for LLM-based data preparation. Its integration with high-performance inference backends (vLLM, SGLang) and focus on production-ready pipelines fills a critical gap in the LLM data engineering ecosystem.
优势
- Comprehensive LLM-based operators for data cleaning, synthesis, and processing
- Integration with high-performance inference backends (vLLM, SGLang)
- Production-ready with CI, tests, Docker, and PyPI releases
- Active community with 52 contributors and strong maintenance
- Multiple interfaces including Gradio UI and programmatic API
局限性
- No examples directory provided despite having a Colab notebook
- Relatively new project with limited long-term track record
- Focused specifically on LLM-based processing which may not suit all data prep scenarios
使用场景
- Preparing and cleaning training data for LLM fine-tuning
- Generating synthetic datasets using LLMs
- Building automated data processing pipelines for NLP workflows
- Data quality filtering and validation for ML training
目标用户: ML engineers, data scientists, and researchers who need to prepare, clean, or generate training data for LLMs and other AI models
学习曲线: 中
分析模型:LongCat-2.0 | 分析时间:1 个月前