OpenDCAI / DataFlow

An all-in-one LLM-powered data preparation framework that streamlines data cleaning, generation, and processing with vLLM and SGLang backends.

活跃维护 Apache 2.0 Python Tracked
7.8k 1.1k 11 天前
CIDockerPyPI
data data-cleaning data-pipelines data-processing data-science data-synthesis llms operators data-agent sglang-bankend vllm-backend gradio-interface quick-data-processing

星标趋势

数据积累中,暂无足够数据生成趋势图

AI 分析

项目摘要

DataFlow is an all-in-one data preparation framework that leverages LLMs for data cleaning, generation, and processing. It provides operators and pipelines with support for vLLM and SGLang backends, featuring a Gradio interface and comprehensive tooling for LLM data workflows.

为什么值得关注

With nearly 7k stars and 52 contributors in under 2 years, DataFlow has rapidly become a go-to solution for LLM-based data preparation. Its integration with high-performance inference backends (vLLM, SGLang) and focus on production-ready pipelines fills a critical gap in the LLM data engineering ecosystem.

优势

  • Comprehensive LLM-based operators for data cleaning, synthesis, and processing
  • Integration with high-performance inference backends (vLLM, SGLang)
  • Production-ready with CI, tests, Docker, and PyPI releases
  • Active community with 52 contributors and strong maintenance
  • Multiple interfaces including Gradio UI and programmatic API

局限性

  • No examples directory provided despite having a Colab notebook
  • Relatively new project with limited long-term track record
  • Focused specifically on LLM-based processing which may not suit all data prep scenarios

使用场景

  • Preparing and cleaning training data for LLM fine-tuning
  • Generating synthetic datasets using LLMs
  • Building automated data processing pipelines for NLP workflows
  • Data quality filtering and validation for ML training
目标用户: ML engineers, data scientists, and researchers who need to prepare, clean, or generate training data for LLMs and other AI models
学习曲线:
分析模型:LongCat-2.0 | 分析时间:1 个月前