run-llama / liteparse

A fast, open-source document parser with local OCR and multi-format output for AI developers.

活跃维护 Apache 2.0 Rust Tracked
12.2k 849 2 天前
CIDocker
document-ocr document-processing ocr ocr-recognition pdf pdf-parser text-extraction

星标趋势

数据积累中,暂无足够数据生成趋势图

AI 分析

项目摘要

LiteParse is an open-source, fast document parser focused on PDF processing with local execution. It provides spatial text parsing with OCR capabilities, supporting multiple output formats and languages for AI-driven document workflows.

为什么值得关注

Notable for its speed, local processing without cloud dependencies, and flexible OCR system, making it a valuable tool for developers building AI applications that require document parsing.

优势

  • Fast local processing with PDFium
  • Flexible OCR system with built-in Tesseract and API support
  • Multi-platform and multi-language support for Rust, Node.js, Python, and browser

局限性

  • Lacks automated tests, which may affect reliability and production readiness
  • Complex documents like dense tables or scanned PDFs may require the cloud-based LlamaParse for optimal results

使用场景

  • Text extraction from PDFs for RAG pipelines
  • OCR processing for scanned documents and images
  • Document complexity assessment for routing in AI systems
目标用户: AI developers and data engineers working on document processing pipelines
学习曲线:
分析模型:mimo-v2.5-pro | 分析时间:2 个月前