run-llama / liteparse
A fast, open-source document parser with local OCR and multi-format output for AI developers.
活跃维护 Apache 2.0 Rust Tracked
12.2k 849 2 天前
CIDocker
document-ocr document-processing ocr ocr-recognition pdf pdf-parser text-extraction
星标趋势
数据积累中,暂无足够数据生成趋势图
AI 分析
项目摘要
LiteParse is an open-source, fast document parser focused on PDF processing with local execution. It provides spatial text parsing with OCR capabilities, supporting multiple output formats and languages for AI-driven document workflows.
为什么值得关注
Notable for its speed, local processing without cloud dependencies, and flexible OCR system, making it a valuable tool for developers building AI applications that require document parsing.
优势
- Fast local processing with PDFium
- Flexible OCR system with built-in Tesseract and API support
- Multi-platform and multi-language support for Rust, Node.js, Python, and browser
局限性
- Lacks automated tests, which may affect reliability and production readiness
- Complex documents like dense tables or scanned PDFs may require the cloud-based LlamaParse for optimal results
使用场景
- Text extraction from PDFs for RAG pipelines
- OCR processing for scanned documents and images
- Document complexity assessment for routing in AI systems
目标用户: AI developers and data engineers working on document processing pipelines
学习曲线: 中
分析模型:mimo-v2.5-pro | 分析时间:2 个月前