mylxsw/extractor

extractor is an HTTP service used to convert PDF, Markdown, HTML, Docx, Xlsx, CSV and other files into plain text output. It is used in RAG implementation to read external documents for vectorization.

/ 100

Experimental

No commits in the last 6 months.

Stale 6m No Package No Dependents

Maintenance 0 / 25

Adoption 4 / 25

Maturity 9 / 25

Community 10 / 25

How are scores calculated?

Stars

Forks

Language

Python

License

MIT

Category

file-content-extraction

Last pushed

Mar 01, 2024

Commits (30d)

GitHub

File Content Extraction · 61 tools

Get this data via API

curl "https://pt-edge.onrender.com/api/v1/quality/rag/mylxsw/extractor"

Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.

Higher-rated alternatives

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR...

kreuzberg-dev/kreuzberg

A polyglot document intelligence framework with a Rust core. Extract text, metadata, and...

yfedoseev/pdf_oxide

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown...

opendataloader-project/opendataloader-pdf

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

NanoNets/docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking...

Explore RAG Tools

All categories Trending RAG directory Insights