I just tried this and the results look very off. I put in the MiniSat[1] Paper, and the result appears catastrophically wrong: <a href="https://imgur.com/a/i3RHwcQ" rel="nofollow">https://imgur.com/a/i3RHwcQ</a><p>[1]: <a href="http://minisat.se/downloads/MiniSat.pdf" rel="nofollow">http://minisat.se/downloads/MiniSat.pdf</a>
I currently use <a href="https://github.com/firecrawl/anydoc" rel="nofollow">https://github.com/firecrawl/anydoc</a> in my PDF pipelines.
In most cases it performs well. I'll test your lib to compare.
Really nice work, thanks for sharing. The table and formula extraction look great.
Great work, thanks for sharing.
I built this because extracting PDFs into Markdown/JSON often loses
reading order, tables, formulas, figures, and their original locations.<p>The goal is a lightweight document extraction pipeline that preserves
document structure and bounding boxes while exporting to Markdown, JSON,
Excel and Word.<p>I'm also working on structure-aware semantic chunking for RAG, so retrieved
chunks can retain their section, page and exact visual location in the PDF.<p>The project is open source and I'd love feedback on the architecture,
extraction quality, and useful use cases.