5 comments

  • antonly1 hour ago
    I just tried this and the results look very off. I put in the MiniSat[1] Paper, and the result appears catastrophically wrong: <a href="https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;i3RHwcQ" rel="nofollow">https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;i3RHwcQ</a><p>[1]: <a href="http:&#x2F;&#x2F;minisat.se&#x2F;downloads&#x2F;MiniSat.pdf" rel="nofollow">http:&#x2F;&#x2F;minisat.se&#x2F;downloads&#x2F;MiniSat.pdf</a>
  • phenomen1 hour ago
    I currently use <a href="https:&#x2F;&#x2F;github.com&#x2F;firecrawl&#x2F;anydoc" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;firecrawl&#x2F;anydoc</a> in my PDF pipelines. In most cases it performs well. I&#x27;ll test your lib to compare.
    • coinfused1 hour ago
      Do you know if this one can also crop the PDF to a region of interest before converting to markdown?
  • JaumeGar1 hour ago
    Really nice work, thanks for sharing. The table and formula extraction look great.
  • archeantus1 hour ago
    Great work, thanks for sharing.
  • beatrizalmeidaf3 hours ago
    I built this because extracting PDFs into Markdown&#x2F;JSON often loses reading order, tables, formulas, figures, and their original locations.<p>The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word.<p>I&#x27;m also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF.<p>The project is open source and I&#x27;d love feedback on the architecture, extraction quality, and useful use cases.
    • thatcherc2 hours ago
      This looks fantastic! The table and formula extraction features are especially interesting. My immediate question is: can this be integrated into Zotero? Most of the PDFs I read are research papers and extracting tables and formulas directly from my zotero collection would be super handy.