feat: replace pymupdf image extraction with pdf2image + Heron layout detection

PDF image extraction now renders pages with pdf2image, detects figure
regions using docling-layout-heron (RT-DETRv2), and crops only the
detected pictures. Removes full-page fallback for text-only pages.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
2026-06-02 22:38:31 +08:00
parent bbf6b921af
commit 6395ee3b49
7 changed files with 381 additions and 55 deletions
+1
View File
@@ -72,4 +72,5 @@ htmlcov/
officefile
data/1法律
data/data_backup