feat: replace pymupdf image extraction with pdf2image + Heron layout detection
PDF image extraction now renders pages with pdf2image, detects figure regions using docling-layout-heron (RT-DETRv2), and crops only the detected pictures. Removes full-page fallback for text-only pages. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -72,4 +72,5 @@ htmlcov/
|
||||
|
||||
officefile
|
||||
data/1法律
|
||||
data/data_backup
|
||||
|
||||
|
||||
Reference in New Issue
Block a user