Skip to content

Recover truncated inline-image PDF text per page with optional PyMuPDF - #2556

Open
Çağdaş Yürekli (cagdasyurekli) wants to merge 2 commits into
microsoft:mainfrom
cagdasyurekli:fix/review-pdf-20260926
Open

Çağdaş Yürekli (cagdasyurekli) wants to merge 2 commits into
microsoft:mainfrom
cagdasyurekli:fix/review-pdf-20260926

Conversation

@cagdasyurekli

Copy link
Copy Markdown
Contributor

PDF text can stop at an inline image while conversion still returns a plausible prefix. This adds optional, local, per-page recovery through markitdown[pdf,pdf-recovery].

The recovery backend inspects decoded images, including compressed page streams, and replaces a plain page only when its extracted words strictly extend the primary word sequence. Empty primary pages are eligible. Pages recognized as forms/tables retain their existing Markdown, and there is no whole-document length gate or replacement. Backend absence/failure keeps the primary extraction. PyMuPDF is not added to the all or pdf extras.

Relates to #1870; alternative to #1883, #1889 and #2092, incorporating the per-page preservation and compressed-stream review concerns. This is not a claim that every affected PDF is recovered: reordered text and form/table pages are deliberately left alone, and there is no OCR.

Validation:

  • Three real-PDF regressions fail on main and pass here: uncompressed inline content, compressed inline content, and a mixed document with over 2,048 characters of healthy prose plus a Markdown table.
  • The synthetic fixture uses an ASCII85 inline image with a malformed bare ~ terminator that PyMuPDF tolerates while pdfminer consumes following text. Fixture provenance is included. This demonstrates the parser-divergence case, not that all valid inline images cause truncation.
  • Nine focused tests pass, including empty primary text, rejecting unrelated longer output, optional-dependency absence, and backend failure.
  • Full core suite: 975 passed, 43 skipped on Python 3.12 with PyMuPDF 1.28.2. Remote/API tests skipped. Black 23.7.0 and git diff --check pass.

Trade-off: installing the optional backend adds a second local PDF inspection pass. Without it, the normal extraction path is retained.

@cagdasyurekli

Copy link
Copy Markdown
Contributor Author

Updated the branch with current main (4cc9fa1) using a locally verified merge.

Validation on Python 3.12 with PyMuPDF 1.28.2: Black 23.7.0 and git diff --check pass; full core pytest: 1059 passed, 57 skipped under the repository CI skip policy. The optional PDF recovery regression tests ran successfully.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant