fix(knowledge): honor cancellation during PDF extraction - #532
Open
sylvesterkaczmarek wants to merge 2 commits into
Open
fix(knowledge): honor cancellation during PDF extraction#532sylvesterkaczmarek wants to merge 2 commits into
sylvesterkaczmarek wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Honor the existing knowledge-base cancellation signal during PDF text extraction.
Fixes #531.
Reproduction / evidence
prepareKnowledgeBase()already accepts anAbortSignaland checks it during discovery and before staging each document. On currentmain, PDF handling callsextractPdf(document, bytes)without passing that signal.Inside
extractPdf(), pdf.js loads the document and then iterates through every page. There is no cancellation check in that loop. A cancellation that arrives after PDF loading begins can therefore leave the caller waiting while remaining pages are processed.Reproduction:
prepareKnowledgeBasewith an abort signal.mainhas no extraction-level signal check, so page processing continues.Expected: stop before processing further pages and propagate the caller's cancellation reason.
The regression added in this PR triggers cancellation only when the first extraction-specific signal check is reached, so it exercises the missing boundary rather than the existing discovery/staging checks.
Root cause
The signal was not propagated from
prepareKnowledgeBase()intoextractPdf().Fix
loadingTask.destroy()finallypath.Tests / validation
Added focused regression coverage for cancellation after PDF extraction begins.
The branch was created from upstream
mainat99c85613b0c4b8202b33dfbd80f41884fb9eac11. Production changes are 13 additions and 3 deletions plus the regression file.Full repository tests cannot be run in this execution environment because the repository cannot be cloned here. Pushed-head CI remains the authoritative full-suite validation.
Risk
Low. Normal successful PDF extraction is unchanged except for signal checks. Existing non-cancellation PDF failures keep their current diagnostic behavior.