GA evidence appendix: render a verify-page report as pasteable review wikitext (PRD-0016)#154
Conversation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
…ale allows - Issue 1: truncate_claim() now compares char_count instead of byte length to avoid spurious ellipsis on multibyte claims - Issue 2: BlockFailure.reason now routed through escape_verbatim() for safe wikitext embedding per PRD-0016 hard rule - Issue 3: sanitize_reason() rewritten to use char-safe string operations (find() on char boundaries) instead of byte-offset slicing that panicked on multibyte inputs - Issue 4: Removed five stale #[allow(dead_code)] annotations (escape_verbatim, ref_label, format_utc_date, Bucket, bucket_for) — all consumed now - Issue 5: Collapsed nested if at line 285 via let-chain syntax instead of suppressing collapsible_if lint - Issue 6: Replaced cast_possible_truncation with u32::try_from(...).unwrap_or(0) for safe type conversion - Issue 7: Fixed u8 arithmetic overflow in panel comparison via u32::from() casts (winner_votes * 2) - Issue 8: Fixed full_report() fixture: changed supported from 2→3 to match three Judged(Supported) findings - Issue 9: Added regression tests for multibyte safety (truncate_claim, sanitize_reason) and escape validation Test results: 24/24 pass Clippy: Clean (0 warnings/errors) Fmt: Verified Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
Post-merge with main (PRs 147/148): fixture literals gain the book-lane contract fields, and the skipped renderer distinguishes the contract's BookSource skips (identifier found, no usable catalog record) from plain non-URL skips, with matching explainer vocabulary. Architecture map regenerated (20 crates). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fa8f9478ac
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if let Some(archive) = &finding.archive_of { | ||
| output.push_str(archive.as_str()); | ||
| } | ||
| output.push_str("] (live link: ["); | ||
| output.push_str(finding.provenance.url.as_str()); |
There was a problem hiding this comment.
Swap archive and live URLs in recovered output
For any citation recovered via archive fallback, try_archive_fallback sets alt_request.source_url to the archive URL and stamps archive_of with the original unreachable live URL (crates/sp42-citation/src/citation/page.rs:231-255). This renderer uses archive_of as the URL to update the citation to and labels provenance.url as the live link, so recovered appendix lines tell reviewers to replace the citation with the dead URL while calling the actual archive the live link. Please render finding.provenance.url as the archive target and finding.archive_of as the live URL.
Useful? React with 👍 / 👎.
…rchive_of the dead live URL The recovered-via-archive line told reviewers to update citations to the dead URL and labeled the archive as the live link (fields swapped; Codex P1, PR 154). Renderer, in-code fixtures, and the demo report now follow the contract: repair handles target the archive that was actually read, with the dead live URL named as what it replaces. Semantics pinned by a never-cite-the-dead-URL test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 62ff880842
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| output.push('.'); | ||
| } | ||
|
|
||
| if u32::from(finding.agreement.winner_votes) * 2 <= u32::from(finding.agreement.panel_size) { |
There was a problem hiding this comment.
Gate panel-split copy on a real panel
For no-model disagreement findings, such as the book-lane book_searched_not_supported_outcome that records PanelAgreement::new(0, 0), this condition is true (0 * 2 <= 0) and the appendix says the review panel split even though no panel voted. That mislabels deterministic search-inside misses as low-confidence model disagreements; require a meaningful panel before applying the minority/no-majority annotation.
Useful? React with 👍 / 👎.
| output.push_str(": "); | ||
| output.push_str(match skip.reason { | ||
| sp42_citation::SkippedReason::NonUrlSource => crate::copy::SKIPPED_NON_URL, | ||
| sp42_citation::SkippedReason::BookSource => crate::copy::SKIPPED_BOOK_UNRESOLVED, |
There was a problem hiding this comment.
Distinguish failed book lookups from misses
When a skipped book ref comes from an Open Library transport failure, SkippedReason::BookSource is still used while the actual LookupFailed vs NotFound detail lives in report.book_resolutions. Rendering every BookSource skip as SKIPPED_BOOK_UNRESOLVED tells reviewers the identifier matched no usable catalog record even when the lookup never completed, so saved reports lose the failure disclosure and imply a catalog miss incorrectly.
Useful? React with 👍 / 👎.
- The panel-split annotation requires a real panel: PanelAgreement(0, 0) no-model findings (deterministic book outcomes) are never labeled low-confidence. - BookSource skips join book_resolutions by ref id: a LookupFailed transport error renders as a tool failure, never as a catalog miss. Explainer vocabulary extended to match. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3e0e09b8ad
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| render_bucket(&mut output, report, Bucket::DeadLink); | ||
| render_bucket(&mut output, report, Bucket::Unreadable); | ||
| render_bucket(&mut output, report, Bucket::Unconfirmed); | ||
| render_bucket(&mut output, report, Bucket::Supported); |
There was a problem hiding this comment.
Render book resolution outcomes
When a URL-less book citation resolves to an Open Library record but has no exact searchable scan, verify_page emits a SourceUnavailable/Unusable finding plus report.book_resolutions carrying the scan state (similar_only/none/unknown). This renderer only emits the finding buckets here and never renders a Books section, so that case becomes a generic unreadable-source line and the reviewer loses whether the book was resolved and why it could not be searched. Please render book_resolutions alongside the GA appendix’s book/offline section.
Useful? React with 👍 / 👎.
| if lookup_failed_refs.contains(skip.ref_id.as_str()) { | ||
| crate::copy::SKIPPED_BOOK_LOOKUP_FAILED |
There was a problem hiding this comment.
Match book lookup failures per source
Fresh evidence: the new fix keys failed lookups only by ref_id. extract_use_sites emits one BookUseSite per book source for a URL-less ref while reusing the same ref_id; if one cited book in that ref has LookupFailed and another has NotFound, both SkippedRef entries satisfy this contains(ref_id) check and the clean catalog miss is rendered as a tool failure. Join to the matching BookResolution (for example by ref id plus identifiers/order) before choosing the copy.
Useful? React with 👍 / 👎.
Keeps render_ga_appendix under the line budget after the lookup-failed join, without an allow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8d0d3b30d7
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| .unwrap_or(after_prefix.len()); | ||
|
|
||
| let cite_id = &remaining[pos..pos + "cite_ref-".len() + token_end]; | ||
| result.push_str(&ref_label(cite_id, 0)); |
There was a problem hiding this comment.
Avoid rewriting every failed unnamed ref as #1
When an extraction failure reason contains an unnamed cite id (for example extract_use_sites emits strings like ref cite_ref-64 has no resolvable claim text), this always calls ref_label(cite_id, 0). Numeric cite ids fall back to ordinal + 1, so any later failed unnamed ref is rendered as ref #1, pointing the reviewer at the wrong citation instead of the failing one.
Useful? React with 👍 / 👎.
…okup join, distinct failure labels 1. Books consulted section: render book_resolutions (resolved outcomes only) with edition title and scan state (exact, similar, none, unknown) via new render_books_section() fn, positioned after skipped section before extraction failures. Adds 6 copy constants for scan-state vocabulary; keeps render_ga_appendix at 98 lines (under 100-line budget by extracting fn). 2. Per-source lookup-failure join: replace ref_id-only key with (ref_id AND identifiers) matching. For URL-less refs citing multiple books, each skip now joins to the matching BookResolution (not all LookupFaileds), so NotFound and LookupFailed outcomes render their distinct copy. Falls back to ref_id-only when skip.book_sources is empty. 3. Extraction-failure labels: unnamed cite_ref-<digits> now render as "an unnamed ref (block N)" using BlockFailure.block_ordinal (not global "ref #1"), so distinct failures stay distinct. Named refs stay on derived-name path. Sanitize_reason now takes block_ordinal parameter. Test coverage: added distinct_unnamed_refs_in_same_output_render_distinct_labels(), per_source_lookup_join_distinguishes_multiple_books_in_same_ref(), books_consulted_section_renders_resolved_editions(), plus block-anchored helper tests. All tests pass; cargo fmt, clippy -D warnings, check-design-system.sh, check-links.sh all pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2cdd129a17
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| sp42_citation::SkippedReason::BookSource => { | ||
| if let Some(sp42_citation::BookResolutionOutcome::LookupFailed { .. }) = | ||
| skip_outcome(skip) | ||
| { | ||
| crate::copy::SKIPPED_BOOK_LOOKUP_FAILED |
There was a problem hiding this comment.
Include identifiers for skipped book sources
When a single URL-less ref names multiple books, verify_book_use_site can produce multiple BookSource skips with the same ref_id, while the per-book identifiers are carried in SkippedRef.book_sources. This branch renders only the shared ref label plus generic failure/miss copy, so two failed/missed books in the same ref become indistinguishable and the appendix drops the only data that tells the reviewer which catalog lookup each line refers to. Include the book identifiers and/or cited page on BookSource skip lines.
Useful? React with 👍 / 👎.
Two missed/failed books under one ref were indistinguishable; skip lines now carry their reader-facing identifiers (ISBN/OCLC/LCCN/OL) and cited page (Codex round 4, PR 154). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c209182738
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| let ordinal = u32::try_from(skip.block_ordinal).unwrap_or(0); | ||
| output.push_str(&ref_label(&skip.ref_id, ordinal)); |
There was a problem hiding this comment.
Avoid using block indexes as ref numbers
When a skipped citation has an unnamed MediaWiki id such as cite_ref-9, this passes SkippedRef.block_ordinal into ref_label, but that field is the paragraph/block index from extraction, not the citation or use-site ordinal. For example, an anonymous non-URL ref in the first block renders as ref #1 even if the skipped ref id is cite_ref-9, so the GA appendix can point reviewers at the wrong citation; use block-anchored wording for skipped unnamed refs or carry the actual ref ordinal.
Useful? React with 👍 / 👎.
| pub const SKIPPED_BOOK_LOOKUP_FAILED: &str = | ||
| "cites a book whose catalog lookup did not complete — a tool failure, not a catalog miss"; |
There was a problem hiding this comment.
Remove pass/fail wording from lookup skips
When an Open Library transport error produces a BookSource skip, render_skipped_section emits this string, so the appendix contains the word failure. That violates the renderer's documented no pass/fail wording invariant for exactly the lookup-failed path that was added here, and it reintroduces verdict-sounding language for a tool transport problem; prefer neutral wording such as a lookup not completing or a tool limitation.
Useful? React with 👍 / 👎.
| let token_end = after_prefix | ||
| .find(|c: char| c.is_whitespace()) | ||
| .unwrap_or(after_prefix.len()); | ||
|
|
||
| let cite_id = &remaining[pos..pos + "cite_ref-".len() + token_end]; |
There was a problem hiding this comment.
Stop treating punctuated unnamed refs as ref #1
Fresh evidence: page verification errors are formatted as verify failed for {ref_id}: {error}, so an unnamed id is seen here as cite_ref-64: because the tokenizer stops only at whitespace. The colon makes rest non-numeric, which sends the value through ref_label(cite_id, 0) and renders it as ref #1 instead of the block-anchored unnamed-ref label, so later failed anonymous refs are again misidentified; strip trailing punctuation or parse the numeric prefix before this check.
Useful? React with 👍 / 👎.
- Skipped unnamed refs render block-anchored, never numbered from the paragraph index. - Failure-reason tokenizing strips trailing ASCII punctuation (cite_ref-64: no longer falls to the fabricated ref-number path), byte-safe on multibyte text. - The lookup-not-completed copy family drops pass/fail vocabulary everywhere (copy, scan-state, explainer), and the word-scan regression now renders every copy branch so the gap cannot reopen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2f6f97688c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| output.push_str(&escape_verbatim(&truncate_claim(&finding.claim, 80))); | ||
| output.push_str(" — ["); | ||
| output.push_str(finding.provenance.url.as_str()); | ||
| output.push_str("] (quote located)\n"); |
There was a problem hiding this comment.
Render book-scan provenance on findings
When a book citation is verified through Internet Archive search-inside, the report sets finding.book_scan with scanned_page, cited_page, and note; this line reduces a supported book finding to just “quote located.” If whole-book search finds the passage on page 32 while the citation says page=42, the GA appendix loses that page-mismatch disclosure and makes an imprecise page citation look cleanly supported, so the book-scan details/note should be rendered on finding lines.
Useful? React with 👍 / 👎.
| let ordinal = u32::try_from(resolution.block_ordinal).unwrap_or(0); | ||
| output.push_str(&ref_label(&resolution.ref_id, ordinal)); |
There was a problem hiding this comment.
Avoid using block ordinal as a book ref number
Fresh evidence: the skipped-ref path now avoids this, but the Books consulted section still passes BookResolution.block_ordinal to ref_label as if it were the citation ordinal. For unnamed book refs such as cite_ref-9, a book cited in block 0 renders as ref #1 even when the actual citation marker is not #1, pointing reviewers at the wrong footnote; use block-anchored wording here too or carry the real use-site/ref ordinal.
Useful? React with 👍 / 👎.
…section labels - Findings verified against an IA scan disclose a scanned-vs-cited page difference and any search note, on supported, unreadable, and dead-link lines alike. - The Books consulted section stops fabricating ref numbers from block ordinals; unnamed refs render block-anchored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5b23765802
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if let (Some(scanned), Some(cited)) = (scan.scanned_page, scan.cited_page.as_deref()) | ||
| && scanned.to_string() != cited | ||
| { | ||
| parts.push(crate::copy::book_scan_pages(scanned, cited)); |
There was a problem hiding this comment.
Escape cited book pages in scan annotations
When a book citation's cited page differs from the scanned page, cited_page comes directly from the cite template and can contain arbitrary wikitext; this path formats it through book_scan_pages and later writes the returned annotation without escape_verbatim. A page value such as 42</nowiki>{{...}} would break out of the pasted appendix even though other verbatim fields are nowiki-escaped, so escape the cited page before putting it in the annotation.
Useful? React with 👍 / 👎.
| output.push_str(" — ["); | ||
| output.push_str(finding.provenance.url.as_str()); | ||
| output.push(']'); |
There was a problem hiding this comment.
Append book-scan notes outside supported lines
Fresh evidence: book_scan_annotation is now only appended in the dead-link, unreadable, and supported render paths. Book search-inside results can also land in this disagreement path, for example NoMatches creates a NotSupported finding with a book_scan.note of “searched, no matching passage”, and model Partial/NotSupported book findings carry scanned-vs-cited page provenance; those appendix lines currently drop the only disclosure that the book scan was searched and whether the cited page matched. Please append the book-scan annotation in the disagreement and unconfirmed render paths too.
Useful? React with 👍 / 👎.
… disagreements - Book-scan annotations escape the template-sourced cited page (hostile wikitext regression added) and render on disagreement and unconfirmed lines too (searched-no-passage notes were dropped there). - Architecture map regenerated (CI freshness gate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 46fcc2852b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| for failure in &report.extraction_failures { | ||
| output.push_str("* "); | ||
| let rewritten = sanitize_reason(&failure.reason, failure.block_ordinal); | ||
| output.push_str(&escape_verbatim(&rewritten)); |
There was a problem hiding this comment.
Neutralize failure wording in extraction reasons
When a URL or book use-site verification errors, verify_page records BlockFailure.reason strings such as verify failed for {ref_id}: ... and book verify failed for ...; this line only rewrites cite IDs and then emits the reason verbatim. Those reports therefore produce GA appendix bullets containing failed, which breaks the renderer's no pass/fail wording guarantee and makes a tool error read like a citation verdict, so the reason should be mapped through neutral copy before rendering.
Useful? React with 👍 / 👎.
| )); | ||
| } | ||
| if let Some(note) = &scan.note { | ||
| parts.push(escape_verbatim(note)); |
There was a problem hiding this comment.
Scrub transport-failure wording from scan notes
When Internet Archive metadata or search-inside requests hit a transport error, prepare_book_grounding stores error.to_string() in book_scan.note, and HttpClientError formats those as transport failed: .... Appending that note verbatim here means dead/unreadable book lines can still leak banned failure wording despite the copy invariant; render these tool notes through neutral wording rather than the raw error string.
Useful? React with 👍 / 👎.
…ontract strings Failure reasons and book-scan notes are tool-authored strings; the copy module now rewrites their verdict-sounding words (failed/failure/pass) into neutral phrasing before rendering, per the appendix's whole-output wording invariant. The maximal word-scan regression exercises real contract phrasings. Also fixes the rustfmt drift CI caught. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d8c72fd77a
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if name.is_empty() || !seq.chars().all(|c| c.is_ascii_digit()) { | ||
| return fallback; | ||
| } | ||
| format!("ref \"{}\"", name.replace('_', " ")) |
There was a problem hiding this comment.
Escape derived ref names before emitting labels
When a named ref id contains markup-like characters, this returns that derived name directly (for example cite_ref-{{x}}_1-0 becomes a live ref "{{x}}" label), and every bucket writes the label into pasteable wikitext without escape_verbatim. Because ref names come from article content or saved report JSON while claims/quotes are explicitly escaped, a named ref with wikitext in its name can transclude or break the pasted appendix; encode or nowiki-wrap the derived name before returning the label.
Useful? React with 👍 / 👎.
- Ref names ride the cite id and are article-authored: labels now escape them like every verbatim field, with a hostile-name regression (Codex round 9, PR 154). - The architecture-map freshness check prints a unified diff on failure: CI and local disagree on staleness at identical commits, and the log needs to show the actual delta to make that diagnosable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8ea7b6966f
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| BookIdentifier::Isbn(value) => format!("ISBN {value}"), | ||
| BookIdentifier::Oclc(value) => format!("OCLC {value}"), | ||
| BookIdentifier::Lccn(value) => format!("LCCN {value}"), | ||
| BookIdentifier::Olid(value) => format!("Open Library id {value}"), |
There was a problem hiding this comment.
When render-report deserializes a saved PageVerificationReport, BookIdentifier variants are accepted from JSON directly, so values do not necessarily pass the constructor validation used by live extraction. These raw format! arms are later pasted into the GA wikitext for skipped book refs without escape_verbatim; a crafted saved report with an ISBN value like {{...}} would transclude or break markup even though claims and pages are nowiki-escaped. Escape the identifier value before returning the label.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7115877302
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| } | ||
| CitationVerdict::SourceUnavailable => match finding.source_unavailable_reason { | ||
| Some(SourceUnavailableReason::Unusable) => Bucket::Unreadable, | ||
| Some(SourceUnavailableReason::Unreachable) | None => Bucket::DeadLink, |
There was a problem hiding this comment.
Keep book scan outages out of dead-link bucket
When a resolved book's Internet Archive metadata/search-inside request cannot be reached, ground_resolved_book's BookGroundingPreparation::Unreachable arm creates a SourceUnavailable finding with book_scan provenance and a status-0 provenance; this branch then files it under DeadLink. In that scenario the GA appendix tells reviewers the citation link may be dead even though only the scan/search tool request failed, so route book-scan transport outages to the tool-limitation/book path instead of the dead-link bucket.
Useful? React with 👍 / 👎.
| // Footer | ||
| output.push_str("\n----\n"); | ||
| output.push_str("''"); | ||
| output.push_str(&report.title); |
There was a problem hiding this comment.
Escape the report title in the footer
When render-report renders a saved report, report.title is deserialized file data but is written directly into live italic wikitext here. A title/value containing wiki markup such as '' or {{...}} can break the footer formatting or transclude when the appendix is pasted, while the other verbatim report fields are nowiki-escaped; escape the title before appending it.
Useful? React with 👍 / 👎.
render-report input is deserialized data: the footer title and book identifier values escape like every verbatim field, and a book-scan outage buckets as a tool limitation, never a dead article link. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8f154e9c4e
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| report.stats.source_unavailable_unreachable, | ||
| plural(report.stats.source_unavailable_unreachable, "link", "links"), | ||
| report.stats.source_unavailable_unusable, |
There was a problem hiding this comment.
Keep book-scan outages out of the summary
When a resolved book's Internet Archive metadata/search request is unreachable, bucket_for now correctly routes the book_scan finding to the unreadable/tool-limitation bucket, but the summary still prints report.stats.source_unavailable_unreachable as dead links. tally_stats counts those status-0 book-scan outages as source_unavailable_unreachable, so the appendix can say there is a dead link even though the only outage was the scan/search tool request; derive the dead/unreadable counts with the same book-scan override used for buckets or adjust the stats before rendering.
Useful? React with 👍 / 👎.
The stats line's dead-link and unreadable counts now come from the same bucket_for the sublists use, so a book-scan outage counted as unreachable in raw stats can never make the summary disagree with the lines below it (Codex round 12, PR 154). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 792d393592
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if !availability.exact.is_empty() | ||
| && availability.exact.iter().any(|item| item.ocaid.is_some()) | ||
| { | ||
| crate::copy::BOOK_SCAN_EXACT | ||
| } else if !availability.similar.is_empty() { |
There was a problem hiding this comment.
Distinguish ungroundable exact scans
When Open Library reports an exact-edition scan but the scan identity could not be recovered (ScanAvailability.exact is non-empty and all ocaids are None), the references code deliberately keeps that as an exact item while groundable_scan() returns None. This condition treats it as if no exact scan exists, so with no similar items the Books consulted section says the book is “not scanned in the tool's catalog” even though the actual reason is that the reported exact scan could not be grounded; reviewers lose the correct explanation for why the citation was not searched.
Useful? React with 👍 / 👎.
… state An exact scan with no recoverable identity was reported as 'not scanned'; the Books section now distinguishes it as cataloged but unsearchable (Codex round 13, PR 154). Explainer vocabulary synced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cbae9c6018
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if let Some(title) = &edition.title { | ||
| output.push_str(&escape_verbatim(title)); | ||
| } else { | ||
| output.push_str("(untitled book)"); |
There was a problem hiding this comment.
Include identifiers for consulted books
When a resolved Open Library edition has no title (or two resolved cited books share the same title), this Books consulted line renders only (untitled book)/the title plus scan state, dropping the BookResolution identifiers and cited page that distinguish which book was actually consulted. In that scenario the appendix can contain indistinguishable book lines and gives the reviewer no way to map the scan-state disclosure back to the cited source; include the resolved/cited identifier and page here as the skipped-book section does.
Useful? React with 👍 / 👎.
| if unnamed { | ||
| format!("an unnamed ref (block {block_ordinal})") | ||
| } else { | ||
| ref_label(ref_id, 0) |
There was a problem hiding this comment.
Anchor unparseable skipped ref ids by block
Fresh evidence: sp42-parsoid preserves marker ids verbatim and its unnamed-ref fixture uses a non-numeric id like cite_ref-test-1 with no ref name. For skipped/book refs with that shape, this call passes ordinal 0 into ref_label; the parse misses the named-ref pattern and renders ref #1, even when the skipped citation is later in the article. Since these paths do not have a real use-site ordinal, parse-miss cite_ref-* labels should stay block-anchored instead of fabricating footnote #1.
Useful? React with 👍 / 👎.
…ck-anchored unparseable ids - Books consulted lines carry identifiers and cited page, keeping same-title or untitled resolutions distinguishable. - Label helpers refactored: the named-parse is its own fallible helper, and skip labels anchor unparseable ids by block instead of fabricating ref numbers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5cd1efcf62
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| pub fn disagreement_verdict(level: SupportLevel) -> &'static str { | ||
| match level { | ||
| SupportLevel::NotSupported => { | ||
| "the source and this claim disagree — the panel found no support for the claim in the source" |
There was a problem hiding this comment.
Avoid attributing no-model book misses to the panel
When Internet Archive search-inside returns no matches, book_searched_not_supported_outcome creates a NotSupported finding with PanelAgreement::new(0, 0), so no model panel voted. This copy is still used for that finding and tells reviewers that “the panel found” no support, misrepresenting a deterministic book-search miss as a panel judgment; branch the wording for no-panel findings or avoid panel attribution here.
Useful? React with 👍 / 👎.
| output.push_str(": the tool fetched ["); | ||
| output.push_str(finding.provenance.url.as_str()); | ||
| output.push_str("] but read "); |
There was a problem hiding this comment.
Use outage-specific copy for unreachable book scans
When ground_resolved_book hits BookGroundingPreparation::Unreachable, it emits a SourceUnavailable finding with book_scan and status 0, and bucket_for routes every book_scan finding into this unreadable path. In that scenario the appendix says the tool fetched the scan URL and read an unusable page, even though the metadata/search request never completed; render a separate book-scan outage message for book_scan + unreachable/status-0 findings.
Useful? React with 👍 / 👎.
…rue copy - Deterministic book-search misses (no panel voted) say the source was searched, never that a panel judged it. - Unreachable book scans render as could-not-reach, never as fetched-but- unusable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8efba5362c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| output.push_str(": supported — "); | ||
| output.push_str(&escape_verbatim(&truncate_claim(&finding.claim, 80))); | ||
| output.push_str(" — ["); | ||
| output.push_str(finding.provenance.url.as_str()); |
There was a problem hiding this comment.
Escape URLs before embedding external links
When a citation URL contains a literal ] in its query or fragment, writing as_str() inside MediaWiki's [url] syntax closes the external link early; for example, a saved report URL like https://example.test/?q=]''x would leave ''x as live wikitext in the pasted appendix. The same raw-URL pattern is used in the other buckets too, so route these through a wikitext-link escaping helper (at least percent-encode ]) before emitting them.
Useful? React with 👍 / 👎.
origin/main does not compile: PR #153 added SkippedReason::UnresolvedShortCite while PR #154 added an exhaustive match over SkippedReason in ga_appendix.rs; they merged independently and the semantic conflict landed on main. This adds the missing arm (with disclosed-not-guessed copy) so this branch builds; main needs the same one-arm fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QB2VfS9GXyTzVw3FayQZxf
A GA reviewer who runs verify-page gets a machine-shaped report; to use it on a review page they hand-transcribe findings, losing the grounding (quotes, ref addressing) and usually paraphrasing honest evidence into verdict-sounding language. This adds the renderer that turns a report into the wikitext appendix a reviewer pastes onto Talk:Article/GAn — evidence, never verdicts.
New domain crate
sp42-assessment(the assessment domain's first): one pure function,PageVerificationReport → wikitext, deterministic given a render timestamp. The criterion-2 section orders findings by consequence for the review — claim–source disagreements first, then archive recoveries, dead links, tool-unreadable sources, unconfirmed supports, a compact supported spot-check record, and first-class skips/extraction failures. Wording invariants are tests, not conventions: no pass/fail language anywhere, mismatch framing for disagreements, non-exact grounding always visible, no raw contract identifiers, and every verbatim field nowiki-escaped against hostile markup including embedded</nowiki>terminators. A footer carries the revision, render date, version, framing line, and a what-is-this explainer link for the cold reader.CLI surfaces:
verify-page --format ga-appendix(bridge convenience) and a new purerender-report <file>that re-renders a savedverify-page --format jsonreport with no server, session, or network — the replayable core. The format enum is command-local soga-appendixcannot appear on commands that can't produce one.Two rendering decisions beyond the PRD's literal sublist spec, recorded in the PRD changelog for strike-or-keep:
Fixtures include a real-article-shaped saved report and a hostile-markup fixture; DoD items 15/16 are satisfied structurally (both surfaces call the same pure function; there are no clients to mock) per the changelog note. Branch is merged with current main, including reconciliation onto the book lane (fixture fields; distinct copy for the contract's book-source skips).
Design: docs/design-plans/2026-07-10-ga-appendix-renderer.md. The reader-facing copy is drafted for the alpha review with real GA reviewers ({{GAList}} native-idiom question included there).
🤖 Generated with Claude Code
https://claude.ai/code/session_01P2RnaXqFdmpu93tzrRbSaj