Repository navigation
feat(awards): report what each grant produced - #538
Merged
Merged
Conversation
Adds /awards and /award/<funder>/<award>: the DOIs grouped by the award that funded them. Nothing surfaced this before, though the data has been there since funder identifiers landed - 3,840 awards across 1,290 DOIs, 478 of them naming more than one output. Grouped by funder and award together. An award number is only unique within its funder: 2014862 is an NSF grant and a plausible identifier anywhere else. NIH deposits the same grant under several spellings. The leading digit is the application type (1 new, 2 renewal, 5 continuation) and the trailing suffix is the budget period, so 1U19NS104648, 5U19NS104648 and "U19 NS104648" are one grant and are folded together. The heading shows whichever spelling was deposited most often, and the drill-down lists them all, because somebody reconciling this against a grants database needs to know the deposits were not uniform. Two things are deliberately not treated as grants: Filler. N/A, none, "Janelia Research Campus" and the like appear in the award field on 73 entries. Grouping on those would invent a grant that funded two dozen unrelated papers. Program names. 230 awards carry a name rather than a number - Investigator, Odyssey Award, Fellowship. These are real awards and are still listed, but many people hold one, so the DOIs beneath them are not one grant's output. They are excluded from the multi-output count and say so on their own page. Counting them as grants overstated that figure by 36. The award page re-reads the full documents by DOI rather than rendering what the grouping loop collected. standard_doi_table goes through DL.get_title and DL.get_publishing_date, which read the raw registrar fields - Crossref's published/issued, DataCite's date structures - rather than the jrc_ ones, so a projection written for the grouping loop leaves both empty. The index itself projects only doi and jrc_funder, since it counts and groups and renders no record of its own. The funder page links to its own awards, and /awards?funder=<id> narrows to one funder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds /awards and /award//: the DOIs grouped by the award that funded them. Nothing surfaced this before, though the data has been there since funder identifiers landed - 3,840 awards across 1,290 DOIs, 478 of them naming more than one output.
Grouped by funder and award together. An award number is only unique within its funder: 2014862 is an NSF grant and a plausible identifier anywhere else.
NIH deposits the same grant under several spellings. The leading digit is the application type (1 new, 2 renewal, 5 continuation) and the trailing suffix is the budget period, so 1U19NS104648, 5U19NS104648 and "U19 NS104648" are one grant and are folded together. The heading shows whichever spelling was deposited most often, and the drill-down lists them all, because somebody reconciling this against a grants database needs to know the deposits were not uniform.
Two things are deliberately not treated as grants:
Filler. N/A, none, "Janelia Research Campus" and the like appear in the award field on 73 entries. Grouping on those would invent a grant that funded two dozen unrelated papers.
Program names. 230 awards carry a name rather than a number - Investigator, Odyssey Award, Fellowship. These are real awards and are still listed, but many people hold one, so the DOIs beneath them are not one grant's output. They are excluded from the multi-output count and say so on their own page. Counting them as grants overstated that figure by 36.
The award page re-reads the full documents by DOI rather than rendering what the grouping loop collected. standard_doi_table goes through DL.get_title and DL.get_publishing_date, which read the raw registrar fields - Crossref's published/issued, DataCite's date structures - rather than the jrc_ ones, so a projection written for the grouping loop leaves both empty. The index itself projects only doi and jrc_funder, since it counts and groups and renders no record of its own.
The funder page links to its own awards, and /awards?funder= narrows to one funder.
@stuarteberg