Repository navigation
feat(maintenance): add orcid name-variant maintenance under sync/maintenance - #534
Merged
Merged
Conversation
…tenance Two programs that correct and enrich given-name lists in the orcid collection, plus the libraries they share and tests for both write paths. DOI author matching compares publisher metadata against each entry in a record's given list, and publishers are inconsistent about the period, so a name held only as "Gerald M. Rubin" silently fails to match everyone who deposited "Gerald M Rubin". Both spellings are now stored. add_orcid_name_variants.py pulls variants orcid.org holds that we do not fix_middle_names.py permutes what is already there (2.0.0) dis_name_lib.py shared name rules dis_review_lib.py shared interactive reviewer These are not nightly synchronization: nothing upstream depends on them and a skipped run costs nothing, which is why they sit apart from sync/bin. fix_middle_names.py moves here from etl/bin and is rewritten. The earlier version wrote the whole given array back with $set, which silently dropped anything apply_orcids.py had added since the read; it also guarded per record rather than per name, left the period unescaped in a regex, and was ASCII-only. Those duplicate entries are still in the data - this run found and collapsed four. Writes use $addToSet, never a whole-array $set, because apply_orcids.py and these two programs all write the same lists. The one write that needs the whole array, collapsing duplicates under --dedupe, re-reads the record immediately beforehand. Fixes a defect in the review path found while testing it: interactive_session returned early whenever there were no name candidates, without consulting the pending dedupe list, so --review --dedupe --write silently skipped them. The check fifteen lines further down already treated pending dedupes as reason enough to proceed. Current data has exactly this shape - no name candidates and four records to collapse. tests/test_name_maintenance.py covers both write paths offline, asserting the update shape directly: a single-writer test cannot tell $addToSet from $set, and that distinction is the whole safety argument here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two programs that correct and enrich given-name lists in the orcid collection, plus the libraries they share and tests for both write paths.
DOI author matching compares publisher metadata against each entry in a record's given list, and publishers are inconsistent about the period, so a name held only as "Gerald M. Rubin" silently fails to match everyone who deposited "Gerald M Rubin". Both spellings are now stored.
add_orcid_name_variants.py pulls variants orcid.org holds that we do not
fix_middle_names.py permutes what is already there (2.0.0)
dis_name_lib.py shared name rules
dis_review_lib.py shared interactive reviewer
These are not nightly synchronization: nothing upstream depends on them and a skipped run costs nothing, which is why they sit apart from sync/bin.
fix_middle_names.py moves here from etl/bin and is rewritten. The earlier version wrote the whole given array back with $set, which silently dropped anything apply_orcids.py had added since the read; it also guarded per record rather than per name, left the period unescaped in a regex, and was ASCII-only. Those duplicate entries are still in the data - this run found and collapsed four.
Writes use $addToSet, never a whole-array $set, because apply_orcids.py and these two programs all write the same lists. The one write that needs the whole array, collapsing duplicates under --dedupe, re-reads the record immediately beforehand.
Fixes a defect in the review path found while testing it: interactive_session returned early whenever there were no name candidates, without consulting the pending dedupe list, so --review --dedupe --write silently skipped them. The check fifteen lines further down already treated pending dedupes as reason enough to proceed. Current data has exactly this shape - no name candidates and four records to collapse.
tests/test_name_maintenance.py covers both write paths offline, asserting the update shape directly: a single-writer test cannot tell $addToSet from $set, and that distinction is the whole safety argument here.
@stuarteberg