Skip to content

feat(maintenance): add orcid name-variant maintenance under sync/maintenance - #534

Merged
robsv merged 1 commit into
mainfrom
feature-orcid-name-maintenance
Oct 7, 2026
Merged

robsv merged 1 commit into
mainfrom
feature-orcid-name-maintenance

Conversation

@robsv

@robsv robsv commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Two programs that correct and enrich given-name lists in the orcid collection, plus the libraries they share and tests for both write paths.

DOI author matching compares publisher metadata against each entry in a record's given list, and publishers are inconsistent about the period, so a name held only as "Gerald M. Rubin" silently fails to match everyone who deposited "Gerald M Rubin". Both spellings are now stored.

add_orcid_name_variants.py pulls variants orcid.org holds that we do not
fix_middle_names.py permutes what is already there (2.0.0)
dis_name_lib.py shared name rules
dis_review_lib.py shared interactive reviewer

These are not nightly synchronization: nothing upstream depends on them and a skipped run costs nothing, which is why they sit apart from sync/bin.

fix_middle_names.py moves here from etl/bin and is rewritten. The earlier version wrote the whole given array back with $set, which silently dropped anything apply_orcids.py had added since the read; it also guarded per record rather than per name, left the period unescaped in a regex, and was ASCII-only. Those duplicate entries are still in the data - this run found and collapsed four.

Writes use $addToSet, never a whole-array $set, because apply_orcids.py and these two programs all write the same lists. The one write that needs the whole array, collapsing duplicates under --dedupe, re-reads the record immediately beforehand.

Fixes a defect in the review path found while testing it: interactive_session returned early whenever there were no name candidates, without consulting the pending dedupe list, so --review --dedupe --write silently skipped them. The check fifteen lines further down already treated pending dedupes as reason enough to proceed. Current data has exactly this shape - no name candidates and four records to collapse.

tests/test_name_maintenance.py covers both write paths offline, asserting the update shape directly: a single-writer test cannot tell $addToSet from $set, and that distinction is the whole safety argument here.

@stuarteberg

…tenance

Two programs that correct and enrich given-name lists in the orcid
collection, plus the libraries they share and tests for both write paths.

DOI author matching compares publisher metadata against each entry in a
record's given list, and publishers are inconsistent about the period, so
a name held only as "Gerald M. Rubin" silently fails to match everyone who
deposited "Gerald M Rubin". Both spellings are now stored.

  add_orcid_name_variants.py  pulls variants orcid.org holds that we do not
  fix_middle_names.py         permutes what is already there (2.0.0)
  dis_name_lib.py             shared name rules
  dis_review_lib.py           shared interactive reviewer

These are not nightly synchronization: nothing upstream depends on them and
a skipped run costs nothing, which is why they sit apart from sync/bin.

fix_middle_names.py moves here from etl/bin and is rewritten. The earlier
version wrote the whole given array back with $set, which silently dropped
anything apply_orcids.py had added since the read; it also guarded per
record rather than per name, left the period unescaped in a regex, and was
ASCII-only. Those duplicate entries are still in the data - this run found
and collapsed four.

Writes use $addToSet, never a whole-array $set, because apply_orcids.py and
these two programs all write the same lists. The one write that needs the
whole array, collapsing duplicates under --dedupe, re-reads the record
immediately beforehand.

Fixes a defect in the review path found while testing it: interactive_session
returned early whenever there were no name candidates, without consulting the
pending dedupe list, so --review --dedupe --write silently skipped them. The
check fifteen lines further down already treated pending dedupes as reason
enough to proceed. Current data has exactly this shape - no name candidates
and four records to collapse.

tests/test_name_maintenance.py covers both write paths offline, asserting the
update shape directly: a single-writer test cannot tell $addToSet from $set,
and that distinction is the whole safety argument here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@robsv robsv self-assigned this Oct 7, 2026
@robsv robsv added the enhancement New feature or request label Oct 7, 2026
@robsv
robsv merged commit 43c1046 into main Oct 7, 2026
0 of 2 checks passed
@robsv
robsv deleted the feature-orcid-name-maintenance branch October 7, 2026 16:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant