Skip to content

feat(orcid): report authors whose deposited ORCID names someone else - #532

Merged
robsv merged 1 commit into
mainfrom
feature-orcid-name-mismatch
Oct 7, 2026
Merged

robsv merged 1 commit into
mainfrom
feature-orcid-name-mismatch

Conversation

@robsv

@robsv robsv commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Adds /orcid_mismatch under Authorship. It walks every author entry carrying an ORCID that is on the Janelia roster and reports the ones where the deposited name shares no token with the name we hold for that ORCID - the credit has landed on the wrong Janelian.

Found while chasing an extra jrc_author on 10.7554/elife.93659: eLife's deposit carries one person's ORCID on a different author's entry, so every refresh credited someone who is not on the paper.

Comparing names naively gives 137 hits, nearly all of them one person written two ways. Three things cut that to 6 real findings:

  • Names are accent-folded through NFKD, so Loesche and Losche match.
  • An inverted "Family, Given" deposit is uninverted before comparison.
  • A single shared token is enough to call it the same person. That covers a swapped name order, a middle initial, a dropped accent and a misspelt surname - all the publisher rendering one person differently, rather than naming another.

The ORCID itself is matched exactly; rapidfuzz only ranks the output, so the report still works if the library is missing.

The page says plainly that removing the credit will not make it stay removed: the error is in the deposit, and update_dois.py re-reads the same ORCID on the next refresh. The fix is a publisher correction.

@stuarteberg

Adds /orcid_mismatch under Authorship. It walks every author entry
carrying an ORCID that is on the Janelia roster and reports the ones
where the deposited name shares no token with the name we hold for that
ORCID - the credit has landed on the wrong Janelian.

Found while chasing an extra jrc_author on 10.7554/elife.93659: eLife's
deposit carries one person's ORCID on a different author's entry, so
every refresh credited someone who is not on the paper.

Comparing names naively gives 137 hits, nearly all of them one person
written two ways. Three things cut that to 6 real findings:

- Names are accent-folded through NFKD, so Loesche and Losche match.
- An inverted "Family, Given" deposit is uninverted before comparison.
- A single shared token is enough to call it the same person. That
  covers a swapped name order, a middle initial, a dropped accent and a
  misspelt surname - all the publisher rendering one person differently,
  rather than naming another.

The ORCID itself is matched exactly; rapidfuzz only ranks the output, so
the report still works if the library is missing.

The page says plainly that removing the credit will not make it stay
removed: the error is in the deposit, and update_dois.py re-reads the
same ORCID on the next refresh. The fix is a publisher correction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@robsv robsv self-assigned this Oct 7, 2026
@robsv robsv added the enhancement New feature or request label Oct 7, 2026
@robsv
robsv merged commit 2f74c53 into main Oct 7, 2026
2 checks passed
@robsv
robsv deleted the feature-orcid-name-mismatch branch October 7, 2026 15:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant