Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
A Pseudonymized Corpus of Occupational Health Narratives for Clinical Entity Recognition in Spanish
Pontificia Universidad Catolica de Chile, Santiago, Chile.
Stockholm University, Faculty of Social Sciences, Department of Computer and Systems Sciences.ORCID iD: 0000-0001-8988-8226
Pontificia Universidad Catolica de Chile, Santiago, Chile.
Universidad de Chile, Santiago, Chile.
Show others and affiliations
Number of Authors: 92024 (English)In: BMC Medical Informatics and Decision Making, E-ISSN 1472-6947, no 24, article id 204Article in journal (Refereed) Published
Abstract [en]

Despite the high creation cost, annotated corpora are indispensable for robust natural language processing systems. In the clinical field, in addition to annotating medical entities, corpus creators must also remove personally identifiable information (PII). This has become increasingly important in the era of large language models where unwanted memorization can occur. This paper presents a corpus annotated to anonymize personally identifiable information in 1,787 anamneses of work-related accidents and diseases in Spanish. Additionally, we applied a previously released model for Named Entity Recognition (NER) trained on referrals from primary care physicians to identify diseases, body parts, and medications in this work-related text. We analyzed the differences between the models and the gold standard curated by a physician in detail. Moreover, we compared the performance of the NER model on the original narratives, in narratives where personal information has been masked, and in texts where the personal data is replaced by another similar surrogate value (pseudonymization). Within this publication, we share the annotation guidelines and the annotated corpus.

Place, publisher, year, edition, pages
2024. no 24, article id 204
Keywords [en]
Natural language processing, Privacy, Named entity recognition, Corpus annotation
National Category
Natural Language Processing
Research subject
Computer and Systems Sciences
Identifiers
URN: urn:nbn:se:su:diva-232090DOI: 10.1186/s12911-024-02609-wISI: 001275573100002PubMedID: 39049027Scopus ID: 2-s2.0-85199343231OAI: oai:DiVA.org:su-232090DiVA, id: diva2:1885702
Available from: 2024-07-24 Created: 2024-07-24 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Vakili, Thomas

Search in DiVA

By author/editor
Vakili, Thomas
By organisation
Department of Computer and Systems Sciences
In the same journal
BMC Medical Informatics and Decision Making
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 221 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf