Source-linked AI summary

A Comparison of Word Embeddings for the Biomedical Natural Language Processing

Yanshan Wang, Sijia Liu, Naveed Afzal, Majid Rastegar-Mojarad, Liwei Wang, Feichen Shen, Paul Kingsbury, Hongfang Liu

arXiv:1802.00400v3cs.IR

TL;DR

Biomedical word embeddings had been trained from varied text resources, but their comparative value for biomedical NLP was not well established. This paper evaluates embeddings from clinical notes, biomedical publications, Wikipedia, and news using qualitative, intrinsic, and extrinsic tests, finding that domain-trained embeddings better capture medical semantics while no source consistently ranks best across downstream tasks.

  • Problem

    Comparative evidence on word embeddings trained from different biomedical and general text resources was limited.

  • Method

    The study qualitatively and quantitatively evaluates embeddings trained from clinical notes, biomedical publications, Wikipedia, and news.

  • Results

    Clinical-note and biomedical-publication embeddings better capture medical semantics, while no embedding source consistently ranks best across downstream biomedical NLP applications.

  • Takeaways & Limitations

    Adding word embeddings as extra features improves results on most downstream biomedical NLP tasks, but biomedical-domain training does not guarantee superior performance.

  • Takeaways & Limitations

    The study examined only embeddings trained on the EHR corpus described by the authors.

Abstract

from arXiv · show

Word embeddings have been widely used in biomedical Natural Language Processing (NLP) applications as they provide vector representations of words capturing the semantic properties of words and the linguistic relationship between words. Many biomedical applications use different textual resources (e.g., Wikipedia and biomedical articles) to train word embeddings and apply these word embeddings to downstream biomedical applications. However, there has been little work on evaluating the word embeddings trained from these resources.In this study, we provide an empirical evaluation of word embeddings trained from four different resources, namely clinical notes, biomedical publications, Wikipedia, and news. We performed the evaluation qualitatively and quantitatively. For the qualitative evaluation, we manually inspected five most similar medical words to a given set of target medical words, and then analyzed word embeddings through the visualization of those word embeddings. For the quantitative evaluation, we conducted both intrinsic and extrinsic evaluation. Based on the evaluation results, we can draw the following conclusions. First, the word embeddings trained on clinical notes and biomedical publications can capture the semantics of medical terms better, and find more relevant similar medical terms, and are closer to human experts' judgments, compared to these trained on Wikipedia and news. Second, there does not exist a consistent global ranking of word embedding quality for downstream biomedical NLP applications. However, adding word embeddings as extra features will improve results on most downstream tasks. Finally, the word embeddings trained on biomedical domain corpora do not necessarily have better performance than those trained on other general domain corpora for any downstream biomedical NLP tasks.

I. INTRODUCTION · II. RELATED WORK

The paper motivates evaluating biomedical word embeddings across clinical, publication, Wikipedia, and news corpora because prior evaluations largely focused on general NLP. It extends limited biomedical-domain work through broader intrinsic, qualitative, and downstream-task comparisons.

  • I. INTRODUCTION: Word embeddings represent semantic properties and linguistic relationships and are widely used as features in NLP and biomedical applications.Biomedical uses include named entity recognition, synonym extraction, relation extraction, information retrieval, and abbreviation disambiguation.
  • I. INTRODUCTION: Biomedical embeddings can be trained from internal task corpora or external resources, but it remained unclear whether task-specific training is necessary.This question is especially important clinically because little electronic health record data is publicly available, whereas biomedical literature is more accessible.
  • I. INTRODUCTION: The study addresses a gap in evaluating embeddings from biomedical textual resources for biomedical NLP applications.The authors note little prior work evaluating embeddings trained from these resources and assess downstream clinical information extraction, biomedical information retrieval, and relation extraction.
  • I. INTRODUCTION: The study evaluates embeddings trained from clinical notes, biomedical publications, Wikipedia, and news using qualitative and quantitative methods.Qualitative analysis inspects similar medical terms and visualizes 377 medical words; quantitative analysis includes intrinsic and extrinsic evaluation.
  • II. RELATED WORK: Prior embedding studies mainly compared models on general-domain semantic, syntactic, and application-oriented benchmarks, with CBOW often among the strongest performers.Related evaluations covered fourteen benchmark datasets and downstream tasks; other work found dependency-based embeddings strongest on NLP tasks and combinations significantly improved results.
  • II. RELATED WORK: Prior guidance recommends evaluating both syntactic and semantic properties with tasks that approximate real-world applications, yet few studies addressed biomedical NLP.Most earlier studies evaluated general rather than biomedical NLP tasks.
  • II. RELATED WORK: Before this study, Pakhomov et al. compared clinical-note, biomedical-publication, and GloVe embeddings on medical similarity, retrieval, and word-sense disambiguation.They found biomedical-publication embeddings captured semantics on par with clinical-note embeddings.
  • II. RELATED WORK: This work extends prior biomedical evaluation with four medical-term datasets, qualitative inspection, and more downstream applications using shared-task data.The extensions target medical-term semantics, qualitative analysis, and broader biomedical NLP applications.

III. WORD EMBEDDINGS AND PARAMETER SETTINGS

The study used skip-gram word2vec with negative sampling to learn word embeddings. Embedding dimensions and training parameters were selected based on intrinsic-evaluation performance, prior studies, or public availability.

  • Model choice: The study used word2vec with the skip-gram architecture, chosen because neither skip-gram nor CBOW was known to outperform the other.Word2vec was selected based on reported performance on general NLP tasks.
  • Training method: Negative sampling approximated the full-vocabulary conditional probability by sampling a few output words and updating embeddings for that smaller sample.The negative-sampling distribution generated k negative samples according to term frequency.
  • Parameter settings: The selected embedding dimensions were 100 for EHR, 60 for MedLit, 100 for GloVe, and 300 for Google News.EHR and MedLit dimensions were chosen using intrinsic-evaluation performance, while 300 was the only publicly available Google News dimension.
  • Parameter settings: For EHR and MedLit training, the window size was 5, minimum word frequency was 7, and negative sampling was set to 5.These parameters were selected based on previous studies.

IV. DATA AND TEXT PRE-PROSESSING

The study evaluates embeddings trained from EHR clinical notes and MedLit biomedical literature alongside Google News and GloVe resources. EHR and MedLit underwent minimal normalization, with additional cleaning tailored to clinical-note structure.

  • Comparison embeddings: Google News and GloVe provide general-domain comparison embeddings, with vocabularies of 3 million and 400k words, respectively.Google News vectors were trained with word2vec, while GloVe used Wikipedia and Gigaword data.
  • General preprocessing: MedLit and EHR preprocessing removed punctuation, lowercased text, replaced digits with “7,” and normalized hyphen-connected words.MedLit additionally removed website URLs, email addresses, and Twitter handles.
  • EHR-specific preprocessing: EHR-specific cleaning removed low-information Family history and Vital Signs sections, expanded contractions, and deleted metadata, headers, dates, contact details, measurements, and punctuation.These steps addressed incomplete clinical narratives and removed non-contextual information from embedding training.

V. QUALITATIVE EVALUATION

Qualitative inspection of the five nearest medical terms shows that EHR and MedLit embeddings capture medical semantics more accurately and relevantly than GloVe and Google News. EHR emphasizes clinical language, whereas MedLit reflects biomedical research usage.

  • V. QUALITATIVE EVALUATION: The evaluation ranked vocabulary terms by cosine similarity and manually inspected the five highest-ranked neighbors for selected disorder, symptom, and drug words.The study selected eight target words across the three medical categories and used word embeddings trained from four resources.
  • V. QUALITATIVE EVALUATION: EHR finds clinically relevant synonyms, comorbidities, symptoms, and medications, including mellitus for diabetes, orthopnea for dyspnea, and opiate for opioid.For diabetes, EHR also identifies cholesterolemia, dyslipidemia, and uncontrolled; for dyspnea, it finds palpitations, exertional, and doe.
  • V. QUALITATIVE EVALUATION: MedLit retrieves biomedical-research terms and mechanisms, such as cardiovascular and obesity-related terms for diabetes and mechanism-of-action terms for opioid.For dyspnea, MedLit also finds related symptoms, a synonym, a relevant disorder, and a term associated with rhonchi.
  • V. QUALITATIVE EVALUATION: EHR and MedLit capture medical-term semantics better than GloVe and Google News, producing more relevant similar medical terms across disorders, symptoms, and drugs.The comparison used manually inspected top-five neighbors for selected medical words; examples include diabetes, dyspnea, opioid, and aspirin.
  • V. QUALITATIVE EVALUATION: EHR reflects clinical narratives, including morphological variants and typos, whereas MedLit retrieves terms primarily from biomedical research perspectives.Examples of EHR variants and typos include melitis, caner, and thraot.

VI. QUANTITATIVE EVALUATION · A. Intrinsic Evaluation

The study quantitatively evaluated biomedical word embeddings using intrinsic medical-term similarity benchmarks and extrinsic downstream biomedical NLP tasks. In intrinsic evaluation, EHR-trained embeddings aligned most closely with expert judgments overall, while MedLit was comparatively strong on UMNSRS and general-domain embeddings performed similarly but worse overall.

  • VI. QUANTITATIVE EVALUATION: The quantitative evaluation combined intrinsic similarity measurement with extrinsic testing on downstream biomedical NLP tasks.Intrinsic evaluation measured semantic similarity between medical terms, whereas extrinsic evaluation used four published datasets and downstream tasks.
  • A. Intrinsic Evaluation: The evaluation covered expert scoring schemes ranging from categorical relatedness ratings to continuous similarity judgments.MayoSRS used a four-point scale, while UMNSRS used continuous annotations from medical residents.
  • A. Intrinsic Evaluation: For each medical-term pair, the study computed embedding similarity and compared it with human judgments using Pearson correlation.fastText-derived vectors were used for medical terms missing from the embedding vocabulary.
  • A. Intrinsic Evaluation: Out-of-vocabulary medical terms were represented with fastText-style character trigram vectors aggregated from shared n-grams.Each word with at least three characters was represented as a bag of character trigrams, whose vectors were averaged for OOV terms.
  • A. Intrinsic Evaluation: EHR-trained embeddings matched human experts’ medical-term similarity judgments most closely overall, outperforming other resources significantly in the reported comparisons.MedLit performed worse than EHR overall but was comparatively strong on UMNSRS; GloVe and Google News were inferior to EHR and MedLit and similar to each other.

B. Extrinsic Evaluation

The extrinsic evaluation measured how word embeddings affect biomedical NLP tasks. It covered clinical information extraction and biomedical information retrieval, including two clinical IE tasks.

  • B. Extrinsic Evaluation: The evaluation assessed word embeddings on specific biomedical NLP tasks.Extrinsic evaluation was used to measure their task-level impact.
  • B. Extrinsic Evaluation: Three biomedical NLP tasks were included, namely clinical IE and biomedical IR.
  • B. Extrinsic Evaluation: Two clinical IE tasks were used to evaluate the word embeddings.

1) Clinical Information Extraction:

Word embeddings improved clinical information extraction, with EHR-trained embeddings strongest on both local fracture and shared smoking-status tasks. However, general-domain embeddings could remain competitive when task terminology overlaps their training data.

  • Cross-task comparison: Although EHR embeddings were best locally, Wikipedia embeddings were comparable, while Google News embeddings were close to EHR on smoking extraction and not significantly inferior.Google News was worst on fracture extraction but had comparable F1 and better recall on smoking-status extraction.
  • Shared smoking-status IE task: Embedding features outperformed term-frequency features on the shared smoking-status task, consistent with the local institutional results.The authors attribute this improvement to semantic information captured by word embeddings.
  • Shared smoking-status IE task: EHR-trained embeddings achieved the best shared-task performance, reaching an F1 score of 0.900 and indicating that effective embeddings can transfer across clinical institutions.The task extracted smoking status from discharge records, and EHR embeddings were strongest despite the data coming from another institution.

2) Biomedical Information Retrieval:

On the TREC 2016 CDS biomedical literature retrieval task, word-embedding query expansion failed to improve retrieval and sometimes worsened infAP and MAP. EHR- and MedLit-trained embeddings performed slightly better than GloVe and Google News, but without statistical significance.

  • Biomedical Information Retrieval:: The TREC 2016 CDS track used EHR-derived query topics covering Diagnosis, Test, and Treatment information needs.Ten topics were provided for each category, with note, description, and summary fields.
  • Biomedical Information Retrieval:: Retrieval performance was measured using infNDCG, infAP, P@10, and MAP.These metrics assess ranking quality, retrieval effectiveness, top-10 relevance, and mean average precision.
  • Biomedical Information Retrieval:: Word-embedding query expansion failed to improve retrieval and worsened performance for infAP and MAP on TREC 2016 CDS.The comparison evaluated embeddings trained from four resources for query expansion.
  • Biomedical Information Retrieval:: EHR and MedLit embeddings performed slightly better than GloVe and Google News, without statistical significance (p<0.01).The result indicates no significant improvement from using embeddings trained from different resources for biomedical information retrieval.

3) Relation Extraction:

On DDIExtraction 2013, Google News embeddings achieved the best overall performance, while corpus-specific results varied between DrugBank and MedLine. General-English semantics appear important for identifying drug interactions, although the overall advantage was not statistically conclusive.

  • 3) Relation Extraction:: The task extracted drug–drug interaction pairs from sentence-level annotations in DrugBank and MedLine abstracts within the DDIExtraction 2013 corpus.Sentences could contain two or more drugs, requiring automatic identification of interacting drug pairs.
  • 3) Relation Extraction:: General-English terms such as “interact” were especially informative for classifying drug interactions, favoring embeddings trained on Google News.In “Acarbose may interact with metformin,” the interaction term is crucial but is not medical terminology.
  • 3) Relation Extraction:: MedLit embeddings performed best on DrugBank, whereas Google News embeddings performed best on MedLine.These results show that embeddings trained on the same corpus as evaluation texts are not necessarily superior.

VII. CONCLUSION AND DISCUSSION

The study evaluates biomedical word embeddings trained from clinical, biomedical, Wikipedia, and news corpora using qualitative, intrinsic, and extrinsic tests. Domain-specific embeddings better captured medical semantics, but general-domain embeddings often performed comparably on downstream tasks and remain practical when clinical corpora are inaccessible.

  • Conclusion: EHR and MedLit embeddings captured medical semantics and relevant similar terms better than GloVe and Google News, with similarity judgments closer to human experts.EHR reflected clinical language, whereas MedLit emphasized terminology and perspectives from medical research.
  • Conclusion: There was no consistent global ranking of embeddings across downstream biomedical NLP applications, although adding embeddings as extra features improved results on most tasks.The downstream applications were clinical information extraction, biomedical information retrieval, and relation extraction.
  • Conclusion: Biomedical-domain embeddings did not necessarily outperform general-domain embeddings, so Wikipedia and news embeddings can remain viable when domain-specific corpora are difficult to access.Local institutional embeddings might nevertheless perform better for local institutional NLP tasks.
  • Future work: The study plans broader evaluations across biomedical applications, language characteristics, healthcare institutions, EHR systems, and sublanguage portability.Future applications include medical named entity recognition and clinical note summarization.
  • Limitations: The EHR analysis used data from only Mayo Clinic, potentially biasing conclusions because EHR quality may vary across institutions.The authors are exploring privacy-preserving techniques to obtain embeddings from multiple sites.
  • Limitations: Generalizability is limited because only two public pre-trained embeddings and one shared-task dataset for each biomedical IR and RE task were evaluated.Additional publicly available embeddings could be assessed in future work.

APPENDIX A INTRINSIC EVALUATION OF WORD EMBEDDINGS WITH DIFFERENT DIMENSIONS.

The appendix reports intrinsic evaluations of word embeddings across different dimensions and visualizes embeddings trained on four corpora. Table IX compares embedding similarity scores with human-expert judgments across four datasets.

  • Intrinsic evaluation: Table IX reports Pearson correlations between similarity scores from embeddings with different dimensions and human-expert scores on four datasets.
  • Visualization of word embeddings: The appendix visualizes word embeddings trained on EHR, MedLit, GloVe, and Google News.
Loading 1802.00400v3…