How different is different? Systematically identifying distribution shifts and their impacts in NER datasets.

When processing natural language, we are frequently confronted with the problem of distribution shift. For example, using a model trained on a news corpus to subsequently process legal text exhibits reduced performance. While this problem is well-known, to this point, there has not been a systematic...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 2; pp. 1111 - 1151
Autores principales: Li, Xue, Groth, Paul
Formato: Artículo
Publicado: Springer Nature Jun2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240049&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 185240049
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 59
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        185240049
        10.1007/s10579-024-09754-8
      ppf: 1111
      ppct: 40
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2MB
      tig:
        atl: How different is different? Systematically identifying distribution shifts and their impacts in NER datasets.
      aug:
        au:
          Li, Xue
          Groth, Paul
        affil: https://ror.org/04dkp9463 Informatics Institute, University of Amsterdam, Science Park, 1098 XH, Amsterdam, Netherlands
      su:
        Task performance
        Corpora
      sug:
        subj:
          Task performance
          Corpora
      keyword:
        Distribution shift
        Named entity recognition
      ab: When processing natural language, we are frequently confronted with the problem of distribution shift. For example, using a model trained on a news corpus to subsequently process legal text exhibits reduced performance. While this problem is well-known, to this point, there has not been a systematic study of detecting shifts and investigating the impact shifts have on model performance for NLP tasks. Therefore, in this paper, we detect and measure two types of distribution shift, across three different representations, for 12 benchmark Named Entity Recognition datasets. We show that both input shift and label shift can lead to dramatic performance degradation. For example, fine-tuning on a wide spectrum dataset (OntoNotes) and testing on an email dataset (CEREC) that shares labels leads to a 63-points drop in F1 performance. Overall, our results indicate that the measurement of distribution shift can provide guidance to the amount of data needed for fine-tuning and whether or not a model can be used "off-the-shelf" without subsequent fine-tuning. Finally, our results show that shift measurement can play an important role in NLP model pipeline definition.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N