Accurate and efficient general-purpose boilerplate detection for crawled web corpora.

Removal of boilerplate is one of the essential tasks in web corpus construction and web indexing. Boilerplate (redundant and automatically inserted material like menus, copyright notices, navigational elements, etc.) is usually considered to be linguistically unattractive for inclusion in a web corp...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 51; no. 3; pp. 873 - 890
Autor principal: Schäfer, Roland
Formato: Artículo
Publicado: Springer Nature Sep2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=124484779&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 124484779
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2017
      vid: 51
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        124484779
        10.1007/s10579-016-9359-2
      ppf: 873
      ppct: 17
      formats:
        fmt:
          @attributes:
            type: P
            size: 658KB
      tig:
        atl: Accurate and efficient general-purpose boilerplate detection for crawled web corpora.
      aug:
        au: Schäfer, Roland
        affil: Deutsche und niederländische Philologie , Freie Universität Berlin , Habelschwerdter Allee 45 14195 Berlin Germany
      su:
        Corpora
        Language & languages
        Computer software
        Websites
        Alphabets
      sug:
        subj:
          Corpora
          Language & languages
          Computer software
          Websites
          Alphabets
      keyword:
        Boilerplate
        Corpus construction
        Non-destructive corpus normalization
        Web corpora
      ab: Removal of boilerplate is one of the essential tasks in web corpus construction and web indexing. Boilerplate (redundant and automatically inserted material like menus, copyright notices, navigational elements, etc.) is usually considered to be linguistically unattractive for inclusion in a web corpus. Also, search engines should not index such material because it can lead to spurious results for search terms if these terms appear in boilerplate regions of the web page. In this paper, I present and evaluate a supervised machine-learning approach to general-purpose boilerplate detection for languages based on Latin alphabets using Multi-Layer Perceptrons (MLPs). It is both very efficient and very accurate (between 95 % and $$99\,\%$$ correct classifications, depending on the input language). I show that language-specific classifiers greatly improve the accuracy of boilerplate detectors. The single features used for the classification are evaluated with regard to the merit they contribute to the classification. Furthermore, I show that the accuracy of the MLP is on a par with that of a wide range of other classifiers. My approach has been implemented in the open-source texrex web page cleaning software, and large corpora constructed using it are available from the COW initiative, including the CommonCOW corpora created from CommonCrawl datasets.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N