Constructing specialised corpora through analysing domain representativeness of websites.

The role of the Web for text corpus construction is becoming increasingly significant. However, the contribution of the Web is largely confined to building a general virtual corpus or low quality specialised corpora. In this paper, we introduce a new technique called SPARTAN for constructing special...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 45; no. 2; pp. 209 - 242
Autores principales: Wong, Wilson, Liu, Wei, Bennamoun, Mohammed
Formato: Artículo
Publicado: Springer Nature May2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=60133437&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 60133437
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: May2011
      vid: 45
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        60133437
        10.1007/s10579-011-9141-4
      ppf: 209
      ppct: 33
      formats:
        fmt:
          @attributes:
            type: P
            size: 472KB
      tig:
        atl: Constructing specialised corpora through analysing domain representativeness of websites.
      aug:
        au:
          Wong, Wilson
          Liu, Wei
          Bennamoun, Mohammed
        affil: School of Computer Science and Software Engineering, The University of Western Australia, Crawley 6009 Australia
      su:
        Rankings of websites
        World Wide Web
        Virtual communities
        Search engines
        Corpora
        Computational linguistics
        Electronic information resource searching
      sug:
        subj:
          Rankings of websites
          World Wide Web
          Virtual communities
          Search engines
          Corpora
          Computational linguistics
          Electronic information resource searching
      keyword:
        Boilerplate removal
        Corpus construction
        Specialised corpus
        Term recognition
        Virtual corpus
        Web-derived corpus
        Website ranking
      ab: The role of the Web for text corpus construction is becoming increasingly significant. However, the contribution of the Web is largely confined to building a general virtual corpus or low quality specialised corpora. In this paper, we introduce a new technique called SPARTAN for constructing specialised corpora from the Web by systematically analysing website contents. Our evaluations show that the corpora constructed using our technique are independent of the search engines employed. In particular, SPARTAN-derived corpora outperform all corpora based on existing techniques for the task of term recognition.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N