Comparing methods for the syntactic simplification of sentences in information extraction.

This article describes research aimed at improving the accuracy of an information extraction (IE) system by treating coordinate structures systematically. Commas, coordinating conjunctions, and adjacent comma–conjunction pairs are considered to be potential indicators of coordination in natural lang...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 26; no. 4; pp. 371 - 389
Autor principal: Evans, Richard J.
Formato: Artículo
Publicado: Oxford University Press / USA Dec2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=66887560&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 66887560
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Dec2011
      vid: 26
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        66887560
        10.1093/llc/fqr034
      ppf: 371
      ppct: 18
      formats:
        fmt:
          @attributes:
            type: P
            size: 207KB
      tig:
        atl: Comparing methods for the syntactic simplification of sentences in information extraction.
      aug:
        au: Evans, Richard J.
        affil: University of Wolverhampton, UK
      su:
        Semantics (Philosophy)
        Sentences (Grammar)
        Data mining
        Conjunctions (Grammar)
        Natural language processing
        Coordinate constructions (Linguistics)
      sug:
        subj:
          Semantics (Philosophy)
          Sentences (Grammar)
          Data mining
          Conjunctions (Grammar)
          Natural language processing
          Coordinate constructions (Linguistics)
      ab: This article describes research aimed at improving the accuracy of an information extraction (IE) system by treating coordinate structures systematically. Commas, coordinating conjunctions, and adjacent comma–conjunction pairs are considered to be potential indicators of coordination in natural language. A recursive algorithm is implemented which converts sentences containing classified potential coordinators into sequences of simple sentences. Several approaches to the classification of potential coordinators are presented, one exploiting memory-based learning, another exploiting the publicly available Stanford parser, and a hybrid approach that classifies commas and conjunctions using the former system and comma–conjunction pairs using the latter. The article describes the initial set of features developed for exploitation by the memory-based classifier and presents optimization of that classifier. A baseline system is also described. The sentence simplification module was exploited by an IE system. With regard to the automatic classifiers that form the basis for simplification, comparative evaluation demonstrated that IE can be performed with greatest accuracy when exploiting the hybrid classifier. It also demonstrated that a simple baseline classifier induces improved accuracy when compared to systems that ignore the presence of coordinate structures in input sentences. The article presents an analysis of the errors made by the different sentence simplification modules and the IE system that exploits them. Directions for future research are suggested.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N