A survey of methods to ease the development of highly multilingual text mining applications.

Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in principle, be developed for many languages, most text analysis tools have only...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 46; no. 2; pp. 155 - 177
Autor principal: Steinberger, Ralf
Formato: Artículo
Publicado: Springer Nature Jun2012
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80202962&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 80202962
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2012
      vid: 46
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        80202962
        10.1007/s10579-011-9165-9
      ppf: 155
      ppct: 22
      formats:
        fmt:
          @attributes:
            type: P
            size: 593KB
      tig:
        atl: A survey of methods to ease the development of highly multilingual text mining applications.
      aug:
        au: Steinberger, Ralf
        affil: European Commission, Joint Research Centre (JRC), Via Fermi 2749 21027 Ispra Italy
      su:
        Quotation
        Multilingual communication
        Multilingual computing
        Data mining
        Machine learning
        Surveys
      sug:
        subj:
          Quotation
          Multilingual communication
          Multilingual computing
          Data mining
          Machine learning
          Surveys
      keyword:
        Algorithms
        Cross-lingual projection
        Information extraction
        Media monitoring
        Methods
        Multilinguality
        Quotation recognition
        Rule-based
        Saving effort
        Sentiment analysis
        String similarity calculation
        Summarisation
        Text mining
      ab: Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in principle, be developed for many languages, most text analysis tools have only been applied to small sets of languages because the development effort per language is large. Self-training tools obviously alleviate the problem, but even the effort of providing training data and of manually tuning the results is usually considerable. In this paper, we gather insights by various multilingual system developers on how to minimise the effort of developing natural language processing applications for many languages. We also explain the main guidelines underlying our own effort to develop complex text mining software for tens of languages. While these guidelines-most of all: extreme simplicity-can be very restrictive and limiting, we believe to have shown the feasibility of the approach through the development of the Europe Media Monitor (EMM) family of applications (). EMM is a set of complex media monitoring tools that process and analyse up to 100,000 online news articles per day in between twenty and fifty languages. We will also touch upon the kind of language resources that would make it easier for all to develop highly multilingual text mining applications. We will argue that-to achieve this-the most needed resources would be freely available, simple, parallel and uniform multilingual dictionaries, corpora and software tools.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2012. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2012
    holdings:
      @attributes:
        islocal: N