UDAT: Compound quantitative analysis of text using machine learning.

Computing machines allow quantitative analysis of large databases of text, providing knowledge that is difficult to obtain without using automation. This article describes Universal Data Analysis of Text (UDAT) —a text analysis method that extracts a large set of numerical text content descriptors f...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 36; no. 1; pp. 187 - 209
Autor principal: Shamir, Lior
Formato: Artículo
Publicado: Oxford University Press / USA Apr2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=150091615&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 150091615
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2021
      vid: 36
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        150091615
        10.1093/llc/fqaa007
      ppf: 187
      ppct: 22
      formats:
        fmt:
          @attributes:
            type: P
            size: 664KB
      tig:
        atl: UDAT: Compound quantitative analysis of text using machine learning.
      aug:
        au: Shamir, Lior
        affil: Kansas State University , USA
      su:
        Microsoft Windows (Operating system)
        Machine learning
        Quantitative research
        Text files
        Pattern recognition systems
        Word frequency
        Linux operating systems
      sug:
        subj:
          Microsoft Windows (Operating system)
          Machine learning
          Quantitative research
          Text files
          Pattern recognition systems
          Word frequency
          Linux operating systems
      ab: Computing machines allow quantitative analysis of large databases of text, providing knowledge that is difficult to obtain without using automation. This article describes Universal Data Analysis of Text (UDAT) —a text analysis method that extracts a large set of numerical text content descriptors from text files and performs various pattern recognition tasks such as classification, similarity between classes, correlation between text and numerical values, and query by example. Unlike several previously proposed methods, UDAT is not based on frequency of words and links between certain key words and topics. The method is implemented as an open-source software tool that can provide detailed reports about the quantitative analysis of sets of text files, as well as exporting the numerical text content descriptors in the form of comma-separated values files to allow statistical or pattern recognition analysis with external tools. It also allows the identification of specific text descriptors that differentiate between classes or correlate with numerical values and can be applied to problems related to knowledge discovery in domains such as literature and social media. UDAT is implemented as a command-line tool that runs in Windows, and the open source is available and can be compiled in Linux systems. UDAT can be downloaded from http://people.cs.ksu.edu/∼lshamir/downloads/udat.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N