Automatic genre identification: a survey: Automatic genre identification: a survey: T. Kuzman, N. Ljubešić.

Automatic genre identification (AGI) is a text classification task focused on genres, i.e., text categories defined by the author's purpose, common function of the text, and the text's conventional form. Obtaining genre information has been shown to be beneficial for a wide range of disciplines, inc...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 1; pp. 537 - 571
Autores principales: Kuzman, Taja, Ljubešić, Nikola
Formato: Literature Review
Publicado: Springer Nature Mar2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=183750656&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 183750656
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2025
      vid: 59
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        183750656
        10.1007/s10579-023-09695-8
      ppf: 537
      ppct: 34
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: Automatic genre identification: a survey: Automatic genre identification: a survey: T. Kuzman, N. Ljubešić.
      aug:
        au:
          Kuzman, Taja
          Ljubešić, Nikola
        affil:
          https://ror.org/01hdkb925 Department of Knowledge Technologies, Jožef Stefan Institute, Jamova cesta 39, 1000, Ljubljana, Slovenia
          https://ror.org/01hdkb925 Jožef Stefan International Postgraduate School, Jamova cesta 39, 1000, Ljubljana, Slovenia
          https://ror.org/05njb9z20 Faculty of Computer and Information Science, University of Ljubljana, Večna pot 113, 1000, Ljubljana, Slovenia
      su:
        Natural language processing
        Automatic identification
        Information technology security
        Computational linguistics
        Corpora
      sug:
        subj:
          Natural language processing
          Automatic identification
          Information technology security
          Computational linguistics
          Corpora
      keyword:
        Automatic genre identification
        Communication and Culture Linguistics
        Genre datasets
        Genre schemata
        Information and Computing Sciences Artificial Intelligence and Image Processing Language
        Survey paper
        Text genre
        Web genre
      ab: Automatic genre identification (AGI) is a text classification task focused on genres, i.e., text categories defined by the author's purpose, common function of the text, and the text's conventional form. Obtaining genre information has been shown to be beneficial for a wide range of disciplines, including linguistics, corpus linguistics, computational linguistics, natural language processing, information retrieval and information security. Consequently, in the past 20 years, numerous researchers have collected genre datasets with the aim to develop an efficient genre classifier. However, their approaches to the definition of genre schemata, data collection and manual annotation vary substantially, resulting in significantly different datasets. As most AGI experiments are dataset-dependent, a sufficient understanding of the differences between the available genre datasets is of great importance for the researchers venturing into this area. In this paper, we present a detailed overview of different approaches to each of the steps of the AGI task, from the definition of the genre concept and the genre schema, to the dataset collection and annotation methods, and, finally, to machine learning strategies. Special focus is dedicated to the description of the most relevant genre schemata and datasets, and details on the availability of all of the datasets are provided. In addition, the paper presents the recent advances in machine learning approaches to automatic genre identification, and concludes with proposing the directions towards developing a stable multilingual genre classifier.
      pubtype: Academic Journal
      doctype: Literature Review
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N