Normalized dataset for Sanskrit word segmentation and morphological parsing.

Sanskrit processing has seen a surge in the use of data-driven approaches over the past decade. Various tasks such as segmentation, morphological parsing, and dependency analysis have been tackled through the development of state-of-the-art models despite working with relatively limited datasets com...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 2; pp. 1279 - 1331
Autores principales: Krishnan, Sriram, Kulkarni, Amba, Huet, Gérard
Formato: Artículo
Publicado: Springer Nature Jun2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240029&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 185240029
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 59
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        185240029
        10.1007/s10579-024-09724-0
      ppf: 1279
      ppct: 52
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: Normalized dataset for Sanskrit word segmentation and morphological parsing.
      aug:
        au:
          Krishnan, Sriram
          Kulkarni, Amba
          Huet, Gérard
        affil:
          https://ror.org/04a7rxb17 Department of Sanskrit Studies, University of Hyderabad, Hyderabad, Telangana, India
          https://ror.org/02kvxyf05 INRIA Paris Center, Paris, France
      su:
        Computational linguistics
        Corpora
        Vocabulary
        Language & languages
      sug:
        subj:
          Computational linguistics
          Corpora
          Vocabulary
          Language & languages
      keyword:
        Datasets
        Morphological parsing
        Sanskrit computational linguistics
        Word segmentation
      ab: Sanskrit processing has seen a surge in the use of data-driven approaches over the past decade. Various tasks such as segmentation, morphological parsing, and dependency analysis have been tackled through the development of state-of-the-art models despite working with relatively limited datasets compared to other languages. However, a significant challenge lies in the availability of annotated datasets that are lexically, morphologically, syntactically, and semantically tagged. While syntactic and semantic tags are preferable for later stages of processing such as sentential parsing and disambiguation, lexical and morphological tags are crucial for low-level tasks of word segmentation and morphological parsing. The Digital Corpus of Sanskrit (DCS) is one notable effort that hosts over 650,000 lexically and morphologically tagged sentences from around 250 texts but also comes with its limitations at different levels of a sentence like chunk, segment, stem and morphological analysis. To overcome these limitations, we look at alternatives such as Sanskrit Heritage Segmenter (SH) and Saṃsādhanī tools, that provide information complementing DCS' data. This work focuses on enriching the DCS dataset by incorporating analyses from SH, thereby creating a dataset that is rich in lexical and morphological information. Furthermore, this work also discusses the impact of such datasets on the performances of existing segmenters, specifically the Sanskrit Heritage Segmenter.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N