Normalized dataset for Sanskrit word segmentation and morphological parsing.
Sanskrit processing has seen a surge in the use of data-driven approaches over the past decade. Various tasks such as segmentation, morphological parsing, and dependency analysis have been tackled through the development of state-of-the-art models despite working with relatively limited datasets com...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 2; pp. 1279 - 1331 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240029&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 185240029 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2025 vid: 59 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 185240029 10.1007/s10579-024-09724-0 ppf: 1279 ppct: 52 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.7MB tig: atl: Normalized dataset for Sanskrit word segmentation and morphological parsing. aug: au: Krishnan, Sriram Kulkarni, Amba Huet, Gérard affil: https://ror.org/04a7rxb17 Department of Sanskrit Studies, University of Hyderabad, Hyderabad, Telangana, India https://ror.org/02kvxyf05 INRIA Paris Center, Paris, France su: Computational linguistics Corpora Vocabulary Language & languages sug: subj: Computational linguistics Corpora Vocabulary Language & languages keyword: Datasets Morphological parsing Sanskrit computational linguistics Word segmentation ab: Sanskrit processing has seen a surge in the use of data-driven approaches over the past decade. Various tasks such as segmentation, morphological parsing, and dependency analysis have been tackled through the development of state-of-the-art models despite working with relatively limited datasets compared to other languages. However, a significant challenge lies in the availability of annotated datasets that are lexically, morphologically, syntactically, and semantically tagged. While syntactic and semantic tags are preferable for later stages of processing such as sentential parsing and disambiguation, lexical and morphological tags are crucial for low-level tasks of word segmentation and morphological parsing. The Digital Corpus of Sanskrit (DCS) is one notable effort that hosts over 650,000 lexically and morphologically tagged sentences from around 250 texts but also comes with its limitations at different levels of a sentence like chunk, segment, stem and morphological analysis. To overcome these limitations, we look at alternatives such as Sanskrit Heritage Segmenter (SH) and Saṃsādhanī tools, that provide information complementing DCS' data. This work focuses on enriching the DCS dataset by incorporating analyses from SH, thereby creating a dataset that is rich in lexical and morphological information. Furthermore, this work also discusses the impact of such datasets on the performances of existing segmenters, specifically the Sanskrit Heritage Segmenter. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|