ТЕКСТОМЕТРИЈА У КОРПУСНОЈ ЛИНВИСТИЦИ: СРПСКИ ФУДБАЛСКИ КОРПУС срФудКо

The corpus srFudKo is the first football corpus of the journalistic genre in Serbian. It represents a partition of the larger FudKo, which is formed by two sections: the Spanish Language corpus esFudKo and the Serbian Language corpus srFudKo, compiled and processed through the unique methodology of...

Full description

Bibliographic Details
Published in:Nasleđe no. 59; pp. 111 - 125
Main Author: Лазаревић, Јелена В.
Format: Article
Published: University of Kragujevac, Faculty of Philology & Arts 2024
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=182598335&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 182598335
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        18201768
        LB0E
      jtl: Nasleđe
      issn: 18201768
      maglogo: N
    pubinfo:
      dt: 2024
      iid: 59
      pid: 45667
      pub: University of Kragujevac, Faculty of Philology & Arts
    artinfo:
      ui:
        182598335
        10.46793/NasKg2459.111L
      ppf: 111
      ppct: 14
      formats:
      tig:
        atl: ТЕКСТОМЕТРИЈА У КОРПУСНОЈ ЛИНВИСТИЦИ: СРПСКИ ФУДБАЛСКИ КОРПУС срФудКо
      aug:
        au: Лазаревић, Јелена В.
        affil: Универзитет у Београду Филолошки факултет Докторске академске студије
      su:
        Corpora
        Word frequency
        Serbian language
        Collocation (Linguistics)
        Spanish language
      sug:
        subj:
          Corpora
          Word frequency
          Serbian language
          Collocation (Linguistics)
          Spanish language
      keyword:
        Corpus Linguistics
        Serbian Football Corpus srFudKo
        textometry
        the language of football
        корпусна лингвистика
        српски фудбалски корпус срФудКо
        текстометрија
        језик фудбала
      ab:
        The corpus srFudKo is the first football corpus of the journalistic genre in Serbian. It represents a partition of the larger FudKo, which is formed by two sections: the Spanish Language corpus esFudKo and the Serbian Language corpus srFudKo, compiled and processed through the unique methodology of textometry in Corpus Linguistics. The aim of the paper is showcasing the formation of srFudKo, analyzing the specificities of the language of football utilized in media articles through statistic and textometric analysis, through the insight into the use of the given tools. The corpus of football articles in srFudKo has been compiled from five Serbian websites: B92, Blic, Mondo, Politika, and Sportklub. It has been prepared as a collection of XML data files, organized by year and website of origin. A total of 11.117 articles were distributed in 37 databases. The domain corpus was processed through tokenization, tagging of the word type and lemmatization. The corpus sr FudKo contains 10.100.553 tokens, out of which 8.618.426 represent words, while 1.482.127 represent interpunction. The average length of an article is 1068 words. The distribution of certain word frequencies in srFudKo testifies to the fact that the Football Corpus contains more proper names and numbers than the Standard Serbian Language Corpus to a significant extent. The specificities of srFudKo are: emoticons of various kinds, as well as different symbols which pertain to the category of corpus word length 1, word length 5 as the most frequent within the corpus, marking a high index of usage specificity of word type in various websites, the uniqueness of concordances and collocations.
        Корпус срФудКо је први фудбалски корпус на српском језику чији је жанр новински. Представља партицију корпуса ФудКо који се састоји из две целине: корпуса шпанског есФудКо и корпуса текстова на српском срФудКо, сачињених и обрађених једин- ственом методологијом текстометрије корпусне лингвистике. Циљ рада је приказивање формирања корпуса срФудКо, испити- вање специфичности језика фудбала који се користи у новинар- ским чланцима путем статистичких и текстометријских анализа кроз увид у коришћење алата за ове анализе. Корпус текстова о фудбалу на српском језику срФудКо прикупљен је са пет српских веб-портала: „Б92", „Блиц", „Мондо", „Политика" и „Спортклуб". Припремљен је као колекција XML датотека, организованих по годинама и по порталима са којих су чланци преузети, тако да је 11.117 чланака распоређено у 37 датотека. Доменски корпус је обрађен поступцима токенизације, тагирања врстом речи и лема- тизацијом. Корпус срФудКо садржи 10.100.553 токена, од чега су 8.618.426 речи а 1.482.127 интерпункцијски знаци. Просечна дужина чланака је 1068 речи. Дистрибуција фреквенција поје- диних врста речи срФудКо сведочи да у корпусу фудбала има знатно више властитих имена и бројева него у корпусу стан- дардног српског језика. Специфичности срФудКо су: емотикони и различите врсте симбола који припадају категорији дужине корпусне речи 1, речи дужине 5 као најфреквентније у корпусу, висок индекс специфичности употребе врста речи на различитим порталима, особеност конкорданци и колокација.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: Serbian
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N