Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain.
The increasing use of artificial intelligence methods in the legal field has sparked interest in applying Natural Language Processing techniques to handle legal tasks and reduce the workload of these professionals. However, the availability of legal corpora in Portuguese, especially for the Brazilia...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 2; pp. 1685 - 1705 |
|---|---|
| Autores principales: | , , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240056&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 185240056 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2025 vid: 59 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 185240056 10.1007/s10579-024-09762-8 ppf: 1685 ppct: 20 formats: fmt: – @attributes: type: T – @attributes: type: P size: 855KB tig: atl: Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain. aug: au: Siqueira, Felipe A. Vitório, Douglas Souza, Ellen Santos, José A. P. Albuquerque, Hidelberg O. Dias, Márcio S. Silva, Nádia F. F. de Carvalho, André C. P. L. F. Oliveira, Adriano L. I. Bastos-Filho, Carmelo affil: https://ror.org/036rp1748 Institute of Mathematical Sciences and Computation, University of São Paulo, São Carlos, São Paulo, Brazil https://ror.org/047908t24 Federal University of Pernambuco, Recife, Pernambuco, Brazil https://ror.org/047908t24 Rural Federal University of Pernambuco, Serra Talhada, Pernambuco, Brazil https://ror.org/00gtcbp88 University of Pernambuco, Recife, Pernambuco, Brazil https://ror.org/024pz1v04 Federal University of Catalão, Catalão, Goiás, Brazil https://ror.org/0039d5757 Federal University of Goiás, Goiânia, Goiás, Brazil su: Natural language processing Artificial intelligence Portuguese language Corpora Legal language Brazil sug: subj: Brazil Natural language processing Artificial intelligence Portuguese language Corpora Legal language keyword: Communication and Culture Linguistics Corpus Governmental domain Information and Computing Sciences Artificial Intelligence and Image Processing Law and Legal Studies Law Language Legal domain ab: The increasing use of artificial intelligence methods in the legal field has sparked interest in applying Natural Language Processing techniques to handle legal tasks and reduce the workload of these professionals. However, the availability of legal corpora in Portuguese, especially for the Brazilian legal domain, is limited. Existing resources offer some legal data but lack comprehensive coverage. To address this gap, we present Ulysses Tesemõ, a large corpus specifically built for the Brazilian legal domain. The corpus consists of over 3.5 million files, totaling 30.7 GiB of raw text, collected from 159 sources encompassing judicial, legislative, academic, news, and other related data. The data was collected by scraping public information from governmental websites, emphasizing contents generated over the past two decades. We categorized the obtained files into 30 distinct categories, covering various branches of the Brazilian government and different types of texts. The corpus retains the original content with minimal data transformations, addressing the scarcity of Portuguese legal corpora and providing researchers with a valuable resource for advancing in the research area. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|