Detecting explicit lyrics: a case study in Italian music.
Preventing the reproduction of songs whose textual content is offensive or inappropriate for kids is an important issue in the music industry. In this paper, we investigate the problem of assessing whether music lyrics contain content unsuitable for children (a.k.a., explicit content). Previous work...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 2; pp. 849 - 868 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=163826586&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 163826586 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2023 vid: 57 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 163826586 10.1007/s10579-022-09595-3 ppf: 849 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 915KB tig: atl: Detecting explicit lyrics: a case study in Italian music. aug: au: Rospocher, Marco affil: Università degli studi di Verona, Lungadige Porta Vittoria, 41, 37129, Verona, Italy su: Italian music Natural language processing Language models Music education Song lyrics sug: subj: Italian music Natural language processing Language models Music education Song lyrics keyword: Convolutional neural networks Explicit content detection Italian language Neural language models Text classification ab: Preventing the reproduction of songs whose textual content is offensive or inappropriate for kids is an important issue in the music industry. In this paper, we investigate the problem of assessing whether music lyrics contain content unsuitable for children (a.k.a., explicit content). Previous works that have computationally tackled this problem have dealt with English or Korean songs, comparing the performance of various machine learning approaches. We investigate the automatic detection of explicit lyrics for Italian songs, complementing previous analyses performed on different languages. We assess the performance of many classifiers, including those–not fully exploited so far for this task–leveraging neural language models, i.e., rich language representations built from textual corpora in an unsupervised way, that can be fine-tuned on various natural language processing tasks, including text classification. For the comparison of the different systems, we exploit a novel dataset we contribute, consisting of approximately 34K songs, annotated with labels indicating explicit content. The evaluation shows that, on this dataset, most of the classifiers built on top of neural language models perform substantially better than non-neural approaches. We also provide further analyses, including: a qualitative assessment of the predictions produced by the classifiers, an assessment of the performance of the best performing classifier in a few-shot learning scenario, and the impact of dataset balancing. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|