Offensive keyword extraction based on the attention mechanism of BERT and the eigenvector centrality using a graph representation

dc.contributor.authorPeña-Sarracén, Gretel Liz de laes_ES
dc.contributor.authorRosso, Paoloes_ES
dc.contributor.funderAGENCIA ESTATAL DE INVESTIGACIONes_ES
dc.date.accessioned2022-11-07T19:01:36Z
dc.date.available2022-11-07T19:01:36Z
dc.date.issued2021-08-27es_ES
dc.description.abstract[EN] The proliferation of harmful content on social media affects a large part of the user community. Therefore, several approaches have emerged to control this phenomenon automatically. However, this is still a quite challenging task. In this paper, we explore the offensive language as a particular case of harmful content and focus our study in the analysis of keywords in available datasets composed of offensive tweets. Thus, we aim to identify relevant words in those datasets and analyze how they can affect model learning. For keyword extraction, we propose an unsupervised hybrid approach which combines the multi-head self-attention of BERT and a reasoning on a word graph. The attention mechanism allows to capture relationships among words in a context, while a language model is learned. Then, the relationships are used to generate a graph from what we identify the most relevant words by using the eigenvector centrality. Experiments were performed by means of two mechanisms. On the one hand, we used an information retrieval system to evaluate the impact of the keywords in recovering offensive tweets from a dataset. On the other hand, we evaluated a keyword-based model for offensive language detection. Results highlight some points to consider when training models with available datasets.en_EN
dc.description.accrualMethodSes_ES
dc.description.bibliographicCitationPeña-Sarracén, GLDL.; Rosso, P. (2021). Offensive keyword extraction based on the attention mechanism of BERT and the eigenvector centrality using a graph representation. Personal and Ubiquitous Computing. 1-13. https://doi.org/10.1007/s00779-021-01605-5es_ES
dc.description.referencesAo X, Yu X, Liu D, Tian H (2020) News keywords extraction algorithm based on textrank and classified TF-IDF. In: 2020 international wireless communications and mobile computing (IWCMC). IEEE, pp 1364–1369es_ES
dc.description.referencesBasile V, Bosco C, Fersini E, Debora N, Patti V, Pardo FMR, Rosso P, Sanguinetti M, et al. (2019) Semeval-2019 task 5: multilingual detection of hate speech against immigrants and women in twitter. In: 13th international workshop on semantic evaluation. Association for Computational Linguistics, pp 54–63es_ES
dc.description.referencesBerry MW, Kogan J (2010) Text mining: applications and theory. John Wiley & Sons, New Yorkes_ES
dc.description.referencesBoudin F (2013) A comparison of centrality measures for graph-based keyphrase extraction. In: Proceedings of the sixth international joint conference on natural language processing, pp 834–838es_ES
dc.description.referencesBrin S, Page L (1998) The anatomy of a large-scale hypertextual Web search engine. In: Proceedings of the seventh international conference on World Wide Web, pp 107–117es_ES
dc.description.referencesBüttcher S, Clarke CL, Cormack GV (2016) Information retrieval: implementing and evaluating search engines. Mit Press, Cambridgees_ES
dc.description.referencesCasula C, Aprosio AP, Menini S, Tonelli S (2020) Fbk-dh at semeval-2020 task 12: using multi-channel bert for multilingual offensive language detection. In: Proceedings of the fourteenth workshop on semantic evaluation, pp 1539–1545es_ES
dc.description.referencesChaudhari S, Polatkan G, Ramanath R, Mithal V (2019) An attentive survey of attention models. arXiv:1904.02874es_ES
dc.description.referencesDai W, Yu T, Liu Z, Fung P (2020) Kungfupanda at semeval-2020 task 12: Bert-based multi-task learning for offensive language detection. arXiv:2004.13432es_ES
dc.description.referencesDevlin J, Chang MW, Lee K, Toutanova K (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805es_ES
dc.description.referencesFersini E, Rosso P, Anzovino M (2018) Overview of the task on automatic misogyny identification at IberEval 2018. IberEval@ SEPLN 2150:214–228es_ES
dc.description.referencesFiroozeh N, Nazarenko A, Alizon F, Daille B (2020) Keyword extraction: issues and methods. Nat Lang Eng 26(3):259–291es_ES
dc.description.referencesHasan KS, Ng V (2014) Automatic keyphrase extraction: a survey of the state of the art. In: Proceedings of the 52nd annual meeting of the association for computational linguistics (Volume 1: Long Papers), pp 1262–1273es_ES
dc.description.referencesHu X, Wu B (2006) Automatic keyword extraction using linguistic features. In: Sixth IEEE international conference on data mining-workshops (ICDMW’06). IEEE, pp 19–23es_ES
dc.description.referencesKathait SS, Tiwari S, Varshney A, Sharma A (2017) Unsupervised key-phrase extraction using noun phrases. Int J Comput Appl 162(1)es_ES
dc.description.referencesKaur J, Gupta V (2010) Effective approaches for extraction of keywords. Int J Comput Sci Issues (IJCSI) 7(6):144es_ES
dc.description.referencesKingma DP, Ba J (2014) Adam: A method for stochastic optimization. arXiv:1412.6980es_ES
dc.description.referencesMandl T, Modha S, Majumder P, Patel D, Dave M, Mandlia C, Patel A (2019) Overview of the HASOC track at FIRE 2019: hate speech and offensive content identification in indo-european languages. In: Proceedings of the 11th forum for information retrieval evaluation, pp 14–17es_ES
dc.description.referencesMihalcea R, Tarau P (2004) Textrank: bringing order into text. In: Proceedings of the 2004 conference on empirical methods in natural language processing, pp 404–411es_ES
dc.description.referencesNasar Z, Jaffry SW, Malik MK (2019) Textual keyword extraction and summarization: state-of-the-art. Inf Process Manag 56(6):102088es_ES
dc.description.referencesNewman ME (2008) The mathematics of networks. New Palgrave Encycl Econ 2(2008):1–12es_ES
dc.description.referencesPappagari R, Zelasko P, Villalba J, Carmiel Y, Dehak N (2019) Hierarchical transformers for long document classification. In: 2019 IEEE automatic speech recognition and understanding workshop (ASRU). IEEE, pp 838–844es_ES
dc.description.referencesDe la Pena Sarracén GL, Rosso P (2020) Prhlt-upv at semeval-2020 task 12: Bert for multilingual offensive language detection. In: Proceedings of the fourteenth workshop on semantic evaluation, pp 1605–1614es_ES
dc.description.referencesPitsilis GK, Ramampiaro H, Langseth H (2018) Detecting offensive language in tweets using deep learning. arXiv:1801.04433es_ES
dc.description.referencesPoletto F, Basile V, Sanguinetti M, Bosco C, Patti V (2020) Resources and benchmark corpora for hate speech detection: a systematic review. Lang Resour Eval pp 1–47es_ES
dc.description.referencesRobertson SE, Walker S, Jones S, Hancock-Beaulieu MM, Gatford M, et al. (1995) Okapi at trec-3. Nist Spec Publ 109:109es_ES
dc.description.referencesRosenthal S, Atanasova P, Karadzhov G, Zampieri M, Nakov P (2020) A large-scale semi-supervised dataset for offensive language identification. arXiv:2004.14454es_ES
dc.description.referencesSahrawat D, Mahata D, Kulkarni M, Zhang H, Gosangi R, Stent A, Sharma A, Kumar Y, Shah RR, Zimmermann R (2019) Keyphrase extraction from scholarly articles as sequence labeling using contextualized embeddings. arXiv:1910.08840es_ES
dc.description.referencesUglow H, Zlocha M, Zmyślony S (2019) An exploration of state-of-the-art methods for offensive language detection. arXiv:1903.07445es_ES
dc.description.referencesVashistha N, Zubiaga A (2020) Online multilingual hate speech detection: experimenting with Hindi and English social mediaes_ES
dc.description.referencesVaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. In: Advances in neural information processing systems, pp 5998–6008es_ES
dc.description.referencesWang S, Liu J, Ouyang X, Sun Y (2020) Galileo at semeval-2020 task 12: multi-lingual learning for offensive language identification using pre-trained language models. arXiv:2010.03542es_ES
dc.description.referencesWani AH, Molvi NS, Ashraf SI (2019) Detection of hate and offensive speech in text. In: International conference on intelligent human computer interaction. Springer, pp 87–93es_ES
dc.description.referencesWiedemann G, Yimam SM, Biemann C (2020) Uhh-lt at semeval-2020 task 12: fine-tuning of pre-trained transformer networks for offensive language detection. In: Proceedings of the fourteenth workshop on semantic evaluation, pp 1638–1644es_ES
dc.description.referencesWiegand M, Ruppenhofer J, Kleinbauer T (2019) Detection of abusive language: the problem of biased datasets. In: Proceedings of the 2019 conference of the North American Chapter of the Association for Computational Linguistics: human language technologies, volume 1 (long and short papers), pp 602–608es_ES
dc.description.referencesWitten IH, Paynter GW, Frank E, Gutwin C, Nevill-Manning CG (2005) KEA: practical automated keyphrase extraction. In: Design and usability of digital libraries: case studies in the asia pacific. IGI Global, pp 129–152es_ES
dc.description.referencesZampieri M, Malmasi S, Nakov P, Rosenthal S, Farra N, Kumar R (2019) Predicting the type and target of offensive posts in social media. In: Proceedings of the 2019 conference of the north american chapter of the association for computational linguistics (NAACL), pp 1415–1420es_ES
dc.description.referencesZampieri M, Malmasi S, Nakov P, Rosenthal S, Farra N, Kumar R (2019) Predicting the type and target of offensive posts in social media. arXiv:1902.09666es_ES
dc.description.referencesZampieri M, Malmasi S, Nakov P, Rosenthal S, Farra N, Kumar R (2019) Semeval-2019 task 6: identifying and categorizing offensive language in social media (offenseval). arXiv:1903.08983es_ES
dc.description.referencesZampieri M, Nakov P, Rosenthal S, Atanasova P, Karadzhov G, Mubarak H, Derczynski L, Pitenis Z, Çöltekin Ç (2020) Semeval-2020 task 12: multilingual offensive language identification in social media (offenseval 2020). arXiv:2006.07235es_ES
dc.description.upvformatpfin13es_ES
dc.description.upvformatpinicio1es_ES
dc.identifier.doi10.1007/s00779-021-01605-5es_ES
dc.identifier.issn1617-4909es_ES
dc.identifier.urihttps://riunet.upv.es/handle/10251/189377
dc.languageIngléses_ES
dc.publisherSpringer-Verlages_ES
dc.relation.ispartofPersonal and Ubiquitous Computinges_ES
dc.relation.pasarelaS\450464es_ES
dc.relation.projectIDinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/PGC2018-096212-B-C31/ES/DESINFORMACION Y AGRESIVIDAD EN SOCIAL MEDIA: AGREGANDO INFORMACION Y ANALIZANDO EL LENGUAJE/es_ES
dc.relation.publisherversionhttps://doi.org/10.1007/s00779-021-01605-5es_ES
dc.relation.references10.1109/IWCMC48107.2020.9148491es_ES
dc.relation.references10.18653/v1/S19-2007es_ES
dc.relation.references10.1002/9780470689646es_ES
dc.relation.references10.1016/S0169-7552(98)00110-Xes_ES
dc.relation.references10.18653/v1/2020.semeval-1.201es_ES
dc.relation.references10.18653/v1/2020.semeval-1.272es_ES
dc.relation.references10.1017/S1351324919000457es_ES
dc.relation.references10.3115/v1/P14-1119es_ES
dc.relation.references10.1109/ICDMW.2006.36es_ES
dc.relation.references10.5120/ijca2017913171es_ES
dc.relation.references10.1145/3368567.3368584es_ES
dc.relation.references10.1016/j.ipm.2019.102088es_ES
dc.relation.references10.1109/ASRU46091.2019.9003958es_ES
dc.relation.references10.18653/v1/2020.semeval-1.209es_ES
dc.relation.references10.1007/s10579-020-09502-8es_ES
dc.relation.references10.18653/v1/2021.findings-acl.80es_ES
dc.relation.references10.1007/978-3-030-45442-5_41es_ES
dc.relation.references10.20944/preprints202011.0646.v1es_ES
dc.relation.references10.18653/v1/2020.semeval-1.189es_ES
dc.relation.references10.1007/978-3-030-44689-5_8es_ES
dc.relation.references10.18653/v1/2020.semeval-1.213es_ES
dc.relation.references10.4018/978-1-59140-441-5.ch008es_ES
dc.relation.references10.18653/v1/N19-1144es_ES
dc.relation.references10.18653/v1/S19-2010es_ES
dc.relation.references10.18653/v1/2020.semeval-1.188es_ES
dc.rightsReserva de todos los derechoses_ES
dc.rights.accessRightsAbiertoes_ES
dc.subjectUnsupervised keyword extractiones_ES
dc.subjectOffensive language detectiones_ES
dc.subjectAttention mechanismes_ES
dc.subjectGraph representationes_ES
dc.titleOffensive keyword extraction based on the attention mechanism of BERT and the eigenvector centrality using a graph representationes_ES
dc.typeArtículoes_ES
dc.type.versioninfo:eu-repo/semantics/publishedVersiones_ES
dspace.entity.typePublication
upv.uuid97991f61-ce3c-4022-b530-25524181c9e9es_ES

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
Pena-SarracenRosso - Offensive keyword extraction based on the attention mechanism of BERT and th....pdf
Tamaño:
2 MB
Formato:
Adobe Portable Document Format
Descripción:
Versión editorial