Regionalized models for Spanish language variations based on Twitter
- PMID: 37360263
- PMCID: PMC9979884
- DOI: 10.1007/s10579-023-09640-9
Regionalized models for Spanish language variations based on Twitter
Abstract
Spanish is one of the most spoken languages in the world. Its proliferation comes with variations in written and spoken communication among different regions. Understanding language variations can help improve model performances on regional tasks, such as those involving figurative language and local context information. This manuscript presents and describes a set of regionalized resources for the Spanish language built on 4-year Twitter public messages geotagged in 26 Spanish-speaking countries. We introduce word embeddings based on FastText, language models based on BERT, and per-region sample corpora. We also provide a broad comparison among regions covering lexical and semantical similarities and examples of using regional resources on message classification tasks.
Keywords: Linguistic resources; Semantic space; Spanish Twitter.
© The Author(s), under exclusive licence to Springer Nature B.V. 2023, Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
Figures
References
-
- Alshutayri A, Atwell E. Exploring Twitter as a source of an Arabic dialect corpus. International Journal Of Computational Linguistics (IJCL) 2017;8:37–44.
-
- Anowar F, Sadaoui S, Selim B. Conceptual and empirical comparison of dimensionality reduction algorithms (PCA, KPCA, LDA, MDS, SVD, LLE, ISOMAP, LE, ICA, t-SNE) Computer Science Review. 2021;40:100378. doi: 10.1016/j.cosrev.2021.100378. - DOI
-
- Bojanowski P, Grave E, Joulin A, Mikolov T. Enriching word vectors with subword information. Transactions of The Association For Computational Linguistics. 2017;5:135–146. doi: 10.1162/tacl_a_00051. - DOI
-
- CKennedy, B., Atari, M., Davani, A. M., Yeh, L., Omrani, A., Kim, Y., Coombs, K., Havaldar, S., Portillo-Wightman, G., Gonzalez, E., & Hoover, J. (2022). Introducing the Gab Hate Corpus: defining and applying hate-based rhetoric to social media posts at scale. In Language Resources and Evaluation. Springer.
-
- Cotton E, Sharp J. Spanish in the Americas. Berlin: Georgetown University Press; 1988.
LinkOut - more resources
Full Text Sources