Identification of discursive topics based on textual similarity

Authors

Keywords:

spoken language, computational linguistics, linguistic research, speech

Abstract

A major outcome of linguistic research is the characterization of a community based on observations that, supported by different specialized areas, such as field linguistics, sociolinguistics, ethnolinguistics, and even corpus linguistics, summarize the most representative features of such communities. Most of these features are based on descriptive statistics, in which frequency plays a relevant role: The higher the frequency of a feature, the more representative it will be. In this regard, it is undeniable that any statistical representation represents a global overview of the lexical generalities of a community; however, frequency is not an absolute in terms of its linguistic particularities. In this respect, this study addresses the identification of discursive topics from the recognition of semantic and discursive patterns, which are not necessarily related to a higher statistical frequency. To this end, a vectorial representation process is described. This is based on the application of textual similarity metrics to automatically recognize the underlying topics regarding the interviews stored in the Corpus del Habla de Baja California. The results show a set of topics to characterize, both explicitly and, to a lesser extent, implicitly, the discursive interests of the speakers of the Baja Californian cities of Mexicali, Tijuana, and Ensenada.

Downloads

Download data is not yet available.

References

Adelstein, A. y Boschiroli, V. de los Á. (2021). Semantic Aspects of National Varieties of Spanish in a Dictionary of Neologisms, the Antenario. International Journal of Lexicography, 34(3), 336–357. https://doi.org/10.1093/ijl/ecab010

Agirre, E., Banea, C., Cardie, C., Cer, D., Diab, M., Gonzalez-Agirre, A., Perinan-Pascual, F., Mihalcea, R., y Wiebe, J. (2014). SemEval-2014 Task 10: Multilingual semantic textual similarity. En Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pp. 81–91. Association for Computational Linguistics. https://doi.org/10.3115/v1/S14-2010.

Atar, C. y Erdem, C. (2019). The advantages and disadvantages of corpus linguistics and conversation analysis in second language studies. Proceedings of IX Scientific and Practical Internet Conference of Young Scientists and Students, November. Ukraine.

Bakarov, A. (2018). A Survey of Word Embeddings Evaluation Methods. CoRR, abs/1801.09536. http://arxiv.org/abs/1801.09536.

Bird, S., Klein, E. y Loper, E. (2009). Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. O’Reilly.

Bond C. (2025). Corpus Linguistics as a Research Method in Nursing: A Practical Approach to Analysing Language Data. Journal of advanced nursing, 81(10), 6960–6967. https://doi.org/10.1111/jan.16659.

Boersma, P. (2014). The Use of Praat in Corpus Research. En U. G. and G. K. Jacques Durand (Ed.), The Oxford Handbook of Corpus Phonology. Oxford University Press. https://doi.org/10.1093/OXFORDHB/9780199571932.013.016.

Campillos-Llanos, L., Terroba, A., Zakhir, S., Valverde-Mateos, A. y Capllonch-Carrión, A. (2022). Building a comparable corpus and a benchmark for Spanish medical text simplification. Procesamiento del Lenguaje Natural, (69), 189–196.

Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., y Specia, L. (2017). SemEval-2017 Task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. En Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval 2017), pp. 1–14. Association for Computational Linguistics. https://doi.org/10.18653/v1/S17-2001.

Clark, A., Fox, C. y Lappin, S. (2010). The Handbook of Computational Linguistics and Natural Language Processing. En Blackwell Handbooks in Linguistics. John Wiley & Sons.

Duffé Montalván, A. L. (2016). Estudios sobre el léxico. Peter Lang Verlag. https://www.peterlang.com/document/1053487.

Hausser, R. (1998). Foundations of Computational Linguistics. Springer-Verlag.

Henríquez Ureña, P. (1977). Observaciones sobre el español de América. En J.C. Ghiano (Comp.) Observaciones sobre el español de América y otros estudios filológicos. Academia Argentina de Letras.

Jurafsky, D. y Martin, J. (2007). Speech and Language Processing: An introduction to natural language processing, computational linguistics, and speech recognition. Prentice Hall.

Kulkarni, A. y Pedersen, T. (2005). SenseClusters: Unsupervised Clustering and Labeling of Similar Contexts. Proceedings of the ACL Interactive Poster and Demonstration Sessions, 105–108.

Lee, L. (1997). Similarity-Based Approaches to Natural Language Processing. arXiv.

Lope Blanch, J. M. (1970). Las zonas dialectales de México. Nueva Revista De Filología Hispánica (NRFH), 19(1), 1-11. https://doi.org/10.24201/nrfh.v19i1.436.

López, H. (2015). Proyecto Panhispánico de disponibilidad léxica. http://www.dispolex.com.

Manning, C., y Schütze, H. (1999). Foundations of statistical natural language processing. MIT Press.

Mautner, G. (2019). A research note on corpora and discourse: Points to ponder in research design. Journal of Corpora and Discourse Studies. 2(2). https://doi.org/10.18573/jcads.32

Mendoza, E. (2004). Notas sobre el español del noroeste. Culiacán: El Colegio de Sinaloa.

Mendoza, E. (2006). El español del noroeste mexicano. En Cestero, A., Molina, I., Paredes, F. (Eds.) Estudios sociolingüísticos del español de España y América. Madrid: Arco Libros.

Mikolov, T., Chen, K., Corrado, G. S. y Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. ICLR.

Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S. y Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 3111--3119.

Miller, G. (1995). WordNet: A Lexical Database for English. Communications of the ACM, 38(11), 39–41.

Molina, C., y Sierra, G. (2015). Hacia una normalización de la frecuencia de los corpus CREA y CORDE. Revista Signos. Estudios de Lingüística, 48(89), 307-331. http://dx.doi.org/10.4067/S0718-09342015000300002

Moreno Fernández, F. (2021). Metodología del Proyecto para el estudio sociolingüístico del español de España y de América (PRESEEA). Universidad de Alcalá.

Pedersen, T., Patwardhan, S. y Michelizzi, J. (2004). WordNet::Similarity - Measuring the Relatedness of Concepts. Proceeding of the 9th National Conference on Artificial Intelligence (AAAI-04), 1024–1025.

Reinert, M. (1986). Un logiciel d'analyse lexicale. Cahiers de l'analyse des données, Tome 11 no. 4, pp. 471-481.

Robbins, Arnold. (2001). Effective awk programming (3°). O’Reilly.

Salcedo P., Zambrano C. y Rojas D. (2017). Metodología de Análisis de Disponibilidad Léxica en Alumnos de Pedagogía a través de la Comparación Jerárquica de Lexicones. Formación Universitaria.

Saldívar, R., Rábago, A. y Lozano, E. (2017). Corpus para el análisis de la variación lingüística: el diseño del corpus del habla de Baja California. En Investigación y praxis contemporáneas en torno a las lenguas modernas. Perales, M; Hernández, E. (Coord.). México: UQroo. pp. 250-264.

Thornbury, S. (2012). What can a corpus tell us about discourse? En A. O’Keeffe & M. McCarthy (Eds.), The Routledge handbook of corpus linguistics pp. 270–287). Routledge

Toutanova, K., Klein, D., Manning, C. D. y Singer, Y. (2003). Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network. Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1, 173–180. https://doi.org/10.3115/1073445.1073478.

Venegas, R. (2006). La similitud léxico-semántica en artículos de investigación científica en español: Una aproximación desde el Análisis Semántico Latente. Revista Signos, 39(60), 75–106. https://doi.org/10.4067/S0718-09342006000100004.

Published

2026-08-19

How to Cite

Reyes Pérez, A. (2026). Identification of discursive topics based on textual similarity. I+D Revista De Investigaciones, 21(2), 83–94. Retrieved from https://sievi.udi.edu.co/ojs/index.php/ID/article/view/537

Issue

Section

Artículos científicos