Journal of Computer-Assisted Linguistic Research - Vol 09 (2025)

Tabla de contenidos



Artículos

  • Effectiveness of commercial text embedding models for multilingual, multi-class SaaS software classification: A practical study
  • Media manipulation detection: Challenges and perspectives
  • Model for automatic lexical disambiguation based on a hybrid measure


URI permanente para esta colecciónhttps://riunet.upv.es/handle/10251/230420

Examinar

Envíos recientes

Mostrando 1 - 3 de 3
  • Item type: Artículo , Access status: Abierto ,
    Effectiveness of commercial text embedding models for multilingual, multi-class SaaS software classification: A practical study
    (Universitat Politècnica de València, 2025-11-20) Du, Yu; Lavarec, Erwann; Lalouette, Colin
    [EN] In the rapidly evolving field of Software as a Service (SaaS), the accurate categorization of multilingual SaaS applications represents a significant challenge due to the inherent linguistic diversity and continuous growth in available software categories. This study investigates the application of commercial text embedding models, which transform textual data into numerical representations, for multilingual, large-scale, multi-class software classification tasks. We systematically compare various text embedding models integrated with classification algorithms, examining their predictive performance and transfer learning capabilities across multiple languages. Our experiments demonstrate that these embedding models exhibit substantial robustness and efficacy in both monolingual and cross-lingual classification contexts. Notably, a multi-layer perceptron classifier trained on bilingual datasets (French and English) using OpenAI s text-embedding-3-large embedding model achieved high accuracy (0.90) and F1-score (0.78), even when evaluated on languages not represented in the training corpus. This research not only offers valuable insights for professionals and practitioners in the SaaS sector but also lays the groundwork for further research in advanced applications, crucial for handling the extensive textual data in the contemporary digital marketplace.
  • Item type: Artículo , Access status: Abierto ,
    Media manipulation detection: Challenges and perspectives
    (Universitat Politècnica de València, 2025-11-20) Barkhatova, Elvira
    [EN] The article deals with the issue of mass media manipulation and presents a prototype of an innovative browser extension specifically designed to identify biases and manipulative techniques across various dimensions, including lexical, syntactic, and pragmatic levels. This tool is designed not only to detect subtle forms of manipulation embedded within media narratives but also to empower users with an insightful understanding of these tactics. The primary objective of this program is to inform users regarding the potential threats linked with consuming media content, such as news articles or analytical political materials. By meticulously scrutinizing the language employed in these texts, the extension aims to uncover the underlying agendas that may distort public perception. The strategy involves the incorporation of an optimal combination of features, such as detection of emotionally charged vocabulary, biased framing, unsourced claims, ambiguities, and more. The article analyzes both the advantages and the drawbacks of the current approach and provides suggestions for further improvement of the future program.
  • Item type: Artículo , Access status: Abierto ,
    Model for automatic lexical disambiguation based on a hybrid measure
    (Universitat Politècnica de València, 2025-11-20) Núñez Torres, Fredy; Agencia Estatal de Investigación; European Regional Development Fund; European Commission
    [EN] This research presents the development of a more robust model for measuring semantic similarity than those currently available for solving the problem of word sense disambiguation applied to natural language processing. The model is based both linguistically and statistically on the interaction of two approaches to taxonomic exploration: path-based and information content, through the incorporation of FunGramKB as a sense inventory. In terms of evaluation, the proposed similarity measure consistently generated efficient results from a linguistic perspective in the automatic lexical disambiguation process.