Effectiveness of commercial text embedding models for multilingual, multi-class SaaS software classification: A practical study
Fecha
Autores
Directores
Editores
Otras autorías
Unidades organizativas
Handle
Cita bibliográfica
Titulación
Resumen
[EN] In the rapidly evolving field of Software as a Service (SaaS), the accurate categorization of multilingual SaaS applications represents a significant challenge due to the inherent linguistic diversity and continuous growth in available software categories. This study investigates the application of commercial text embedding models, which transform textual data into numerical representations, for multilingual, large-scale, multi-class software classification tasks. We systematically compare various text embedding models integrated with classification algorithms, examining their predictive performance and transfer learning capabilities across multiple languages. Our experiments demonstrate that these embedding models exhibit substantial robustness and efficacy in both monolingual and cross-lingual classification contexts. Notably, a multi-layer perceptron classifier trained on bilingual datasets (French and English) using OpenAI s text-embedding-3-large embedding model achieved high accuracy (0.90) and F1-score (0.78), even when evaluated on languages not represented in the training corpus. This research not only offers valuable insights for professionals and practitioners in the SaaS sector but also lays the groundwork for further research in advanced applications, crucial for handling the extensive textual data in the contemporary digital marketplace.
