Cruciata, PietroPulizzotto, DavideBeaudry, Catherine2022-11-092022-11-092022-09-209788413960180https://riunet.upv.es/handle/10251/189496[EN] This study offers alternative and promising approaches to word count methods, largely used to develop innovation indicators from unstructured text. We propose a method based on Information Retrieval (IR) and word-embedding models to tackle the semantic ellipsis, one of the main issues of word count methods. We test our IR model by investigating the concept of collaboration and comparing our approach with a baseline corresponding to the keyword search. To ensure the best performances, we use several ways to represent queries and documents in a vector space and three pre-trained word-embedding models. The results prove that our approach can alleviate the semantic ellipsis problem. Indeed, the IR model developed outperforms the classical keyword search in terms of F1-score and Recall. Moreover, we create a combined method that achieves the highest F1-score. These preliminary results can facilitate the creation of reliable innovation indicators from unstructured textual data substituting or complementing survey-based questionnaires.8Reconocimiento - No comercial - Sin obra derivada (by-nc-nd)Text miningNatural Language ProcessingInformation RetrievalInnovation measuresText mining methods for innovation studies: limits and future perspectivesCapĂtulo de libro10.4995/CARMA2022.2022.15076Abierto