Sentence clustering using continuous vector space representation
Fecha
Directores
Editores
Otras autorías
Unidades organizativas
Handle
https://riunet.upv.es/handle/10251/64386
Cita bibliográfica
Chinea Ríos, M.; Sanchis Trilles, G.; Casacuberta Nolla, F. (2015). Sentence clustering using continuous vector space representation. En Pattern Recognition and Image Analysis: 7th Iberian Conference, IbPRIA 2015, Santiago de Compostela, Spain, June 17-19, 2015, Proceedings. Springer International Publishing. 432-440. doi:10.1007/978-3-319-19390-8 49
Titulación
Resumen
In this paper, we present a clustering approach based on the combined use of a continuous vector space representation of sentences and the k-means algorithm. The principal motivation of this proposal is to split a big heterogeneous corpus into clusters of similar sentences. We use the word2vec toolkit for obtaining the representation of a given word as a continuous vector space. We provide empirical evidence for proving that the use of our technique can lead to better clusters, in terms of intra-cluster perplexity and F 1 score.
Descripción
The final publication is available at Springer via http://dx.doi.org/10.1007/978-3-319-19390-8_49
Palabras clave
Fuente
Pattern Recognition and Image Analysis: 7th Iberian Conference, IbPRIA 2015, Santiago de Compostela, Spain, June 17-19, 2015, Proceedings isbn: 978-3-319-19389-2 issn: 0302-9743
Editorial
Springer International Publishing
