Japanese readability assessment using machine learning

Reconocimiento - No comercial - Compartir igual (by-nc-sa)Reconocimiento - No comercial - Compartir igual (by-nc-sa)Reconocimiento - No comercial - Compartir igual (by-nc-sa)

Directores

Editores

Otras autorías

Unidades organizativas

Compartir

Handle

https://riunet.upv.es/handle/10251/206507

Cita bibliográfica

Ivie, T.; Reynolds, R. (2024). Japanese readability assessment using machine learning. En Editorial Universitat Politècnica de València, EuroCALL 2023. CALL for all Languages - Short Papers (pp. 133-138). https://doi.org/10.4995/EuroCALL2023.2023.16989

Titulación

Resumen

[EN] We present a new corpus of Japanese texts, labeled according to six second-language readability levels. We also show the results of experiments training machine-learning classifiers to automatically label new texts according to reading level. The resulting models can be used in language-learning websites and applications to enhance Japanese language learning. The best-performing model, Random Forest, achieved an F1 score of 0.86, with an adjacent accuracy of 0.97. Of the 114 features used, we identify a small subset of five features that are sufficient to achieve an F1 score of 0.74. The corpus, code, and resulting models are free and open-source.¹ ¹ https://github.com/reynoldsnlp/japanese_readability_corpus

Fuente

EuroCALL 2023. CALL for all Languages - Short Papers isbn: 9788413961316

Editorial

Editorial Universitat Politècnica de València

Enlaces relacionados

URL