Ivie, TylerReynolds, Robert2024-07-222024-07-222024-02-129788413961316https://riunet.upv.es/handle/10251/206507[EN] We present a new corpus of Japanese texts, labeled according to six second-language readability levels. We also show the results of experiments training machine-learning classifiers to automatically label new texts according to reading level. The resulting models can be used in language-learning websites and applications to enhance Japanese language learning. The best-performing model, Random Forest, achieved an F1 score of 0.86, with an adjacent accuracy of 0.97. Of the 114 features used, we identify a small subset of five features that are sufficient to achieve an F1 score of 0.74. The corpus, code, and resulting models are free and open-source.¹ ¹ https://github.com/reynoldsnlp/japanese_readability_corpus6Reconocimiento - No comercial - Compartir igual (by-nc-sa)ReadabilityMachine learningJapaneseJapanese readability assessment using machine learningCapítulo de libro10.4995/EuroCALL2023.2023.16989Abierto