Domenech, JosepCarles-Vega, PauMartínez-Varea, Alicia2025-09-292025-09-292025-06-309788413963129https://riunet.upv.es/handle/10251/226633[EN] This paper proposes a systematic framework for integrating large language models (LLMs) into the evaluation of student work. The framework addresses challenges inherent in automated grading, such as ensuring validity, reliability, and minimizing bias, by outlining a structured process that includes prompt design, model selection, evaluation, calibration, and iterative refinement. The approach is designed to be adaptable across diverse educational contexts, supporting both formative and summative assessment needs. This work contributes to the growing literature on AI-driven education, offering practical guidelines and highlighting the need for careful design and continuous validation for high-stakes educational applications.8Reconocimiento - No comercial - Compartir igual (by-nc-sa)Large language modelsArtificial intelligence in educationAI-driven assessmentAutomated gradingAssessment frameworkFormative assessmentPrompt engineeringEducational innovationModelos de lenguaje grandes (LLMsInteligencia artificial en educaciónCalificación automatizadaEvaluación formativaInnovación educativaIngeniería de promptsA Framework for Automated Student Grading Using Large Language ModelsComunicación en congreso10.4995/HEAd25.2025.20152Abierto