Moving beyond traditional usability scales to measure Content Quality, Pedagogical Alignment, and Ethics in the era of Generative AI.
Generative AI (GenAI) tools such as ChatGPT, Gemini, Claude, and DeepSeek are now capable of producing complex digital educational resources—including lectures, assessments, explanations, and personalized learning content.
"While GenAI can enhance efficiency, inclusivity, and personalized learning, AI-generated content may also introduce misinformation, bias, pedagogical misalignment, accessibility issues, or ethical concerns."
Traditional evaluation metrics such as the Systematic Usability Scale (SUS) or linguistic similarity scores (BLEU) fail to capture deeper educational quality. To address this gap, this thesis proposes a validated, multi-dimensional evaluation framework.
Figure 1: The AIGDER Evaluation Framework Structure
The core dimensions measured by the instrument.
Assesses the relevance, accuracy, clarity, and readability of the output. Does the AI provide factually correct information in a format suitable for learning?
Measures how well the content supports learning goals, explains concepts (scaffolding), and provides useful feedback or assessment support.
Evaluates language simplicity, cultural neutrality (bias), device flexibility, and adherence to universal design principles.
Focuses on academic integrity, the disclosure of AI generation, and the user's ability to trust the tool ethically.
Captures the usability experience, efficiency, speed, and overall satisfaction with the tool as a support mechanism.