Thesis Framework

Beyond Usability: A Comprehensive Evaluation Framework for AI-Generated Educational Resources

Moving beyond traditional usability scales to measure Content Quality, Pedagogical Alignment, and Ethics in the era of Generative AI.

Test the Framework

The Context

Generative AI (GenAI) tools such as ChatGPT, Gemini, Claude, and DeepSeek are now capable of producing complex digital educational resources—including lectures, assessments, explanations, and personalized learning content.

"While GenAI can enhance efficiency, inclusivity, and personalized learning, AI-generated content may also introduce misinformation, bias, pedagogical misalignment, accessibility issues, or ethical concerns."

Traditional evaluation metrics such as the Systematic Usability Scale (SUS) or linguistic similarity scores (BLEU) fail to capture deeper educational quality. To address this gap, this thesis proposes a validated, multi-dimensional evaluation framework.

Figure 1: The AIGDER Evaluation Framework Structure

The 5 Pillars

The core dimensions measured by the instrument.

1

Content Quality & Expression

Assesses the relevance, accuracy, clarity, and readability of the output. Does the AI provide factually correct information in a format suitable for learning?

2

Pedagogical Alignment

Measures how well the content supports learning goals, explains concepts (scaffolding), and provides useful feedback or assessment support.

3

Inclusivity & Accessibility

Evaluates language simplicity, cultural neutrality (bias), device flexibility, and adherence to universal design principles.

4

Transparency & Trust

Focuses on academic integrity, the disclosure of AI generation, and the user's ability to trust the tool ethically.

5

Support & Feedback

Captures the usability experience, efficiency, speed, and overall satisfaction with the tool as a support mechanism.