SYNTOLOGY HomeExplorerAtlasCodeMethodologyAboutDevelopersFeedPricing
Paper · 2503.05888 · ACL · 2025

QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation

Meng Jiang, - Madison, Mengxia Yu, Bang Nguyen, Tingting Du, Lawrence Angrave

arXiv · PDF · Open in the Atlas

Code that ran

We lifted 4 functions out of this paper's own repositories and ran 1 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.

RepositoryRoleRan
copy not recorded — 1 of 1
bnguyen5/qg-sms — 0 of 3
FunctionStatusWhere it lives
prompt_to_chatml Ran this paper's copy was not recorded; identical code first harvested from princeton-nlp/llmbar
pointer only · get_code("70d139854ce1775e")
Complete Not yet run bnguyen5/qg-sms/qg-sms/evaluate.py
pointer only (licence: NONE) · get_code("50e8b5f3decb3ef3")
openai_completion Not yet run bnguyen5/qg-sms/qg-sms/evaluate.py
pointer only (licence: NONE) · get_code("03954fe7d0e72ef2")
simulate Not yet run bnguyen5/qg-sms/qg-sms/evaluate.py
pointer only (licence: NONE) · get_code("df5e032a7f9614a1")

Repositories linked to this paper

Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.

Abstract

While the Question Generation (QG) task has been increasingly adopted in educational assessments, its evaluation remains limited by approaches that lack a clear connection to the educational values of test items. In this work, we introduce test item analysis, a method frequently used by educators to assess test question quality, into QG evaluation. Specifically, we construct pairs of candidate questions that differ in quality across dimensions such as topic coverage, item difficulty, item discrimination, and distractor efficiency. We then examine whether existing QG evaluation approaches can effectively distinguish these differences. Our findings reveal significant shortcomings in these approaches with respect to accurately assessing test item quality in relation to student performance. To address this gap, we propose a novel QG evaluation framework, QG-SMS, which leverages Large Language Model for Student Modeling and Simulation to perform test item analysis. As demonstrated in our extensive experiments and human evaluation study, the additional perspectives introduced by the simulated student profiles lead to a more effective and robust assessment of test items.

For agents

The same record, over MCP at https://syntology.ai/mcp:

get_harvested_code_for_paper("2503.05888")
get_code_for_paper("2503.05888")
have("2503.05888")

Connect an agent — have() is free.