CogBench: Evaluating Cognitive-Level Control in LLM Question Generation
Abstract
Educational LLM systems generate assessment questions on demand and depend on the assumption that a model can be controlled to a specified cognitive level. We test the assumption. CogBench is a benchmark for cognitive-level control in LLM question generation, grounded in the Revised Bloom's Taxonomy. Its central design choice is a two-mode protocol that separates cognition from imitation: in standard mode the model is asked to generate at level Lt; in adversarial mode it must do so while restricted to verbs from a different level Lv, demonstrating the requested level structurally rather than lexically. CogBench spans 8 subjects, 6 Bloom levels, and 120 CC-BY OpenStax passages, with ~17,000 graded generations across 13 models (six frontier closed APIs and seven open-source local models), each scored by a deliberately dual evaluator: a 28-rule constraint checker, and the Cognitive Complexity Score (CCS), a fine-tuned BERT classifier at 82.0% exact accuracy that outperforms a four-model LLM-as-judge panel by 7.8 points. We find that (1) adversarial vocabulary collapses strict adherence by 24-56 pp across the 12 models above floor, with the strongest standard performer suffering the largest collapse; (2) Remember-level adherence under adversarial vocabulary is essentially 0% in all 13 models, suggesting Remember lives almost entirely in its surface verbs; (3) the two evaluators rank models differently, including on the leader, a divergence we argue benchmarks should report rather than hide; and (4) the most adversarially robust models on CCS are not frontier flagships but gpt-4o-mini (46.8%) and llama3.1-8b (45.6%). What current LLMs offer educational systems looks like cognitive control and is largely verb-level imitation that resembles it. We release the dataset (CC-BY 4.0; Croissant 1.0 with Responsible-AI fields), the evaluation harness, and a 13-model leaderboard at https://huggingface.co/spaces/cogbench-anonymous/cogbench-leaderboard.