Does the model know when it's right?
Measured on Idapt. Every score below has a receipt: click any verdict for the model's actual answer.
This battery is a LENS over the graded Trap Rate and Trick Riddles items rather than a new question set. Every one of those items asks the model to append a confidence (Answer: X (confidence: NN%)), so every graded answer becomes a (confidence, was-correct) pair.
Binned by confidence decile, those pairs draw a reliability diagram: a well-calibrated model that says "80%" is right about 80% of the time, landing on the diagonal; an overconfident model sits below it. Two summary numbers accompany the diagram — the Brier score (mean squared error of the probability against the 0/1 outcome, lower is better) and the expected calibration error (the average gap between a bin's confidence and its accuracy).
Reported only for models with enough confidence-bearing items to be meaningful.
Contamination: public — the underlying items are the public Trap + Riddle questions; this battery adds no new questions, only a different way to read the same answers.
These items are famous, so any model trained after they circulated may know the answers directly. Read scores as a timeline of what still bites, not a measure of intelligence.