The intuitive answer is wrong. Does the model stop and think?
Measured on Idapt. Every score below has a receipt: click any verdict for the model's actual answer.
Each item is a short word problem with a compelling intuitive answer that is wrong (the bat costs $1.00 — no, $1.05; the patch covers half the lake on day 24 — no, day 47). This is the Cognitive Reflection Test construct (Frederick, 2005): it measures the tendency to override a fast wrong answer with a slower correct one, which is exactly the failure mode a fluent language model is prone to.
This battery is a lens over graded items that already live in the catalog — it introduces no new questions and no new answer keys (every key here was hand-verified for the Trap Rate / Trick Riddles batteries). Models reason freely and commit on a final Answer: line; outcomes are correct / incorrect / abstain / unparseable, each item runs k times, and the score is the majority verdict with a 95% Wilson interval and its n.
Contamination: public. These are famous puzzles, so a model may have seen the answer directly — read a high score as "did not take the bait this time," not as proof of reasoning.
These items are famous, so any model trained after they circulated may know the answers directly. Read scores as a timeline of what still bites, not a measure of intelligence.