Real captured outputs and agent demos from AI models on knowledge prompts. Every score has a receipt.