Real captured outputs and agent demos from AI models on consulting prompts. Every score has a receipt.