Latest — 10 Jul 2026 Compared to What? Every impressive number is quietly missing its bottom half. A three-question test for “97% token reduction” claims.
Teaching an algorithm Sanzo Wada's eye The colours were easy. Teaching the metric how Wada combines them took five tries
Answers Aren't Work Models produce answers. Work is deciding which answers are allowed to touch production.
Prompting Is Not a Safety Boundary A prompt is not a control. If the only thing between the agent and the blast radius is language, you don’t have a boundary.
"Green" Isn't Done A passing suite can still be evidence of nothing. Green means the tests ran, not that the risk is covered.