Rubrics and Evals Are Not Correct from Day One. Grow Them with Failure Logs
A small verification of how coarse evaluation criteria can pass weak LLM output. This article shows how to treat rubrics and evals as operational assets that improve through failure logs.