Write the eval before the prompt: eval-driven development for AI features
TDD's failing-test-first loop, applied to AI: write the eval before the prompt, let the score say when an LLM feature actually works, and stop shipping changes you can't measure.