Dispatch
Evals are CI/CD for AI (and almost nobody writes them)
Everyone builds agents. Almost nobody writes the tests that decide whether they can ship. An agent isn't judged by feel: you put it under gate, like code. Here's what that looks like in practice, indirect injection included, the test everyone is missing.
evals robustesse compliance