Verdict
The counterweight to optimistic agents: it requires fresh command output as evidence before any claim that work is fixed, passing, or complete. Pairing it with a planning or implementation skill removes most false success reports.
Best for
Any change where “it should work” is not good enough and the agent has a command it can actually run.
Keep in mind
It cannot manufacture a verification path where none exists, and on projects without runnable checks it degrades into a reminder rather than proof.
Example task
Before telling me the migration is done, run the test suite and the local smoke script and show me the output of both.