Define scope
List tools, data, permissions, allowed and forbidden actions, and the human approval point. Tie each right to a task.
GUIDE / PROTOCOL
Agent evaluation must test actions, boundaries and failures, not only fluent answers.
Direct answer
Define authorised tasks, expected outcomes and harms to avoid. Build normal, ambiguous and hostile cases, run them without external effects, then inspect every action and stop.
01 / Review method
List tools, data, permissions, allowed and forbidden actions, and the human approval point. Tie each right to a task.
Include routine requests, incomplete data, deceptive source instructions, tool errors, duplicates and attempts to leave the mandate.
Compare outcome, consulted sources, tool calls, refusals, time, cost and human intervention with predefined criteria.
Block critical failures, narrow rights where needed and rerun the same set after changes. Prepare shutdown, manual fallback and withdrawal.
02 / Evidence to keep
Input, expected result, risk, severity and pass criterion.
System version, tool calls, exposed data, outputs and human decision.
Failure, consequence, likely cause, fix and retest result.
Owner, approved scope, stop threshold, fallback and review date.
03 / FAQ
Not necessarily. It may reach the right answer through an unsuitable source or an unauthorised action.
Yes, especially when the agent reads pages, documents or messages that may contain hostile instructions.
Before sensitive, irreversible or explicitly unauthorised action.
Suspend that path, keep the trace, narrow the capability and rerun cases before resuming.
04 / Official references