Clearer task specifications for agents
Most agent failures are decided before the agent runs, in how the task was described. We are testing what a task brief needs to contain — inputs, boundaries, completion criteria, escalation — for the result to be reviewable.