A two-second response
Fictional example: first text at 2 s, tools done at 12 s, acceptance at 90 s. Publish separate stages instead of a two-second completed-task claim.
In-depth practical guides
Time complete tasks, retain failures and separate apparent speed from accepted results.
Updated :

AI-generated illustration · Testing workstation and stop button. Illustrative scene.
First displayed content, end of generation and human acceptance are different events. A useful comparison states timing boundaries and expected output. This protocol ranks no vendor; it organises reproducible local measurements.
Specify request sent, first content, tools completed, output ready or review finished. Use a duration-appropriate clock. Do not subtract clocks on different machines without known synchronisation.
Keep input, instructions, version, permissions and cache state. Separate first and repeated calls. Record external-tool time rather than attributing everything to the model.
Record durations, attempts, retries and timeouts. Publish count and distribution, not just the best result. An abandoned task remains visible even without a completion duration.
Apply the same acceptance criterion. Count corrections and review time. Fast invalid output is not useful time saved; state human effort.
Typical situations for preparing a check. They do not describe completed assignments or actual observations.
Fictional example: first text at 2 s, tools done at 12 s, acceptance at 90 s. Publish separate stages instead of a two-second completed-task claim.
Illustrative trial: 9 successes and 1 timeout. Successful durations alone do not describe all ten requests. Publish success rate and exclusions.
Stage timings linked to quality, failures and acceptance effort.
Yes for perceived responsiveness. Also retain task completion and acceptance to distinguish first display from useful output.
They hide variation and failures. Include trial count, conditions, median or distribution and accepted-result rate. State calculation rules.
TOOLS / REGISTERS
Local timing depends on task, load, network and tools. Percentiles from few trials are fragile. Controlled trials are not general availability promises.