In-depth practical guides

Measure useful AI-agent latency

Time complete tasks, retain failures and separate apparent speed from accepted results.

Updated :

Testing workstation and stop button. Illustrative scene.

AI-generated illustration · Testing workstation and stop button. Illustrative scene.

When does the result become usable?

First displayed content, end of generation and human acceptance are different events. A useful comparison states timing boundaries and expected output. This protocol ranks no vendor; it organises reproducible local measurements.

  1. Define boundaries

    Specify request sent, first content, tools completed, output ready or review finished. Use a duration-appropriate clock. Do not subtract clocks on different machines without known synchronisation.

  2. Fix the scenario

    Keep input, instructions, version, permissions and cache state. Separate first and repeated calls. Record external-tool time rather than attributing everything to the model.

  3. Retain failed trials

    Record durations, attempts, retries and timeouts. Publish count and distribution, not just the best result. An abandoned task remains visible even without a completion duration.

  4. Compare accepted quality

    Apply the same acceptance criterion. Count corrections and review time. Fast invalid output is not useful time saved; state human effort.

Situations and decisions

Typical situations for preparing a check. They do not describe completed assignments or actual observations.

A two-second response

Fictional example: first text at 2 s, tools done at 12 s, acceptance at 90 s. Publish separate stages instead of a two-second completed-task claim.

Timeouts excluded from the mean

Illustrative trial: 9 successes and 1 timeout. Successful durations alone do not describe all ten requests. Publish success rate and exclusions.

Record to retain

  • Timing boundaries and units
  • Versions, input and cache state
  • Durations, failures and retries
  • Quality criterion and human time

Official references

Related protocols

Checks before reaching a conclusion

  • Are boundaries identical?
  • Is cache state recorded?
  • Are failures visible?
  • Is quality judged consistently?

Stage timings linked to quality, failures and acceptance effort.

Frequently asked questions

Should first content be timed?

Yes for perceived responsiveness. Also retain task completion and acceptance to distinguish first display from useful output.

Are two means enough?

They hide variation and failures. Include trial count, conditions, median or distribution and accepted-result rate. State calculation rules.

TOOLS / REGISTERS

Write a record

Local timing depends on task, load, network and tools. Percentiles from few trials are fragile. Controlled trials are not general availability promises.

  • Access date and time zone
  • Exact URL
  • Task / journey step
  • Version / revision
  • Start and end events
  • Stage durations and unit
  • Failures, timeouts and retries
  • Repetitions and dependence between observations
  • Acceptance criterion
  • Limit / uncertainty
  • Decision / next action
  • Review owner
  • Next review