In-depth practical guides

Measure uncertainty in an AI citation rate

Separate observed rate, sample size and statistical uncertainty when interpreting an AI answer observatory.

Updated :

Reports and charts on a data review desk. Illustrative scene with no real data.

Reports and charts on a data review desk. Illustrative scene with no real data.

Is a citation rate precise enough to guide a decision?

A proportion without a denominator conceals its fragility. Four cited answers out of ten and forty out of one hundred have the same observed rate, but different precision under a binomial model. A Wilson interval explores that difference. It does not turn a selected question panel into a representative user sample.

  1. Define success and denominator

    Choose before collection what counts: a link, a citation supporting the claim or an accurate answer. Score each observation as binary under this rule. Document failures, exclusions and inaccessible responses without changing the rule after seeing results.

  2. Retain the protocol

    Fix language, interface, country, model or mode and period. Repeated prompts in one conversation may be dependent. Separate strata when conditions change; repeating one question does not create diverse needs.

  3. Read rate and interval

    The tool applies 95% Wilson for independent binary observations with stable probability. It remains defined with no successes or all successes. Confidence concerns the method’s coverage over comparable repetitions, not the probability that a citation is true.

  4. Decide with care

    Retain the raw rate and counts. If the interval crosses the operational threshold, plan a better defined collection before concluding. Before/after observations using the same questions are paired and require change analysis; this tool does not test a difference.

Calculate a 95% Wilson interval

Calculated in your browser. Files and entered values are not sent to the server. UTF-8

Ready to calculate.

Only under the confirmed assumptions. This interval does not correct panel selection or test a difference between two series.

Situations and decisions

Typical situations for preparing a check. They do not describe completed assignments or actual observations.

Four out of ten

The tool returns 40% and about 16.8–68.7%. Under the stated assumptions, the result is compatible with a broad range of rates. It does not establish universal ranking.

Zero out of ten

The observed rate is zero, but the upper bound is about 27.8%. No citations in this sample does not establish that citation is impossible.

Record to retain

  • Success definition, question panel, exclusions and collection context.
  • Success count k, observation count n and series date.
  • Wilson method, 95% level, assumptions and decision threshold.

Official references

Related protocols

Checks before reaching a conclusion

  • Was success defined before counting?
  • Are observations independent and the protocol stable?
  • Is the result separate from assessment of citation accuracy?

Frequently asked questions

Does 95% mean 95% of answers will be cited?

No. The rate is k/n. The 95% level qualifies the interval method under the stated assumptions.

Do overlapping intervals prove no difference?

No. Overlap is not a difference test, and paired observations require appropriate treatment.