Four out of ten
The tool returns 40% and about 16.8–68.7%. Under the stated assumptions, the result is compatible with a broad range of rates. It does not establish universal ranking.
In-depth practical guides
Separate observed rate, sample size and statistical uncertainty when interpreting an AI answer observatory.
Updated :

Reports and charts on a data review desk. Illustrative scene with no real data.
A proportion without a denominator conceals its fragility. Four cited answers out of ten and forty out of one hundred have the same observed rate, but different precision under a binomial model. A Wilson interval explores that difference. It does not turn a selected question panel into a representative user sample.
Choose before collection what counts: a link, a citation supporting the claim or an accurate answer. Score each observation as binary under this rule. Document failures, exclusions and inaccessible responses without changing the rule after seeing results.
Fix language, interface, country, model or mode and period. Repeated prompts in one conversation may be dependent. Separate strata when conditions change; repeating one question does not create diverse needs.
The tool applies 95% Wilson for independent binary observations with stable probability. It remains defined with no successes or all successes. Confidence concerns the method’s coverage over comparable repetitions, not the probability that a citation is true.
Retain the raw rate and counts. If the interval crosses the operational threshold, plan a better defined collection before concluding. Before/after observations using the same questions are paired and require change analysis; this tool does not test a difference.
Calculated in your browser. Files and entered values are not sent to the server. UTF-8
Only under the confirmed assumptions. This interval does not correct panel selection or test a difference between two series.
Typical situations for preparing a check. They do not describe completed assignments or actual observations.
The tool returns 40% and about 16.8–68.7%. Under the stated assumptions, the result is compatible with a broad range of rates. It does not establish universal ranking.
The observed rate is zero, but the upper bound is about 27.8%. No citations in this sample does not establish that citation is impossible.
No. The rate is k/n. The 95% level qualifies the interval method under the stated assumptions.
No. Overlap is not a difference test, and paired observations require appropriate treatment.