analysis.mean_confidence_interval

analysis.mean_confidence_interval(values, *, ci_level=0.95, method='t')

Confidence interval for a mean, computed over independent replicate values.

Parameters

Name Type Description Default
values array - like The replicate-level values to summarise - typically the "value" column of replication_means’s output. Must be one value per replication, never one value per entity - see Notes. required
ci_level float Confidence level, e.g. 0.95 for a 95% interval. 0.95
method t How the interval is computed. Only Student’s t is supported: with n typically in the 5-30 range for a replicated simulation study, a normal approximation is systematically too narrow (at n=5, t=2.776 vs z=1.96 - 29% too narrow), so t with n - 1 degrees of freedom is used unconditionally, never z. "t"

Returns

Name Type Description
ConfidenceInterval Named tuple of (mean, half_width, lower, upper, n, method). If fewer than 2 non-missing values are given, half_width, lower and upper are NaN and a warning is raised - an interval needs at least 2 points to estimate a spread from.

Raises

Name Type Description
ImportError If scipy is not installed. Install it with pip install vidigi[stats].
ValueError If method is not "t".

See Also

replication_means : Produces the per-replication values this function summarises.

Notes

This must be computed over replication means, never pooled per-entity observations. Entities within a single run are strongly serially correlated - one bad morning makes fifty consecutive waits long together - so pooling treats correlated observations as independent, shrinking the standard error by roughly the square root of the number of entities per run and producing an interval that can be an order of magnitude too narrow. Replications are the independent unit; entities are not.

At n=2, t gives a large interval (t=12.71 at the 95% level) - this is correct and informative given only two replications, not a bug to special-case.

Back to top