analysis.compare_replication_values

analysis.compare_replication_values(
    values_a,
    values_b,
    *,
    label_a='A',
    label_b='B',
    ci_level=0.95,
)

Compare two independent samples of per-replication values.

Computes a confidence interval on each side independently (never a pooled or paired calculation - the two samples are two different scenarios’ replications, not before/after pairs of the same run), plus a Welch’s t-test p-value as a more rigorous companion figure. Typically called on two replication_means(...)["value"] series, or two resource_utilisation(by="run") columns, for the same metric under two different scenarios.

Parameters

Name Type Description Default
values_a array - like Per-replication values for each scenario - one value per replication, never one value per entity. See mean_confidence_interval’s Notes for why pooling per-entity observations here would understate the interval. required
values_b array - like Per-replication values for each scenario - one value per replication, never one value per entity. See mean_confidence_interval’s Notes for why pooling per-entity observations here would understate the interval. required
label_a str Human-readable names for each scenario, carried through to the output. "A", "B"
label_b str Human-readable names for each scenario, carried through to the output. "A", "B"
ci_level float Confidence level for each side’s interval and for the significance test. 0.95

Returns

Name Type Description
ScenarioComparison Named tuple (label_a, label_b, mean_a, mean_b, delta, delta_pct, ci_a, ci_b, ci_overlap, ci_level, p_value, n_a, n_b). - delta : mean_b - mean_a. - delta_pct : delta as a percentage of mean_a; NaN if mean_a is 0. - ci_a, ci_b : each side’s ConfidenceInterval, from mean_confidence_interval. - ci_overlap : True/False if both intervals are defined, else None if either side has fewer than 2 replications. - p_value : two-sided p-value from Welch’s t-test (scipy.stats.ttest_ind(..., equal_var=False), which does not assume the two samples share a variance); NaN if either side has fewer than 2 replications.

Raises

Name Type Description
ImportError If scipy is not installed - see mean_confidence_interval.

See Also

mean_confidence_interval : The confidence interval computed independently on each side. replication_means : Produces the per-replication values this function compares. resource_utilisation : Also produces one value per run, via by="run".

Notes

ci_overlap=False is a safe “these two scenarios differ” signal: two independent confidence intervals failing to overlap is a stricter condition than a two-sample significance test at the same ci_level. ci_overlap=True does not prove “these are the same” - only “not conclusively different by this simple check” - p_value is the more rigorous figure to read alongside it, not a replacement for looking at both delta and the two intervals.

Back to top