analysis.replication_means

analysis.replication_means(
    durations,
    *,
    value_col='duration',
    run_col='run_number',
    what='mean',
    **kwargs,
)

Reduce a per-entity durations frame to one value per replication.

The independent unit for any confidence interval is the replication, not the entity - entities within a run are strongly serially correlated, so this is the function that produces the values mean_confidence_interval must be called on. See the Notes there for why.

Parameters

Name Type Description Default
durations pandas.DataFrame Per-entity durations, e.g. the output of event_durations. required
value_col str Column to aggregate within each run. "duration"
run_col str Column identifying which run each row belongs to. "run_number"
what str The per-replication statistic to compute: one of "mean", "median", "max", "min", "quantile", "std", "var", "sum" - a genuine pandas Series method, callable on a single run’s values. Entity-counting aggregations such as "count" or "unserved_rate" answer a different question (how many entities, not what value) and are not meaningful re-averaged across runs; use TrialLogger.get_event_duration_stat or vidigi.plots.plot_metric_bar(across="entities") for those instead. "mean"
**kwargs dict Additional keyword arguments passed to the chosen statistic, e.g. q=0.9 for what="quantile". {}

Returns

Name Type Description
pandas.DataFrame Columns run_col and "value", one row per run that has at least one non-missing value_col. Rows with a missing (NaN) value_col - an incomplete pairing - are dropped before aggregating, since a missing duration cannot contribute to a run’s statistic.

Raises

Name Type Description
ValueError If what is not one of the supported per-replication statistics.

See Also

mean_confidence_interval : Compute a confidence interval over this function’s output. event_durations : The underlying per-entity durations.

Back to top