analysis.replication_means
analysis.replication_means(
durations,
*,
value_col='duration',
run_col='run_number',
what='mean',
**kwargs,
)Reduce a per-entity durations frame to one value per replication.
The independent unit for any confidence interval is the replication, not the entity - entities within a run are strongly serially correlated, so this is the function that produces the values mean_confidence_interval must be called on. See the Notes there for why.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| durations | pandas.DataFrame | Per-entity durations, e.g. the output of event_durations. |
required |
| value_col | str | Column to aggregate within each run. | "duration" |
| run_col | str | Column identifying which run each row belongs to. | "run_number" |
| what | str | The per-replication statistic to compute: one of "mean", "median", "max", "min", "quantile", "std", "var", "sum" - a genuine pandas Series method, callable on a single run’s values. Entity-counting aggregations such as "count" or "unserved_rate" answer a different question (how many entities, not what value) and are not meaningful re-averaged across runs; use TrialLogger.get_event_duration_stat or vidigi.plots.plot_metric_bar(across="entities") for those instead. |
"mean" |
| **kwargs | dict | Additional keyword arguments passed to the chosen statistic, e.g. q=0.9 for what="quantile". |
{} |
Returns
| Name | Type | Description |
|---|---|---|
| pandas.DataFrame | Columns run_col and "value", one row per run that has at least one non-missing value_col. Rows with a missing (NaN) value_col - an incomplete pairing - are dropped before aggregating, since a missing duration cannot contribute to a run’s statistic. |
Raises
| Name | Type | Description |
|---|---|---|
| ValueError | If what is not one of the supported per-replication statistics. |
See Also
mean_confidence_interval : Compute a confidence interval over this function’s output. event_durations : The underlying per-entity durations.