Pooled, per-replication, and equal-weight-with-zeros — which TrialLogger gives you, and which you want
A TrialLogger holds one event log per replication. Almost every question you ask it — “what was the mean wait?”, “how big did the queue get?”, “how busy were the nurses?” — needs those replications combined into a single number or curve, and there is more than one defensible way to do that.
This page lays out the three conventions vidigi uses, which method applies which, and how to pick.
The three conventions
1. Pooled (entity-weighted)
Every entity’s value from every replication goes into one pile, and the statistic is taken over the pile. Run boundaries are ignored.
A pooled mean is sum(all values) / count(all values). Because a busy replication contributes more entities, it pulls the pooled figure towards itself — each replication is effectively weighted by its throughput, not counted once.
This is an unbiased estimate of the average entity’s experience when replications are independent and comparable, and it is what every release of vidigi before 2.0 gave. It carries no information about how much the answer varied between replications, so there is no confidence interval attached to it.
2. Per-replication (equal weight)
The statistic is computed within each replication first, giving one value per run, and those per-run values are then averaged — a “mean of means”.
Every replication counts once, regardless of how many entities passed through it. This is the standard unit of analysis for a replicated simulation study: the entities inside one run are strongly correlated (one bad morning makes fifty consecutive waits long together), so the replication, not the entity, is the independent observation. It is also the only basis on which a confidence interval across runs is meaningful — see mean_confidence_interval’s notes, and Law & Kelton, Simulation Modeling and Analysis.
3. Equal weight, with genuine zeros
For time-series and resource metrics — queue length at each snapshot, resource occupancy, utilisation — each replication contributes one observation per snapshot (or per resource), averaged with equal weight, and a replication where the queue was empty or the resource unused contributes a real 0, not a missing value.
Averaging only over the replications where “something happened” biases the result upwards: a queue that formed in one run out of four had an average length near its formed length, not a quarter of it. This convention is what plot_queue_size and the resource plots use.
Per-replication mean of per-run ratios, with zeros
optional (error_bars=)
plot_resource_utilisation_over_time(...)
Equal weight per snapshot, with zeros
—
plot_warm_up_diagnostic(...)
Ensemble mean across runs per index
—
The default for a duration statistic stays pooled, so existing code keeps returning the number it always has. Reach for across="runs" / get_event_duration_ci when you want the replication-level view.
Seeing the difference
from vidigi.logging import EventLogger, TrialLoggerdef make_run(run_number, waits): log = EventLogger(run_number=run_number)for entity_id, wait inenumerate(waits, start=1): log.log_arrival(entity_id=entity_id, time=0.0) log.log_departure(entity_id=entity_id, time=float(wait))return log# Two quiet replications of 2 patients, one busy replication of 6.trial = TrialLogger( [ make_run(1, [10, 12]), make_run(2, [20, 22, 24, 26, 28, 30]), make_run(3, [14, 16]), ])print("pooled (across='entities'):", trial.get_event_duration_stat("arrival", "depart"))print("per-replication (across='runs'):", trial.get_event_duration_stat("arrival", "depart", across="runs"))print("interval over replications:", trial.get_event_duration_ci("arrival", "depart"))
pooled (across='entities'): 20.2
per-replication (across='runs'): 17.0
interval over replications: ConfidenceInterval(mean=np.float64(17.0), half_width=np.float64(17.913371790059195), lower=np.float64(-0.9133717900591947), upper=np.float64(34.913371790059195), n=3, method='t')
The pooled figure sits above the per-replication one: the busy replication puts six entities into the pool but is still just one of three replication means. Which is “right” depends on the question — what does a typical patient experience (pooled) versus what does a typical run of this system look like, and how sure am I (per-replication).
The interval here is wide because three replications with this much spread cannot pin the mean down tightly. That is the honest answer, not a defect — get_replication_precision is the tool for deciding how many more replications to run (Hoad, Robinson & Davies, 2010).
Rules of thumb
Reporting a headline number for a study:get_event_duration_ci — the per-replication mean with its interval.
Comparing scenarios: per-replication (across="runs"), so the comparison is between systems, not between entity populations of different sizes; pair it with an interval or plot_metric_bar(across="runs", error_bars="ci").
Describing the experience of an individual entity (e.g. “what fraction waited more than an hour”): pooled is defensible — you genuinely do want every entity to count.
Queue lengths and resource use over time: the plotting functions already apply the equal-weight-with-zeros convention; you do not need to choose.
Deciding how many replications to run:get_replication_precision / plot_replication_analysis.
Deciding how much warm-up to discard:plot_warm_up_diagnostic.