analysis.event_occurrence_rate

analysis.event_occurrence_rate(
    event_log,
    event_name,
    *,
    event_col_name='event',
    n_runs=None,
    ci_level=0.95,
    run_col_name='auto',
)

Proportion of replications in which an event occurs at least once.

A per-run rate - “in what fraction of runs did this happen at all” - distinct from _summarise_durations’s "unserved_rate"/"served_rate", which are per-entity rates within one event pair. Useful for a rare condition that either happens or doesn’t in a given run (e.g. a capacity breach, a specific alarm event), rather than a duration between two events.

Parameters

Name Type Description Default
event_log pandas.DataFrame Long-format event log, e.g. the output of TrialLogger.to_dataframe(). required
event_name str The event to check for, matched against event_col_name. Occurring for any entity, one or more times, counts a run as an occurrence. Unlike event_durations, a name that matches nothing in the log is not an error here - a rare event legitimately occurring zero times in the runs available is exactly the proportion=0.0 case this function exists to report, so it is not distinguished from a typo. Check event_log[event_col_name].unique() if unsure a name is spelled correctly. required
event_col_name str Column holding the event name. "event"
n_runs int The number of runs to use as the denominator. If None (default), falls back to the number of distinct values in the resolved run column across the whole event_log - which still misses a run that logs no rows at all. Passing n_runs explicitly (e.g. len(trial_logger._event_logs)) is the reliable route whenever the true number of runs is known - see TrialLogger.get_event_occurrence_rate. None
ci_level float Confidence level for the interval. 0.95
run_col_name str or None Column identifying which run each row belongs to, used to count how many runs the event occurred in, and (when n_runs is not given) to infer the denominator. See vidigi.analysis.event_durations’s same parameter. "auto"

Returns

Name Type Description
ProportionEstimate Named tuple (proportion, lower, upper, n_runs, n_occurred, ci_level, method). lower/upper come from a Wilson score interval, which - unlike a Student-t interval - stays within [0, 1] and is well behaved near 0 or 1, exactly where a rare-event rate typically sits. Unlike ConfidenceInterval, there is no half_width: a Wilson interval is asymmetric around proportion near the boundaries, so a single half-width would misrepresent it.

Raises

Name Type Description
ValueError If the resolved n_runs is 0 or less, or if ci_level is not in (0, 1).
ImportError If scipy is not installed - see mean_confidence_interval.

See Also

_summarise_durations : The per-entity, per-event-pair analogue (unserved_rate/served_rate).

Notes

Unlike mean_confidence_interval, there is no low-n warning: the Wilson interval is well-defined for any n_runs >= 1, including n_occurred of 0 or n_runs - a mean’s confidence interval needs at least 2 points to estimate a spread, but a proportion’s does not.

Back to top