analysis.event_occurrence_rate
analysis.event_occurrence_rate(
event_log,
event_name,
*,
event_col_name='event',
n_runs=None,
ci_level=0.95,
run_col_name='auto',
)Proportion of replications in which an event occurs at least once.
A per-run rate - “in what fraction of runs did this happen at all” - distinct from _summarise_durations’s "unserved_rate"/"served_rate", which are per-entity rates within one event pair. Useful for a rare condition that either happens or doesn’t in a given run (e.g. a capacity breach, a specific alarm event), rather than a duration between two events.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| event_log | pandas.DataFrame | Long-format event log, e.g. the output of TrialLogger.to_dataframe(). |
required |
| event_name | str | The event to check for, matched against event_col_name. Occurring for any entity, one or more times, counts a run as an occurrence. Unlike event_durations, a name that matches nothing in the log is not an error here - a rare event legitimately occurring zero times in the runs available is exactly the proportion=0.0 case this function exists to report, so it is not distinguished from a typo. Check event_log[event_col_name].unique() if unsure a name is spelled correctly. |
required |
| event_col_name | str | Column holding the event name. | "event" |
| n_runs | int | The number of runs to use as the denominator. If None (default), falls back to the number of distinct values in the resolved run column across the whole event_log - which still misses a run that logs no rows at all. Passing n_runs explicitly (e.g. len(trial_logger._event_logs)) is the reliable route whenever the true number of runs is known - see TrialLogger.get_event_occurrence_rate. |
None |
| ci_level | float | Confidence level for the interval. | 0.95 |
| run_col_name | str or None | Column identifying which run each row belongs to, used to count how many runs the event occurred in, and (when n_runs is not given) to infer the denominator. See vidigi.analysis.event_durations’s same parameter. |
"auto" |
Returns
| Name | Type | Description |
|---|---|---|
| ProportionEstimate | Named tuple (proportion, lower, upper, n_runs, n_occurred, ci_level, method). lower/upper come from a Wilson score interval, which - unlike a Student-t interval - stays within [0, 1] and is well behaved near 0 or 1, exactly where a rare-event rate typically sits. Unlike ConfidenceInterval, there is no half_width: a Wilson interval is asymmetric around proportion near the boundaries, so a single half-width would misrepresent it. |
Raises
| Name | Type | Description |
|---|---|---|
| ValueError | If the resolved n_runs is 0 or less, or if ci_level is not in (0, 1). |
|
| ImportError | If scipy is not installed - see mean_confidence_interval. |
See Also
_summarise_durations : The per-entity, per-event-pair analogue (unserved_rate/served_rate).
Notes
Unlike mean_confidence_interval, there is no low-n warning: the Wilson interval is well-defined for any n_runs >= 1, including n_occurred of 0 or n_runs - a mean’s confidence interval needs at least 2 points to estimate a spread, but a proportion’s does not.