analysis.event_durations
analysis.event_durations(
event_log,
first_event,
second_event,
*,
match='first',
warm_up=0,
entity_col_name='entity_id',
event_col_name='event',
time_col_name='time',
run_col_name='auto',
pathway_col_name='pathway',
keep_incomplete=True,
)Pair occurrences of two events per entity and compute the duration between them.
Replaces a pivot on event_col_name, which raises on any log where an entity revisits a step - a rework loop produces duplicate (entity_id, event) pairs, which pivot cannot place in a single cell - and silently drops pathway.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| event_log | pandas.DataFrame | Long-format event log, e.g. the output of EventLogger.to_dataframe() or TrialLogger.to_dataframe(). |
required |
| first_event | str | The two event names to pair, matched against event_col_name. Must be different. |
required |
| second_event | str | The two event names to pair, matched against event_col_name. Must be different. |
required |
| match | (first, last, occurrence) | How repeated occurrences of the two events for the same entity are paired: - "first": the entity’s earliest first_event with its earliest second_event, regardless of how many times either occurs. - "last": the entity’s latest of each, symmetrically. - "occurrence": the n-th first_event with the n-th second_event, in time order. An entity with an unequal number of the two events has its excess occurrences paired with nothing (NaN on the missing side), and a warning names how many entities are affected. |
"first" |
| warm_up | float | Pairings whose first_time is before warm_up are excluded - the standard truncation rule for duration data (Law & Kelton): a pairing is discarded by when it started, not by whether it later straddles the cutoff. Unlike resource_use_intervals/resource_occupancy_over_time’s warm_up, nothing is censored or clipped mid-interval - a duration is one atomic observation, not a bout that can be partially inside the window. A pairing with no first_time at all (a second_event with no matching first_event) is never excluded by this, since there is no time to compare against warm_up. The default of 0 is a verified no-op. |
0 |
| entity_col_name | str | Column identifying the entity. | "entity_id" |
| event_col_name | str | Column holding the event name. | "event" |
| time_col_name | str | Column holding the event time. | "time" |
| run_col_name | str or None | Column identifying which simulation run each row belongs to; pairing is done within each run rather than across the whole log. "auto" looks for a column named (case-insensitively) one of run, run_number, replication, rep or run_id. Pass an explicit column name to override, or None to disable and pair across the whole frame. If no run column is found, the output run_number column is filled with NA. |
"auto" |
| pathway_col_name | str or None | Column holding the entity’s pathway, carried through to the output. Its absence at the default name is tolerated (the output pathway column is then all NA); an explicit name that is missing raises. |
"pathway" |
| keep_incomplete | bool | If True (default), rows where either the first or second event is missing for a pairing are kept, with NaN in the missing side’s time and in duration. If False, they are dropped. |
True |
Returns
| Name | Type | Description |
|---|---|---|
| pandas.DataFrame | One row per matched pair, with columns entity_id, run_number, pathway, occurrence, first_time, second_time, duration. |
Raises
| Name | Type | Description |
|---|---|---|
| ValueError | If first_event equals second_event; if either is not present in event_col_name; if pathway_col_name names a column that does not exist; or if warm_up is negative. |
Notes
The pairing is an outer join on entity (and run, and occurrence for match="occurrence"), not a left join on first_event. This is what preserves both directions of incompleteness: an entity that reached first_event but never second_event (the familiar “unserved” case), and - equally real, but invisible under a left join - one with a second_event that has no matching first_event.
Where every entity visits each event at most once per run, match="first" reproduces exactly what the old pivot-based calculation produced. Where an entity revisits, the old calculation raised, so there is no prior behaviour to match.
Examples
>>> event_durations(event_log, "treatment_begins", "treatment_ends")