analysis.event_durations

analysis.event_durations(
    event_log,
    first_event,
    second_event,
    *,
    match='first',
    warm_up=0,
    entity_col_name='entity_id',
    event_col_name='event',
    time_col_name='time',
    run_col_name='auto',
    pathway_col_name='pathway',
    keep_incomplete=True,
)

Pair occurrences of two events per entity and compute the duration between them.

Replaces a pivot on event_col_name, which raises on any log where an entity revisits a step - a rework loop produces duplicate (entity_id, event) pairs, which pivot cannot place in a single cell - and silently drops pathway.

Parameters

Name Type Description Default
event_log pandas.DataFrame Long-format event log, e.g. the output of EventLogger.to_dataframe() or TrialLogger.to_dataframe(). required
first_event str The two event names to pair, matched against event_col_name. Must be different. required
second_event str The two event names to pair, matched against event_col_name. Must be different. required
match (first, last, occurrence) How repeated occurrences of the two events for the same entity are paired: - "first": the entity’s earliest first_event with its earliest second_event, regardless of how many times either occurs. - "last": the entity’s latest of each, symmetrically. - "occurrence": the n-th first_event with the n-th second_event, in time order. An entity with an unequal number of the two events has its excess occurrences paired with nothing (NaN on the missing side), and a warning names how many entities are affected. "first"
warm_up float Pairings whose first_time is before warm_up are excluded - the standard truncation rule for duration data (Law & Kelton): a pairing is discarded by when it started, not by whether it later straddles the cutoff. Unlike resource_use_intervals/resource_occupancy_over_time’s warm_up, nothing is censored or clipped mid-interval - a duration is one atomic observation, not a bout that can be partially inside the window. A pairing with no first_time at all (a second_event with no matching first_event) is never excluded by this, since there is no time to compare against warm_up. The default of 0 is a verified no-op. 0
entity_col_name str Column identifying the entity. "entity_id"
event_col_name str Column holding the event name. "event"
time_col_name str Column holding the event time. "time"
run_col_name str or None Column identifying which simulation run each row belongs to; pairing is done within each run rather than across the whole log. "auto" looks for a column named (case-insensitively) one of run, run_number, replication, rep or run_id. Pass an explicit column name to override, or None to disable and pair across the whole frame. If no run column is found, the output run_number column is filled with NA. "auto"
pathway_col_name str or None Column holding the entity’s pathway, carried through to the output. Its absence at the default name is tolerated (the output pathway column is then all NA); an explicit name that is missing raises. "pathway"
keep_incomplete bool If True (default), rows where either the first or second event is missing for a pairing are kept, with NaN in the missing side’s time and in duration. If False, they are dropped. True

Returns

Name Type Description
pandas.DataFrame One row per matched pair, with columns entity_id, run_number, pathway, occurrence, first_time, second_time, duration.

Raises

Name Type Description
ValueError If first_event equals second_event; if either is not present in event_col_name; if pathway_col_name names a column that does not exist; or if warm_up is negative.

Notes

The pairing is an outer join on entity (and run, and occurrence for match="occurrence"), not a left join on first_event. This is what preserves both directions of incompleteness: an entity that reached first_event but never second_event (the familiar “unserved” case), and - equally real, but invisible under a left join - one with a second_event that has no matching first_event.

Where every entity visits each event at most once per run, match="first" reproduces exactly what the old pivot-based calculation produced. Where an entity revisits, the old calculation raised, so there is no prior behaviour to match.

Examples

>>> event_durations(event_log, "treatment_begins", "treatment_ends")
Back to top