analysis.entity_metric_by_arrival
analysis.entity_metric_by_arrival(
event_log,
first_event,
second_event,
*,
arrival_event='arrival',
match='first',
entity_col_name='entity_id',
event_col_name='event',
time_col_name='time',
run_col_name='auto',
pathway_col_name='pathway',
keep_incomplete=True,
)Pair two events per entity, as event_durations does, and attach each entity’s arrival time alongside the resulting duration.
Answers a different question from event_durations alone: not “how long did this interval take”, but “does this duration vary depending on when the entity arrived at the system” - a non-stationary arrival process or a time-of-day/load effect, for example. Underlies plot_metric_vs_arrival_time.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| event_log | pandas.DataFrame | Long-format event log, e.g. the output of EventLogger.to_dataframe() or TrialLogger.to_dataframe(). |
required |
| first_event | str | The two events to pair - see event_durations. |
required |
| second_event | str | The two events to pair - see event_durations. |
required |
| arrival_event | str | The event marking an entity’s arrival at the system. Deliberately independent of first_event/second_event: it can coincide with first_event (e.g. measuring time from arrival itself), but does not have to - the arrival time is looked up separately regardless of which two events the duration itself is measured between. |
"arrival" |
| match | (first, last, occurrence) | How repeated occurrences of first_event/second_event are paired - see event_durations. Does not affect the arrival-time lookup, which always uses the entity’s earliest occurrence of arrival_event regardless of match - an entity ordinarily arrives once. |
"first" |
| entity_col_name | str | Column identifying the entity. | "entity_id" |
| event_col_name | str | Column holding the event name. | "event" |
| time_col_name | str | Column holding the event time. | "time" |
| run_col_name | str or None | Column identifying which simulation run each row belongs to - see event_durations. |
"auto" |
| pathway_col_name | str or None | Column holding the entity’s pathway, carried through to the output - see event_durations. |
"pathway" |
| keep_incomplete | bool | If True (default), rows with no complete first_event/second_event pairing are kept, with NaN duration - see event_durations. |
True |
Returns
| Name | Type | Description |
|---|---|---|
| pandas.DataFrame | One row per matched (first_event, second_event) pair - the same granularity as event_durations - with columns entity_id, run_number, pathway, occurrence, first_time, second_time, duration, arrival_time. An entity with a complete duration pairing but no arrival_event recorded in that run has arrival_time = NaN rather than being dropped. |
Raises
| Name | Type | Description |
|---|---|---|
| ValueError | Everything event_durations raises, plus if arrival_event is not present in event_col_name. |
Notes
The merge joining the duration and arrival frames is keyed on (run_number, entity_id) only, not occurrence - every occurrence-row for one entity (e.g. a rework loop under match="occurrence") shares the same arrival_time, since there is one arrival per entity per run regardless of how many times the metric’s event pair repeats.
See Also
event_durations : The duration pairing this builds on. vidigi.plots.plot_metric_vs_arrival_time : The matching chart.
Examples
>>> entity_metric_by_arrival(event_log, "treatment_wait_begins", "treatment_begins")