analysis.activity_occupancy_stats

analysis.activity_occupancy_stats(
    event_log,
    *,
    every_x_time_units=1,
    warm_up=0,
    limit_duration=None,
    across_runs='average',
    include_queues=True,
    include_resources=True,
    entity_col_name='entity_id',
    time_col_name='time',
    event_type_col_name='event_type',
    event_col_name='event',
    resource_col_name='resource_id',
    run_col_name='auto',
    pathway_col_name=None,
)

Summarise how many entities were present at each step of a process.

For every queue step (event_type == "queue") and every resource step (event_type == "resource_use"), this takes the per-snapshot occupancy series - queue_size_over_time for queues, resource_occupancy_over_time for resources - and reduces it to the mean, minimum, maximum and median number of entities present. The result is one row per step, intended to be merged onto the node table of a directly-follows graph (discover_dfg(occupancy_stats=...)) so a process map can show queue build-up and resource load alongside the frequency and timing statistics it already carries.

Parameters

Name Type Description Default
event_log pandas.DataFrame Long-format event log spanning one or more runs, e.g. the output of EventLogger.to_dataframe() or TrialLogger.to_dataframe(). Pass the raw log, not one already filtered by time - warm_up is applied here the same way reshape_for_animations applies it. required
every_x_time_units float Time granularity for snapshots. The queue path runs reshape_for_animations once per run, so a small value on a long run is the expensive case; the resource path is a cheap interval sweep either way. 1
warm_up float Start of the reported window, forwarded to the underlying functions. 0
limit_duration float End of the reported window. None (default) uses the latest time seen anywhere in the log. None
across_runs (average, pool) How a multi-run log is combined. "average" computes each statistic within each run and then averages those per-run values, so max is the mean of the per-run maxima - the figure expected per replication. "pool" concatenates every (run, snapshot) count and takes one statistic over the pool, so max is the largest queue seen in any run. A single-run log gives the same answer either way. "average"
include_queues bool Set False to skip queue steps entirely - this is what avoids the reshape_for_animations cost. True
include_resources bool Set False to skip resource steps. True
entity_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
time_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
event_type_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
event_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
resource_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
run_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'
pathway_col_name str or None Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. 'entity_id'

Returns

Name Type Description
pandas.DataFrame Columns event, kind ("queue" or "resource"), mean_occupancy, min_occupancy, max_occupancy, median_occupancy. One row per step. Empty (with those columns) if the log has no queue or resource steps.

Notes

  • An event name logged as both a queue and a resource-use step is reported as a queue only, with a warning - the two occupancy questions cannot share one node.
  • Steps that are neither a queue nor a resource-use step (arrival, depart, a resource_use_end label, a custom milestone) get no row; after the merge in discover_dfg their occupancy columns are NaN.
Back to top