analysis.activity_occupancy_stats
analysis.activity_occupancy_stats(
event_log,
*,
every_x_time_units=1,
warm_up=0,
limit_duration=None,
across_runs='average',
include_queues=True,
include_resources=True,
entity_col_name='entity_id',
time_col_name='time',
event_type_col_name='event_type',
event_col_name='event',
resource_col_name='resource_id',
run_col_name='auto',
pathway_col_name=None,
)Summarise how many entities were present at each step of a process.
For every queue step (event_type == "queue") and every resource step (event_type == "resource_use"), this takes the per-snapshot occupancy series - queue_size_over_time for queues, resource_occupancy_over_time for resources - and reduces it to the mean, minimum, maximum and median number of entities present. The result is one row per step, intended to be merged onto the node table of a directly-follows graph (discover_dfg(occupancy_stats=...)) so a process map can show queue build-up and resource load alongside the frequency and timing statistics it already carries.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| event_log | pandas.DataFrame | Long-format event log spanning one or more runs, e.g. the output of EventLogger.to_dataframe() or TrialLogger.to_dataframe(). Pass the raw log, not one already filtered by time - warm_up is applied here the same way reshape_for_animations applies it. |
required |
| every_x_time_units | float | Time granularity for snapshots. The queue path runs reshape_for_animations once per run, so a small value on a long run is the expensive case; the resource path is a cheap interval sweep either way. |
1 |
| warm_up | float | Start of the reported window, forwarded to the underlying functions. | 0 |
| limit_duration | float | End of the reported window. None (default) uses the latest time seen anywhere in the log. |
None |
| across_runs | (average, pool) | How a multi-run log is combined. "average" computes each statistic within each run and then averages those per-run values, so max is the mean of the per-run maxima - the figure expected per replication. "pool" concatenates every (run, snapshot) count and takes one statistic over the pool, so max is the largest queue seen in any run. A single-run log gives the same answer either way. |
"average" |
| include_queues | bool | Set False to skip queue steps entirely - this is what avoids the reshape_for_animations cost. |
True |
| include_resources | bool | Set False to skip resource steps. |
True |
| entity_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| time_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| event_type_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| event_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| resource_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| run_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
| pathway_col_name | str or None | Column names, forwarded to queue_size_over_time / resource_occupancy_over_time. See those functions’ docstrings. |
'entity_id' |
Returns
| Name | Type | Description |
|---|---|---|
| pandas.DataFrame | Columns event, kind ("queue" or "resource"), mean_occupancy, min_occupancy, max_occupancy, median_occupancy. One row per step. Empty (with those columns) if the log has no queue or resource steps. |
Notes
- An event name logged as both a queue and a resource-use step is reported as a queue only, with a warning - the two occupancy questions cannot share one node.
- Steps that are neither a queue nor a resource-use step (
arrival,depart, aresource_use_endlabel, a custom milestone) get no row; after the merge indiscover_dfgtheir occupancy columns areNaN.