analysis.resource_use_intervals
analysis.resource_use_intervals(
event_log,
*,
unclosed='censor',
warm_up=0,
limit_duration=None,
entity_col_name='entity_id',
time_col_name='time',
event_type_col_name='event_type',
event_col_name='event',
resource_col_name='resource_id',
run_col_name='auto',
)Pair resource_use/resource_use_end rows into one interval per bout of use.
Parameters
| Name | Type | Description | Default |
|---|---|---|---|
| event_log | pandas.DataFrame | Long-format event log, e.g. the output of TrialLogger.to_dataframe(). |
required |
| unclosed | (censor, drop) | How to handle a resource_use row with no matching resource_use_end - an entity still holding the resource when the window ends. - "censor" (default): the interval’s end is set to the window end and censored is True. Dropping these understates utilisation exactly when it matters most - entities still holding a resource at the end of a run are disproportionately those in a congested system. - "drop": the interval is excluded entirely. |
"censor" |
| warm_up | float | Start of the analysis window. See Notes. | 0 |
| limit_duration | float | End of the analysis window. None (default) uses the latest time seen anywhere in the trial - not per run, so every run shares the same denominator. See Notes. |
None |
| entity_col_name | str | 'entity_id' |
|
| time_col_name | str | 'entity_id' |
|
| event_type_col_name | str | 'entity_id' |
|
| event_col_name | str | 'entity_id' |
|
| resource_col_name | str | Column names in event_log. |
'resource_id' |
| run_col_name | str or None | Column identifying which run each row belongs to. See event_durations. |
"auto" |
Returns
| Name | Type | Description |
|---|---|---|
| pandas.DataFrame | Columns run_number, entity_id, event (the start row’s event name - the end row’s name is a label, not a key), resource_id, start, end, censored, busy_time. |
Raises
| Name | Type | Description |
|---|---|---|
| ValueError | If unclosed is not "censor" or "drop"; or if resource_col_name is present for some resource_use/resource_use_end rows but not others. |
Notes
Pairing is within (run, entity, resource_id), matched in time order (the n-th start of that combination with its n-th end) - mirroring event_durations, but with no match argument, since a resource bout has no ambiguity to choose between.
An end row with no matching start (a logging defect, not something to compute from) is always dropped, with a warning.
If resource_col_name is missing from every row (dropped by to_dataframe()’s dropna(axis=1, how="all"), or present but entirely null), pairing falls back to (run, entity) only, with a warning: busy_time, mean_in_use and utilisation stay exact, only the per-unit breakdown is lost.
busy_time is clipped to the analysis window last, after pairing: max(0, min(end, window_end) - max(start, warm_up)). window_end is limit_duration if given, else the latest time seen anywhere in the trial - a per-run maximum would give each run a different denominator, making utilisation incomparable across runs and any confidence interval computed over them meaningless.
For a censored row (censored=True), end is the window boundary, not a real completion time - safe for busy_time (an interval genuinely was occupied up to that point), but not for a service-time-type analysis. A naive end - start on a censored row understates nothing since it’s right-truncated - filter censored rows out first, or pass unclosed="drop", before treating start/end as observed bout durations.
This assumes each step’s capacity is constant across the whole analysis window. A resource whose capacity varies within the window (e.g. shift-based staffing) will not be represented correctly by a single capacity value - resource_utilisation’s mean_in_use is still a valid time-average, but utilisation (which divides by one static number) is not.
See Also
resource_utilisation : Aggregates this function’s output into one row per run per group.