spacr.benchmark¶
Measure throughput and memory use to recommend a worker count.
WHY A DEFAULT NEEDS THIS. spaCR’s worker defaults are arithmetic on the core
count – cpu_count() - 4, cpu_count() // 2, -1. Cores are the one
thing that is never the binding constraint on a measurement run: a field is
hundreds of megabytes decompressed, and eight workers on a 16-core laptop
with 16 GB will swap long before they saturate the CPU. The number that
matters is how much MEMORY one worker needs for one representative field, and
that is a property of the plate and the machine, not of the core count.
So: run the real work over a few fields, watch peak RSS and peak VRAM, and divide what is available by what one worker actually took.
This measures; it does not tune. It returns a recommendation and the
evidence behind it, and never writes a setting. A benchmark that silently
changed n_jobs would make a run’s speed depend on when the benchmark last
ran, which is the opposite of reproducible.
Qt-free and torch-optional: VRAM is reported when torch is present with a
CUDA device and reported as None otherwise, which is not the same as 0.
Classes¶
What one benchmark run observed. |
|
A worker count and the retained evidence supporting it. |
Functions¶
|
Memory this machine can actually give to workers. |
|
Run |
|
The benchmark as a few lines a user can read and paste into an issue. |
|
How many workers this machine can actually feed. |
Module Contents¶
- class spacr.benchmark.Measurement[source]¶
What one benchmark run observed.
- Parameters:
items – how many representative units were processed.
seconds – wall clock for all of them.
peak_rss_bytes – the highest resident set size seen, for the whole process. Not per worker – the benchmark runs serially on purpose, so this IS one worker’s requirement.
peak_vram_bytes – peak CUDA allocation, or
Nonewhen there is no CUDA device.Noneand0are different answers: one means “not measured”, the other “measured, and it used none”.baseline_rss_bytes – RSS before the work started, so the caller can tell the interpreter’s own footprint from the work’s.
notes – caveats about warm-up or unavailable measurements that must accompany the numeric result.
- property items_per_second: float[source]¶
Return measured throughput, or NaN for a nonpositive duration.
- class spacr.benchmark.Recommendation[source]¶
A worker count and the retained evidence supporting it.
- Parameters:
workers – final recommended parallel-worker count; always at least one and bounded by the normalized core count, measured memory capacity when usable, and configured maximum.
reason – human-readable explanation of the branch that set
workers: missing measurement, one-worker fallback, memory bound, core bound, or configured maximum.measurement – exact
Measurementsupplied torecommend_workers();Nonemeans no benchmark was available, while a zero-work-footprint measurement is retained but triggers the core-count fallback.cores – effective logical-core ceiling after defaulting from
os.cpu_count()and clamping to at least one.available_bytes – available-memory snapshot before the configured reserve is subtracted; supplied by the caller or measured by spaCR, and zero when unavailable.
- spacr.benchmark.available_memory_bytes() int[source]¶
Memory this machine can actually give to workers.
MemAvailablefrom/proc/meminforather than total: total includes what is already in use, and sizing workers against it is how a run gets OOM-killed at field 900.
- spacr.benchmark.benchmark(work: Callable[[Any], Any], items: Sequence[Any], *, warmup: int = 1) Measurement[source]¶
Run
workoveritemsserially and record what it cost.SERIALLY ON PURPOSE. The question is what ONE worker needs, and running them in parallel measures the sum while hiding the per-worker figure that the recommendation divides by.
- Parameters:
work – callable invoked once for every warm-up and measured item; its return value is ignored because only resource use is measured.
items – ordered workload to process. At least one item is required, and the final item is always kept in the measured set.
warmup – items processed before the clock starts. The first field pays for imports, CUDA context creation and page faults that no later field pays again, and counting it makes a short run look far slower than the plate it is predicting.
- Raises:
ValueError – no items to measure – a benchmark over nothing would return a per-item cost of NaN and a recommendation built on it.
- spacr.benchmark.format_report(measurement: Measurement, recommendation: Recommendation) str[source]¶
The benchmark as a few lines a user can read and paste into an issue.
- Parameters:
measurement – observed serial benchmark costs and throughput.
recommendation – worker recommendation derived from those costs.
- spacr.benchmark.recommend_workers(measurement: Measurement | None = None, *, cores: int | None = None, available_bytes: int | None = None, reserve_bytes: int = 2 * 1024**3, maximum: int = 32) Recommendation[source]¶
How many workers this machine can actually feed.
The rule, in order:
Never more than
cores. More workers than cores is contention.Never more than
available memory - reservedivided by what ONE worker measurably needed. This is the term the core-count defaults omit, and it is usually the binding one.Never fewer than 1, and never more than
maximum.
- Parameters:
reserve_bytes – memory left for everything that is not a worker – the GUI, the page cache the readers depend on, and the operating system. Defaults to 2 GiB.
- Returns:
a
Recommendationcarrying the reason, so a number a user disagrees with can be argued with rather than just overridden.