spacr.resource_log

What spaCR’s process TREE costs while it runs, under its own setting.

WHY THIS IS NOT THE READINGS spaCR ALREADY TAKES.

Every resource figure in the package counts the CALLING process: spacr.fit_resources.host_rss reads /proc/self/statm, spacr.qt.timing reads its own resident size, and the parameter sweep’s floor reads the MACHINE’s free memory, which cannot tell spaCR’s own children from another tenant on a shared box.

spaCR’s heaviest work does not happen in the calling process – spacr.sequencing starts a saver process and spacr.parameter_sweep runs every trial in a child – so the parent looks healthy right up to the moment the out-of-memory reaper takes the run, and afterwards there is nothing to read.

This module sums the process and every descendant, and names each one, so “which trial was large” is a question the record can answer.

WHY IT IS NOT VERBOSE LOGGING. Verbose logging only decides which log records are kept, so it has no account of memory to give. The function tracer that it once installed fired on every call and every return, and it cost twenty times the startup. An account taken through a tracer would describe the traced program rather than the real one, which is exactly the program nobody wants measured. So this samples instead: one psutil read a second, on a daemon thread that is never the GUI thread, on an otherwise unperturbed run. Three states rather than a checkbox, because the useful default is not “off” – the most valuable resource data comes from runs nobody expected to fail.

WHICH NUMBER IS RECORDED. USS where the platform gives it, then PSS, then RSS – and every record NAMES the measure it used, because RSS double-counts the pages a fork shares and would overstate a sweep badly. A number whose definition is unrecorded cannot be compared between two machines.

Nothing here is required to succeed. A child that exits between being enumerated and being read is an expected outcome and not an error: that child is skipped, the rest of the tree is kept, and the count of skipped readings goes in the sample so the record says what it missed. A platform that cannot supply a per-thread time records that it could not, never a zero, because a zero reads as “this thread was free”. Per-thread GPU memory is absent on purpose: a CUDA context belongs to a process, so a per-thread figure would be fiction.

Classes

ResourceSampler

A bounded background record of what the process tree costs.

Functions

describe(→ str)

The peaks as a person reads them, for a log line or a support request.

level_source(→ str)

What decided the level, for a test and for a support request.

preferred_measure(→ Optional[str])

Which memory definition this platform can supply.

read_log(→ Dict[str, Any])

Read a written log back, tolerating a run that was killed mid-line.

resolve_level(→ str)

Which of LEVELS is in force.

summarise(→ Dict[str, Any])

Totals and peaks over recorded samples, and which pid held the peak.

tree_sample(→ Dict[str, Any])

One reading of this process and every descendant.

Module Contents

class spacr.resource_log.ResourceSampler(path: Any = None, level: str | None = None, interval: float = DEFAULT_INTERVAL_SECONDS, capacity: int = DEFAULT_CAPACITY, label: str | None = None, clock=time.time)[source]

A bounded background record of what the process tree costs.

A daemon thread takes one reading every effective interval seconds into a ring buffer of capacity samples. The default settings retain the most recent hour in memory; custom settings retain approximately capacity * interval seconds. A run that lasts a week therefore cannot grow the in-memory series without limit.

The thread is a daemon and is never the GUI thread: it cannot hold the process open at exit and it cannot delay a repaint.

When a path is given, each sample is written as one JSON line and flushed, after a header line naming the level, the measure, the interval and the start. Registering that file against a run is the caller’s job – this class is imported by worker processes that have no artifacts database and no GUI.

At level "off" no thread is started and no file is opened, which is what a thread census before and after a run is entitled to see.

Prepare a sampler without starting it.

Parameters:
  • path – where to write the series, or None to keep it only in memory.

  • level – one of LEVELS, or None to resolve one.

  • interval – seconds between readings, floored at MIN_INTERVAL_SECONDS; together with capacity, this determines the retained time span.

  • capacity – maximum number of in-memory samples to retain, clamped to at least one; older samples are discarded.

  • label – what this record is OF – a run id, a sweep trial – so a file found later can be matched to the work that made it.

  • clock – the time source, passed in so a test can drive it.

Raises:

ValueError – if an explicit level is not one of LEVELS.

__enter__() → ResourceSampler[source]

Start sampling for the duration of a block.

Returns:

this sampler.

__exit__(*exc_info) → bool[source]

Stop sampling, whatever ended the block.

Parameters:

exc_info – the exception the block raised, if any.

Returns:

False, so an exception in the block still propagates.

describe() → str[source]

The peaks as a person reads them.

Returns:

what describe() returns, "" when nothing was recorded.

is_running() → bool[source]

Whether a sampler thread is alive.

Returns:

True while the thread is running.

sample_once() → Dict[str, Any] | None[source]

Take one reading now, keep it and write it.

Public and separate from the loop so a caller – a stage boundary, a test – can take a reading at a moment it chooses rather than waiting for the interval to come round.

Returns:

the sample, or None at level "off".

samples() → List[Dict[str, Any]][source]

Every sample still in the ring buffer, oldest first.

Returns:

a copy, so the caller can read it while sampling continues.

start() → bool[source]

Begin sampling on a daemon thread.

Returns:

whether a sampler thread is now running, which is False at level "off".

stop(timeout: float = 5.0) → bool[source]

Stop sampling, join the thread and close the file.

Parameters:

timeout – seconds to wait for the thread to end.

Returns:

whether no sampler thread remains.

summary() → Dict[str, Any][source]

Totals and peaks over what has been recorded.

Returns:

what summarise() returns, empty when nothing was recorded.

spacr.resource_log.describe(samples: Sequence[Mapping[str, Any]]) → str[source]

The peaks as a person reads them, for a log line or a support request.

Parameters:

samples – records from tree_sample().

Returns:

the lines, or "" when nothing was recorded.

spacr.resource_log.level_source(level: str | None = None) → str[source]

What decided the level, for a test and for a support request.

Parameters:

level – the same argument resolve_level() takes.

Returns:

one of SOURCES.

Raises:

ValueError – if an explicit level is not one of LEVELS.

spacr.resource_log.preferred_measure(process: Any = None) → str | None[source]

Which memory definition this platform can supply.

Parameters:

process – the process to probe, or None for this one.

Returns:

one of MEASURES, or None when nothing can be read.

spacr.resource_log.read_log(path: Any) → Dict[str, Any][source]

Read a written log back, tolerating a run that was killed mid-line.

One JSON object per line is the format that survives a kill: everything written before the kill parses, and the partial last line is dropped rather than making the file unreadable.

Parameters:

path – the file a ResourceSampler wrote.

Returns:

header (empty when the file has none), samples, and unreadable, the number of lines that could not be parsed.

spacr.resource_log.resolve_level(level: str | None = None) → str[source]

Which of LEVELS is in force.

Parameters:

level – an explicit level, or None to resolve one from the environment, then the stored preference, then DEFAULT_LEVEL.

Returns:

one of LEVELS.

Raises:

ValueError – if an explicit level is not one of LEVELS.

spacr.resource_log.summarise(samples: Sequence[Mapping[str, Any]]) → Dict[str, Any][source]

Totals and peaks over recorded samples, and which pid held the peak.

Empty when nothing was recorded – NOT zero, for the reason spacr.fit_resources.peak gives: “nothing was using memory” and “nobody measured” are opposite findings, and a summary that spells the second as the first invites a reader to conclude the run was cheap.

Parameters:

samples – records from tree_sample().

Returns:

samples, measure, missed and pids always; peak_total and peak_total_time when any tree total was read; peak_process naming the pid that held the largest single share; cpu_seconds, the largest CPU total seen in one sample, when any CPU time was read.

spacr.resource_log.tree_sample(level: str | None = None, process: Any = None, now: float | None = None) → Dict[str, Any][source]

One reading of this process and every descendant.

Parameters:
  • level – "summary" or "detailed", resolved from the environment and the preference when None. "detailed" adds per-thread CPU times. "off" governs the background sampler rather than a reading a caller asks for outright, and reads as "summary" here.

  • process – the root of the tree, or None for this process.

  • now – the timestamp to stamp, or None for the wall clock.

Returns:

a record keyed record, time, level, measure, unit, total, processes and missed. total and measure are None when nothing could be read, which is not the same as zero. missed counts processes that vanished or refused to be read while the tree was walked.