spacr.pipeline_graph

The DAG of what produced what, with staleness marked.

spacr.artifacts already knows both halves of this picture and neither half is visible anywhere: upstream_of / downstream_of walk the provenance edges one artifact at a time, and is_stale answers “is this still what it was made from” for one artifact at a time. A user does not have one artifact. They have a project that has been run, partly re-run, and re-run again with a different mask diameter, and the question they actually ask is which of these files can I still believe.

This module turns those per-artifact answers into one whole-project graph:

  • every registered artifact is a Node, carrying its module, kind, path, run id, settings hash, spaCR version and — the point — its Node.state, one of STATE_CURRENT, STATE_STALE or STATE_MISSING, together with the machine causes and the human reasons that spacr.artifacts.Staleness gives;

  • every inputs entry is an Edge, so the shape a user recognises (“mask fed measure fed classify”) is drawn from what actually happened rather than from what the pipeline is supposed to do;

  • PipelineGraph.layers puts each node in a column by its longest distance from a root, which is what makes the drawing readable and is also what a text renderer needs.

Two things are deliberately kept apart:

Stale is not missing. An artifact whose inputs moved on is stale — the number in it no longer follows from the files it names. An artifact that was deleted is missing — an availability problem. spacr.artifacts already refuses to conflate them and so does this: a node can be both, and the Node.state reports missing first because a file that is not there cannot be re-read to check anything else.

The registry is not the only source of truth. A project can have run nothing at all, and a graph of zero nodes tells that user nothing. So module_graph() draws the static DAG from spacr.ports.PORTS — what feeds what, by declaration — and build_graph() records which modules of it have actually produced something. The screen shows the module DAG behind the artifact DAG for exactly that reason: an empty project still gets a picture of the pipeline it is about to run.

Everything here is read-only and headless. Nothing in this module writes to the registry, imports Qt, or raises for a project that has no registry at all — a missing artifacts.db is an empty graph with a note, because “you have not run anything yet” is an answer, not an error.

Public API:

from spacr.pipeline_graph import build_graph, format_graph, module_graph

graph = build_graph("/data/plate7")
print(format_graph(graph))
stale = graph.stale_nodes()
print(module_graph().layers)

Classes

Edge

One provenance edge: source was an input to target.

ModuleGraph

The static module DAG declared by spacr.ports.PORTS.

Node

One registered artifact, with the verdict on whether to believe it.

PipelineGraph

Every artifact of one project and the edges between them.

Functions

build_graph(→ PipelineGraph)

Build the provenance graph of one project, with staleness marked.

format_graph(→ str)

Render a graph as text: one block per column, edges named.

module_graph() → ModuleGraph)

Build the declared module DAG from spacr.ports.PORTS.

stale_summary(→ Dict[str, Any])

Counts and cause tallies for a graph, for a one-line verdict.

to_dot(→ str)

Render a graph as Graphviz DOT, coloured by state.

Module Contents

class spacr.pipeline_graph.Edge[source]

One provenance edge: source was an input to target.

Parameters:
  • source – artifact id that was consumed.

  • target – artifact id that was produced from it.

  • kind – the source’s kind, so an edge can be labelled without a second lookup.

  • dangling – the source id is recorded on the target but is no longer in the registry. The edge is kept — the fact that something was consumed and then forgotten is exactly what makes the target stale, and dropping the edge would hide it.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy of the edge.

class spacr.pipeline_graph.ModuleGraph[source]

The static module DAG declared by spacr.ports.PORTS.

What feeds what by declaration, independent of whether anything has run.

Parameters:
  • modules – every module key in the graph, sorted.

  • edges – (producer, consumer) pairs.

  • layers – modules grouped by longest distance from a source.

  • ran – the subset of modules that has produced a registered artifact in the project this was built alongside; empty for a bare module_graph() call.

next_of(module: str) → Tuple[str, ...][source]

Modules this one can feed, from the declared edges.

Parameters:

module – a module key; a key with no outgoing edge, or not in the graph, gives an empty tuple.

Returns:

the consumer module keys, sorted.

previous_of(module: str) → Tuple[str, ...][source]

Modules that can feed this one, from the declared edges.

Parameters:

module – a module key; a key with no incoming edge, or not in the graph, gives an empty tuple.

Returns:

the producer module keys, sorted.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy of the module graph.

class spacr.pipeline_graph.Node[source]

One registered artifact, with the verdict on whether to believe it.

Parameters:
  • artifact_id – the registry id; the node’s identity in the graph.

  • project – absolute project root.

  • kind – a spacr.ports kind, e.g. "measurements-db".

  • role – the producing module’s port role.

  • module – producing module key.

  • path – absolute path of the file or folder.

  • run_id – the run that produced it, when one was recorded.

  • settings_hash – digest of the material settings.

  • spacr_version – the version that produced it.

  • created_utc – registration time.

  • created_ns – the same instant as time.time_ns(), for ordering.

  • size_bytes – bytes on disk at registration.

  • n_files – files covered.

  • status – the artifact’s own complete / partial / failed.

  • exists – whether the path is still on disk.

  • state – STATE_CURRENT, STATE_STALE or STATE_MISSING.

  • reasons – sentences from spacr.artifacts.Staleness.

  • causes – the matching machine codes, e.g. "upstream-newer".

  • depth – longest distance from a root, i.e. the column to draw it in.

  • inputs – artifact ids this was derived from, as recorded.

__str__() → str[source]

One line: state, module, kind and path.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy of the node.

property label: str[source]

the module and what it produced.

Type:

Short two-line-able label

property stale: bool[source]

True when an input or a setting moved on after this was written.

class spacr.pipeline_graph.PipelineGraph[source]

Every artifact of one project and the edges between them.

Parameters:
  • project – the project root the graph covers; "" for a registry holding several.

  • nodes – every artifact, in draw order (by depth, then time).

  • edges – every provenance edge.

  • layers – node ids grouped by Node.depth.

  • modules – the static module DAG, with ran filled in.

  • generated_utc – when the graph was built.

  • notes – anything a caller should know — no registry, an empty project, a provenance cycle.

  • registry_file – the registry the graph was read from.

__bool__() → bool[source]

True when the graph has at least one artifact.

Spelled out because __len__ alone would make an empty-but-valid graph falsy in a way that reads as “the call failed”, and the two are different: an empty graph is the correct answer for a project that has not been run.

__len__() → int[source]

How many artifacts the graph holds.

by_module() → Dict[str, Tuple[Node, ...]][source]

Nodes grouped by the module that produced them.

downstream(artifact_id: str) → Tuple[Node, ...][source]

Every node reachable from this one, following the edges down.

The “what does re-running this invalidate?” question, answered from the graph already in memory rather than by another registry walk.

Parameters:

artifact_id – the registry id of the starting node. The start itself is not included, and an id with no edges gives an empty tuple.

Returns:

the reachable nodes, ordered by depth, then newest first, then id.

leaves() → Tuple[Node, ...][source]

Nodes nothing was derived from — the current end of the pipeline.

node(artifact_id: str) → Node | None[source]

The node with this id, or None.

Parameters:

artifact_id – the registry id to look up.

roots() → Tuple[Node, ...][source]

Nodes with no registered input — where the project starts.

stale_nodes() → Tuple[Node, ...][source]

Every node that is stale or missing, worst first.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy of the whole graph.

upstream(artifact_id: str) → Tuple[Node, ...][source]

Every node this one was derived from, transitively.

Parameters:

artifact_id – the registry id of the starting node. The start itself is not included, and an id with no edges gives an empty tuple.

Returns:

the ancestor nodes, ordered by depth, then newest first, then id.

spacr.pipeline_graph.build_graph(project: str | os.PathLike | None = None, *, registry: spacr.artifacts.Registry | None = None, settings: Mapping[str, Any] | None = None, limit: int | None = None, all_projects: bool = False) → PipelineGraph[source]

Build the provenance graph of one project, with staleness marked.

Parameters:
  • project – the project root. Ignored when registry is given and all_projects is True.

  • registry – an open spacr.artifacts.Registry; one is opened read-only for project when this is omitted. A project with no registry yields an empty graph carrying a note, never an exception — “nothing has been run here” is a legitimate answer.

  • settings – the settings a caller is about to run with. When given, every node is additionally checked against them, so a graph can show “this would be stale if you ran now” before anything is overwritten. Note that this compares EVERY node’s recorded settings against the one dict, which is what “would my current settings reproduce this?” means.

  • limit – cap on how many artifacts to read, newest first.

  • all_projects – read every project in the registry file rather than one. For the shared registry spacr.artifacts.ARTIFACTS_DB_ENV points at.

Returns:

a PipelineGraph.

spacr.pipeline_graph.format_graph(graph: PipelineGraph, *, width: int = 100) → str[source]

Render a graph as text: one block per column, edges named.

The headless counterpart of the screen, and what the tests read. No trailing newline.

Parameters:
  • graph – the graph to render.

  • width – soft wrap for a path; longer ones are elided in the middle.

spacr.pipeline_graph.module_graph(modules: Sequence[str] | None = None, *, ran: Sequence[str] = ()) → ModuleGraph[source]

Build the declared module DAG from spacr.ports.PORTS.

An edge exists from a to b when b REQUIRES a kind that a produces — the same rule spacr.ports.next_modules() applies, used here for every module at once. Optional consumers are left out on purpose: an optional input is a module that can read something, and drawing those edges turns the picture into a mesh in which the pipeline is no longer visible.

Parameters:
  • modules – restrict to these module keys; default is every key in spacr.ports.PORTS.

  • ran – module keys that have actually produced something, recorded on the result so a drawing can dim the ones that have not.

Returns:

a ModuleGraph.

spacr.pipeline_graph.stale_summary(graph: PipelineGraph) → Dict[str, Any][source]

Counts and cause tallies for a graph, for a one-line verdict.

Parameters:

graph – the graph to summarise.

Returns:

n_nodes, n_edges, n_current, n_stale, n_missing, causes ({code: count}), modules and verdict — one sentence fit for a status bar.

spacr.pipeline_graph.to_dot(graph: PipelineGraph) → str[source]

Render a graph as Graphviz DOT, coloured by state.

Not used by the GUI — it draws itself — but it is what a user pastes into a methods figure, and it is the cheapest way to eyeball a graph while developing. No trailing newline.

Parameters:

graph – the pipeline graph to render. Each node is labelled with its module, kind and path basename and filled by state; dangling edges are drawn dashed red.