spacr.active_learning

Active-learning queue — order unannotated crops by model uncertainty.

Annotation is the bottleneck in every screen. A crop the model already calls with 0.999 teaches the model nothing when a human labels it; the crops worth a person’s afternoon are the ones sitting on the decision boundary. This module turns an already-scored measurements.db into a work queue ordered so the informative crops come first.

Public API

least_confidence(probs) / margin(probs) / entropy(probs)

Per-row uncertainty scores. Pure numpy — no torch, no database.

disagreement(prob_sets)

Spread across an ensemble or a set of MC-dropout passes.

rank_by_uncertainty(probs, measure=…)

Row indices, most uncertain first, deterministically tie-broken.

build_queue(db_path, annotation_column, …)

The queue itself, read straight out of png_list.

queue_rows(queue)

The queue as [(png_path, None), …] — the shape the Annotate screen already paginates (see spacr.qt.annotate_engine.fetch_page()).

format_queue_summary(queue)

The queue’s shape, class balance and caveats, as text.

probabilities_from_logits(logits) / as_probabilities(scores)

Raw-head-output → probability matrix, both head shapes handled.

predict_probabilities(model, batches)

Optional live-model bridge; the only function that touches torch.

Things this module refuses to get wrong

A softmax is not a probability. Everything here is called an uncertainty score, never a confidence. Modern networks are badly calibrated — typically over-confident, and more so the deeper they are (Guo et al., 2017) — so a 0.87 from the head is not “87 % sure”. The scores are used for one thing only: putting crops in an order. Nothing in this module reports a calibrated probability, and neither should anything built on it. See CALIBRATION_NOTE.

Already-annotated crops never enter the queue. NULL in the annotation column is the abstention marker — the crop has not been looked at. 0 is a real class, and a queue that re-serves it wastes exactly the resource this module exists to save. The two are distinguished by IS NULL, never by falsiness. (Same convention as spacr.agreement.)

The two head shapes are handled separately. The classifier head emits either a single logit (binary; needs a sigmoid) or C logits (multiclass; needs a softmax) — see spacr.deep_spacr.apply_model_to_tar(). Pushing a single-logit column through a softmax yields a column of 1.0 and destroys the ordering; pushing C logits through a sigmoid can invert it. Both are silent failures that produce a confident-looking queue full of nonsense, so probabilities_from_logits() branches on the shape and the tests pin both directions.

margin and least_confidence are the same ranking on two classes. For C = 2, 1 − (p₁ − p₂) = 2·min(p, 1−p) = 2·(1 − max p): a linear transform, so identical order and identical ties. They are not two independent choices on a binary screen — they differ only from three classes up, where margin looks at the top two classes while least_confidence looks only at the top one and entropy looks at the whole distribution.

Pure uncertainty ranking collapses onto one region of feature space. The 100 most uncertain crops on a real plate are routinely 100 near copies from the same two wells — one ambiguity, labelled a hundred times, for almost the information of labelling it once. So the queue is diversified by default (round-robin across wells) and format_queue_summary() prints how many wells the queue actually covers. What that costs is stated plainly in build_queue().

Uncertainty sampling skews toward the majority class. On a screen that is 98 % negative, the decision boundary is mostly populated by negatives, so the queue will be too. That is not a bug to hide — it is reported as a class-balance table so the annotator can see it and cap or rebalance if they want to.

Determinism. Same inputs, same seed, same order — including ties. Ties break on row index by default; seed swaps in a seeded permutation, which is still reproducible. Nothing here consults an unseeded RNG.

Nothing in this module imports torch at module scope; the ranking maths is numpy only, so the Qt screen can build a queue without waking a multi-second import chain. predict_probabilities() imports torch lazily and is the only entry point that needs it.

Classes

RoundResult

What one retrain round produced.

StoppingVerdict

Whether the last stretch of annotation bought anything measurable.

Functions

annotation_coverage(→ pandas.DataFrame)

Summarize annotation coverage by class, well, plate and acquisition.

as_probabilities(→ numpy.ndarray)

Read stored scores as an (N, C) probability matrix.

build_queue(→ pandas.DataFrame)

Build the annotation queue: unlabelled crops, most uncertain first.

crops_for_object_keys(→ List[Tuple[str, Optional[int]]])

Resolve object keys to crop rows, in the caller's order.

disagreement(→ numpy.ndarray)

How much several score sets disagree about each crop.

ensure_round_tables(→ None)

Create ROUND_TABLE and ROUND_LOG_TABLE if absent.

entropy(→ numpy.ndarray)

Shannon entropy −Σ p log p of each row.

format_coverage_summary(→ str)

Render annotation_coverage() as text, worst concentration first.

format_learning_curve(→ str)

Render the round-by-round curve and the stopping verdict as text.

format_queue_summary(→ str)

Render a queue's shape, class balance and caveats as plain text.

holdout_report(→ Dict[str, Any])

Held-out metrics, with the confusion matrix they were derived from.

label_rounds(→ pandas.DataFrame)

Per-label round provenance as a frame (empty when never recorded).

learning_curve(→ pandas.DataFrame)

Held-out accuracy per round, oldest first — the curve to watch flatten.

least_confidence(→ numpy.ndarray)

1 − max_c p_c — how much probability mass is not on the winner.

margin(→ numpy.ndarray)

1 − (p₁ − p₂) — closeness of the top two classes, as uncertainty.

next_round(→ int)

The round number the next batch of labels belongs to.

predict_probabilities(→ numpy.ndarray)

Run model over batches and return an (N, C) probability matrix.

probabilities_from_logits(→ numpy.ndarray)

Convert raw classifier-head outputs to an (N, C) probability matrix.

queue_rows(→ List[Tuple[str, Optional[int]]])

The queue as [(png_path, None), …].

rank_by_uncertainty(→ numpy.ndarray)

Row indices ordered most-uncertain-first.

record_labels(→ int)

Stamp each label with the round it was made in.

record_round(→ None)

Append (or replace) one row of the learning curve.

resolve_measure(→ Tuple[str, Callable[..., numpy.ndarray]])

Turn a measure name (or callable) into (name, function).

retrain_round(→ RoundResult)

Fit a model on the labels so far, score every crop, close the loop.

round_features(, nuclei_limit, pathogen_limit)

The measurement features for every crop, indexed by png_path.

should_stop(→ StoppingVerdict)

Has the last label_window labels moved held-out accuracy at all?

uncertainty_scores(→ numpy.ndarray)

Score every row with measure.

Module Contents

class spacr.active_learning.RoundResult(**fields: Any)[source]

What one retrain round produced.

Parameters:
  • fields – named values for the round fields below; omitted list and mapping fields are normalized to empty containers.

  • round_index – the round number recorded.

  • n_labels – labels the model was fitted on.

  • n_new_labels – labels added since the previous round.

  • report – the holdout_report() for this round.

  • split_rule – how the held-out set was drawn, in words.

  • scored – how many crops were re-scored in the database.

  • score_columns – the columns written back.

  • model_path – where the fitted model was saved, if it was.

  • card_path – the model card beside it, if one was written.

  • verdict – the StoppingVerdict after this round.

  • notes – anything the round wants the annotator to know.

Populate supported fields and normalize missing containers.

__repr__() → str[source]

Return the round, label count, and four-decimal held-out accuracy.

summary() → str[source]

One paragraph: the round, its numbers and what to do next.

property accuracy: float[source]

Held-out accuracy of this round.

property per_class: Dict[str, float][source]

{class name: held-out accuracy} for this round.

class spacr.active_learning.StoppingVerdict(stop: bool, reason: str, *, gain: float | None = None, labels_in_window: int = 0, window_from: int | None = None, confident: bool = False, noise: float | None = None, trend: str = 'unknown')[source]

Whether the last stretch of annotation bought anything measurable.

Parameters:
  • stop – the recommendation.

  • reason – one sentence, in the words the screen shows.

  • gain – held-out accuracy change over the window examined.

  • labels_in_window – how many labels that change is attributed to.

  • window_from – the round the window opened at.

  • confident – whether gain is larger than one standard error of the held-out accuracy itself. When it is not, “flat” and “we cannot tell” look identical from the numbers, and this says which you have.

  • noise – one standard error of the latest held-out accuracy, sqrt(p(1-p)/n).

  • trend – 'rising', 'flat', 'falling' or 'unknown'.

Normalize and store the recommendation and its measured evidence.

__bool__() → bool[source]

True when the recommendation is to stop.

__repr__() → str[source]

Return stop, trend, gain, and labels-in-window for diagnostics.

to_dict() → Dict[str, Any][source]

A JSON-friendly copy, for a card or a log.

spacr.active_learning.annotation_coverage(db_path: str, annotation_column: str = 'annotate', table: str = PNG_TABLE, key: str = PNG_KEY, image_type: str | None = None) → pandas.DataFrame[source]

Summarize annotation coverage by class, well, plate and acquisition.

The distribution of labels across experimental units determines whether a classifier can generalize beyond acquisition-specific staining, focus and confluency. A label set concentrated within one well can therefore yield optimistic performance under an object-level random split. This function exposes such concentration before model training or evaluation.

Reads png_list read-only, plus ROUND_TABLE when it is there, so labels can also be attributed to the active-learning round that surfaced them.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – the column the Annotate app writes into.

  • table – crop table (default png_list).

  • key – row key (default png_path).

  • image_type – substring filter on the key, matching the Annotate screen’s own filter. Every count below it, n_rows included, is over the crops that matched — a denominator taken from the whole table would put the numerator and the denominator on two different populations. n_rows_unfiltered keeps the total, and a note says how many were excluded.

Returns:

one row per (plateID, rowID, columnID, class) that has at least one annotation, with n and share — plus the whole breakdown in attrs['spacr_annotation_coverage']: by_class, by_plate, by_well, by_class_plate, by_class_well, by_round, by_source, concentration (per class) and notes.

Raises:
  • ValueError – when the table has no such column — an empty result would read as “nothing annotated yet”, which is a different fact.

  • FileNotFoundError – when the database is not there.

spacr.active_learning.as_probabilities(scores: Any) → numpy.ndarray[source]

Read stored scores as an (N, C) probability matrix.

Unlike probabilities_from_logits() this assumes the values are already probabilities — which is what the pred column of png_list holds, because spacr.deep_spacr.apply_model_to_tar() applied the sigmoid or softmax before writing it.

  • a single column is read as the positive-class probability of a binary problem and expanded to [1 − p, p];

  • (N, C) rows are renormalised if they do not sum to 1;

  • rows that cannot be a distribution (negatives, all-zero, values outside [0, 1] in the single-column case) become all-NaN, so they score NaN and get excluded rather than silently mis-ranked.

Parameters:

scores – array-like of stored probabilities.

Returns:

(N, C) float array.

spacr.active_learning.build_queue(db_path: str, annotation_column: str = 'annotate', pred_column: Any = None, table: str = PNG_TABLE, key: str = PNG_KEY, measure: Any = DEFAULT_MEASURE, limit: int | None = None, diversity: Any = 'well', group_columns: Sequence[str] | None = None, seed: int | None = None, image_type: str | None = None) → pandas.DataFrame[source]

Build the annotation queue: unlabelled crops, most uncertain first.

Reads png_list read-only and returns one row per crop still waiting for a label, ordered so the crops that would teach the model most come first.

Already-annotated crops are excluded. NULL in annotation_column means “not looked at”; anything else — including 0 — means a human committed to a class. If the column does not exist at all, nothing has been annotated yet and every crop is queued (with a note saying so). This is the same abstention convention as spacr.agreement.

The queue is diversified by default, and that costs something. With diversity='well' the ranked crops are dealt round-robin across wells: position 1 is still the single most uncertain crop, but position 2 is the most uncertain crop in a different well, which may be materially less uncertain than the runner-up overall. You give up some per-item uncertainty to stop the annotator labelling the same ambiguity a hundred times — the failure mode of pure uncertainty sampling, where the top 100 crops routinely come from two wells. If limit is smaller than the number of wells, the queue will contain roughly one crop from each of limit wells and none from the rest. Pass diversity='none' for the pure order, 'field'/'plate' for other strata, or group_columns=[…] for your own — including a cluster id you computed from features, which is the more thorough diversification this trades away for not needing a feature matrix.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – column the Annotate app writes into (default 'annotate').

  • pred_column – name, or list of names, of the model-score column(s). None (default) auto-detects: pred_0, pred_1, … style columns for multiclass, else pred. A single column is read as the positive-class probability of a binary problem.

  • table – table holding the crops (default png_list).

  • key – row key (default png_path).

  • measure – 'entropy' (default), 'least_confidence', 'margin', or a callable. On two classes the last two give the same order.

  • limit – keep at most this many crops.

  • diversity – 'well' (default), 'plate', 'row', 'column', 'field', or 'none'.

  • group_columns – explicit columns to stratify over, overriding diversity.

  • seed – tie-breaking seed; None breaks ties on row order. Either way the result is reproducible.

  • image_type – substring filter on png_path — matches the Annotate screen’s own image_type filter (e.g. 'cell').

Returns:

DataFrame with rank (1-based), the key, uncertainty, predicted_class, the probability columns and whatever crop metadata png_list has. Empty (with the same columns) when there is nothing to annotate. Diagnostics live in queue.attrs['spacr_active_learning']; render them with format_queue_summary().

Raises:
  • ValueError – when the table or the prediction column is missing — both mean there is no queue to build, and guessing would produce a plausible-looking wrong order.

  • FileNotFoundError – when the database is not there.

Warning

uncertainty is a ranking score, not a calibrated confidence. See CALIBRATION_NOTE.

spacr.active_learning.crops_for_object_keys(db_path: str, keys: Sequence[str], *, table: str = PNG_TABLE, key: str = PNG_KEY, annotation_column: str | None = None, timelapse: bool = False, image_type: str | None = None) → List[Tuple[str, int | None]][source]

Resolve object keys to crop rows, in the caller’s order.

The database half of the object-routing contract in spacr.qt.linked_selection: a scatter plot or a confusion-matrix cell names objects by spacr.selection.OBJECT_KEY_COLUMNS key, and the Annotate screen has to turn those into the crops it paginates. Kept here rather than in the Qt screen so it is testable without a display, and so a second consumer does not have to reimplement the 'o5'-versus-5 trap in _object_label().

Input order is preserved so priority rankings such as worst errors first remain unchanged. Keys without a corresponding crop are omitted.

Typed keys distinguish objects with the same numeric label in one field. png_list identifies the object type through the populated <type>_id column (PNG_ID_COLUMN_TYPES), so nucleus 1 and pathogen 1 resolve independently. An untyped key selects the first matching crop. If the table does not expose object-type columns, typed lookup falls back to the corresponding untyped key.

Escaped metadata components are resolved in both encoded and raw form. For example, spacr.selection.object_keys() represents a fieldID of 'f_1' as 'f%5F1', while a key assembled from crop-table columns contains the raw underscore.

Parameters:
  • db_path – path to measurements.db.

  • keys – object keys, typed or not. A png_path, a prcfo or a file_name is also accepted, so a caller working from a crop table rather than a measurement table does not need a translation step.

  • table – crop table.

  • key – crop key column.

  • annotation_column – read the existing label too, so an already annotated crop renders with its colour rather than blank.

  • timelapse – the keys carry a timepoint.

  • image_type – substring filter on the crop key.

Returns:

[(png_path, annotation or None), …] in the keys’ order.

spacr.active_learning.disagreement(prob_sets: Any, method: str = 'variance') → numpy.ndarray[source]

How much several score sets disagree about each crop.

The measures above see one model’s opinion and call a 50/50 output “uncertain” whether the model is genuinely torn or merely mis-calibrated. An ensemble — several checkpoints, several folds, or several MC-dropout passes of one model — separates those: crops the members disagree about are where the model class itself is undecided (epistemic uncertainty), which is what a new label actually fixes.

Parameters:
  • prob_sets – sequence of M score arrays, each (N,) or (N, C) over the SAME N crops in the same order (a 3-D (M, N, C) array works too). Each member is coerced with as_probabilities().

  • method –

    • 'variance' (default) — mean across classes of the across-member variance (population variance, ddof=0). 0 when the members agree exactly.

    • 'bald' — mutual information H(mean p) − mean H(p) (Houlsby et al., 2011). 0 when the members agree, regardless of how uncertain they jointly are — so unlike entropy it does not fire on crops that are ambiguous.

Returns:

(N,) scores; NaN where any member is unusable for that crop.

Raises:

ValueError – for fewer than one set, or sets of different shapes — a length mismatch means the members are not aligned to the same crops, and averaging them would be meaningless.

A single set returns all zeros: one opinion cannot disagree with itself. That is a real answer, not a failure, but it means the queue would be in row order, so check the count before using it.

spacr.active_learning.ensure_round_tables(db_path: str) → None[source]

Create ROUND_TABLE and ROUND_LOG_TABLE if absent.

Parameters:

db_path – existing measurements.db in which to create the tables.

Two tables, not one. Per-label provenance and per-round metrics have different cardinalities and different lifetimes: a label keeps its round forever, a round’s held-out accuracy is rewritten if the round is re-fit.

spacr.active_learning.entropy(probs: Any, base: float | None = None, normalize: bool = False) → numpy.ndarray[source]

Shannon entropy −Σ p log p of each row.

The only measure here that uses the whole distribution. Minimum 0 at a one-hot row; maximum log C at the uniform row (log 2 ≈ 0.6931 for two classes in nats). 0 · log 0 is taken as 0.

Parameters:
  • probs – (N,) or (N, C); coerced with as_probabilities().

  • base – logarithm base. None (default) means natural log, so the units are nats; pass 2 for bits.

  • normalize – divide by log C so the maximum is 1. Monotone, so the ranking is unchanged.

Returns:

(N,) uncertainty scores; NaN for unusable rows.

spacr.active_learning.format_coverage_summary(coverage: pandas.DataFrame) → str[source]

Render annotation_coverage() as text, worst concentration first.

Parameters:

coverage – the frame from annotation_coverage().

Returns:

multi-line text, no trailing newline.

spacr.active_learning.format_learning_curve(curve: pandas.DataFrame, verdict: StoppingVerdict | None = None) → str[source]

Render the round-by-round curve and the stopping verdict as text.

Parameters:

curve – round-by-round metrics frame from learning_curve().

spacr.active_learning.format_queue_summary(queue: pandas.DataFrame) → str[source]

Render a queue’s shape, class balance and caveats as plain text.

Reports the numbers that decide whether the queue is worth working through: how much of the screen is already labelled, how many crops could not be scored, the range of the scores, how many wells the queue actually spreads over (the diversity check), and the class balance of the queue next to the balance of the pool it came from — because uncertainty sampling on an imbalanced screen pulls hard toward the majority class’s boundary and the annotator should see that rather than discover it.

Parameters:

queue – frame from build_queue(). Works on a slice or a copy too, falling back to what can be recomputed from the rows when attrs did not survive.

Returns:

multi-line text, no trailing newline.

spacr.active_learning.holdout_report(y_true: Any, probs: Any, classes: Sequence[Any] | None = None) → Dict[str, Any][source]

Held-out metrics, with the confusion matrix they were derived from.

Torch-free, so both the classical-ML round here and spacr.deep_spacr.held_out_report() can use one implementation and a model card written by either says the same thing in the same shape.

Every derived figure is exactly the standard function of the matrix — accuracy = trace / total, per_class[c] = M[c, c] / M[c, :].sum() — so a reader can recompute the card rather than trust it.

n is the number of rows the matrix actually contains, and the supports sum to it. A row whose true class the head has no column for is counted as an error rather than dropped: the matrix grows to hold it, and the missing column stays empty because the head can never predict that class. Reporting n over one population and accuracy over another is the one thing a model card must not do — a three-class held-out set scored by a binary head used to report a perfect score on a set the model got two of three right, and the accuracy == trace / total invariant still checked out because the row was missing from both sides of it.

Parameters:
  • y_true – integer class ids, shape (N,). Negative ids are not classes; they are excluded, counted in n_unscored and explained in notes rather than silently folded into the total.

  • probs – (N,) positive-class probabilities, or (N, C) rows.

  • classes – class names in head order.

Returns:

n, n_unscored, num_classes, head_classes, classes, accuracy, f1_macro, per_class_accuracy, class_support, predicted_support, confusion_matrix, notes.

Raises:

ValueError – when there are not as many score rows as labels — the two are not aligned, and every figure below would be a comparison of one object’s label with another’s prediction.

spacr.active_learning.label_rounds(db_path: str, annotation_column: str = 'annotate') → pandas.DataFrame[source]

Per-label round provenance as a frame (empty when never recorded).

Parameters:

db_path – path to the measurements.db to read.

spacr.active_learning.learning_curve(db_path: str, annotation_column: str = 'annotate') → pandas.DataFrame[source]

Held-out accuracy per round, oldest first — the curve to watch flatten.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – the column the rounds labelled into.

Returns:

a frame with one row per round: round, finished_utc, n_labels, n_new_labels, n_holdout, holdout_accuracy, holdout_f1_macro, per_class (dict), split_rule, model_type, model_path, card_path, notes (list), and the derived gain (accuracy change since the previous round). Empty (with those columns) when no round has been recorded.

spacr.active_learning.least_confidence(probs: Any, normalize: bool = False) → numpy.ndarray[source]

1 − max_c p_c — how much probability mass is not on the winner.

Minimum 0 at a one-hot row; maximum 1 − 1/C at the uniform row. It looks only at the top class, so on 3+ classes it cannot tell [0.5, 0.5, 0.0] from [0.5, 0.25, 0.25]; entropy() can.

Parameters:
  • probs – (N,) positive-class probabilities or (N, C) rows. Coerced with as_probabilities().

  • normalize – rescale by C / (C − 1) so the maximum is 1. A monotone rescaling — it never changes the ranking, only how the number reads.

Returns:

(N,) uncertainty scores; NaN for unusable rows.

spacr.active_learning.margin(probs: Any) → numpy.ndarray[source]

1 − (p₁ − p₂) — closeness of the top two classes, as uncertainty.

This returns an uncertainty score, not the margin: a small margin (the two leading classes neck and neck) is a large return value, so it is oriented like every other measure here. Minimum 0 at a one-hot row, maximum 1 at any row whose top two classes tie.

On two classes this is a linear transform of least_confidence() — 1 − (p₁ − p₂) = 2·(1 − max p) — so the two produce the identical order and the identical ties. They are one choice, not two, until C ≥ 3, where margin ignores everything below the runner-up and least-confidence ignores everything below the winner.

Parameters:

probs – (N,) or (N, C); coerced with as_probabilities().

Returns:

(N,) uncertainty scores in [0, 1]; NaN for unusable rows.

spacr.active_learning.next_round(db_path: str, annotation_column: str = 'annotate') → int[source]

The round number the next batch of labels belongs to.

Parameters:

db_path – path to the measurements.db whose round log is queried.

Round 0 is “before any model was retrained from inside Annotate” — the labels that seeded the loop. The first retrain produces round 1.

spacr.active_learning.predict_probabilities(model: Callable[[Any], Any], batches: Iterable[Any], device: Any = None, from_logits: bool = True) → numpy.ndarray[source]

Run model over batches and return an (N, C) probability matrix.

A convenience for scoring crops that are not in the database yet. The queue does not need this — build_queue() works from the pred column that spacr.deep_spacr.merge_predictions_into_db() already wrote, which is the normal path.

torch is imported inside this function, and only to get no_grad/device handling; if it is not importable the batches are iterated and the model is called directly. Nothing else in this module touches torch.

Parameters:
  • model – any callable mapping a batch to raw head outputs. A torch.nn.Module is put in eval() mode first if it has one.

  • batches – iterable of batches. A batch that is a (inputs, …) tuple has its first element passed to the model, matching the loaders in spacr.deep_spacr.

  • device – optional torch device to move inputs/model to.

  • from_logits – treat outputs as raw logits and apply probabilities_from_logits() (the default — a model head emits logits). Set False if the model already outputs probabilities, which then go through as_probabilities().

Returns:

(N, C) probability matrix in batch order.

spacr.active_learning.probabilities_from_logits(logits: Any) → numpy.ndarray[source]

Convert raw classifier-head outputs to an (N, C) probability matrix.

The head shape decides the link function, and getting this wrong is silent:

  • (N,) or (N, 1) — a single-logit binary head. Sigmoid, then expanded to [1 − p, p] so every measure sees two classes. Pushing this through a softmax instead would return a column of 1.0 — every crop maximally certain, the whole ordering gone.

  • (N, C) with C ≥ 2 — a C-logit head. Row-wise softmax. Sigmoiding these instead reads each logit in isolation and can invert the order: [5, 5] is a perfect 50/50 tie under softmax but looks like a confident 0.993 under a sigmoid of column 1.

Both branches are exactly what spacr.deep_spacr.apply_model_to_tar() and spacr.deep_spacr.evaluate_model_performance() do at inference time.

Parameters:

logits – array-like or torch tensor of raw head outputs.

Returns:

(N, C) float array whose rows sum to 1 (C ≥ 2).

Raises:

ValueError – for a scalar or a 3-D input.

Note

The output is a probability vector, not a calibrated probability. See CALIBRATION_NOTE.

spacr.active_learning.queue_rows(queue: pandas.DataFrame, key: str = PNG_KEY) → List[Tuple[str, int | None]][source]

The queue as [(png_path, None), …].

Exactly the shape spacr.qt.annotate_engine.fetch_page() and spacr.qt.annotate_engine.fetch_filtered_paths() return, so the Annotate screen can page through a queue with no other change. The annotation is always None — every crop in a queue is unlabelled by construction.

Parameters:
  • queue – frame from build_queue().

  • key – the path column (default png_path).

Returns:

list of (path, None) tuples in queue order.

spacr.active_learning.rank_by_uncertainty(probs: Any, measure: Any = DEFAULT_MEASURE, limit: int | None = None, seed: int | None = None, scores: Any | None = None) → numpy.ndarray[source]

Row indices ordered most-uncertain-first.

Ordering is total and reproducible:

  • primary key — the uncertainty score, descending;

  • ties — row index ascending by default, or a permutation seeded with seed when one is given. Both are deterministic: the same inputs and the same seed always give the same order. Ties are the normal case on a screen where thousands of crops score exactly 0.5, so an unseeded shuffle there would reshuffle the annotator’s queue on every refresh.

  • NaN scores sort last, always, and never in front of a real one. They are ranked rather than dropped so the returned indices stay a permutation of range(N); build_queue() drops them and says how many.

Parameters:
  • probs – (N,) or (N, C) probabilities.

  • measure – name from UNCERTAINTY_MEASURES or a callable.

  • limit – keep only the first limit indices.

  • seed – seed for tie-breaking; None breaks ties on index.

  • scores – pre-computed scores to rank instead of recomputing from probs (used by build_queue(), and by anything ranking a disagreement() score).

Returns:

(N,) (or (limit,)) int array of row indices.

spacr.active_learning.record_labels(db_path: str, annotation_column: str, labels: Dict[str, Any], round_index: int, source: str = 'manual') → int[source]

Stamp each label with the round it was made in.

Called by the Annotate screen every time it flushes a batch. The round a label came from is what makes early-round bias auditable: the first round’s labels are drawn from whatever ordering existed before any model had seen this screen, and if 80 % of a class’s labels carry round 0 then the “active learning” was mostly not active.

first_round is preserved across re-labelling while round follows the current value, so both “when was this crop first looked at” and “which round set the label it has now” survive a correction.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – the column the labels were written into.

  • labels – {png_path: class or None}. None is a cleared label and is recorded as such rather than dropped — a crop that was looked at and deliberately left blank is not the same as one never seen.

  • round_index – the round these labels belong to.

  • source – how the crop reached the annotator — 'manual', 'queue', or a caller’s own tag.

Returns:

number of rows written.

spacr.active_learning.record_round(db_path: str, annotation_column: str, round_index: int, **fields: Any) → None[source]

Append (or replace) one row of the learning curve.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – the column this round labelled into.

  • round_index – the round number.

  • fields – any of n_labels, n_new_labels, n_holdout, holdout_accuracy, holdout_f1_macro, per_class (dict), split_rule, model_type, model_path, card_path, measure, diversity, notes (list).

spacr.active_learning.resolve_measure(measure: Any) → Tuple[str, Callable[..., numpy.ndarray]][source]

Turn a measure name (or callable) into (name, function).

Parameters:

measure – a key of UNCERTAINTY_MEASURES, or any callable f(probs) -> (N,).

Raises:

ValueError – for an unknown name, listing the valid ones.

spacr.active_learning.retrain_round(db_path: str, annotation_column: str = 'annotate', *, features: pandas.DataFrame | None = None, model_type: str = 'logistic_regression', group_by: str = 'well', holdout: float = 0.25, seed: int = 0, min_labels: int = 8, round_index: int | None = None, table: str = PNG_TABLE, key: str = PNG_KEY, image_type: str | None = None, write_scores: bool = True, save_model: bool = True, model_dir: str | None = None, write_card: bool = True, label_window: int = 50, min_gain: float = 0.003, measure: Any = DEFAULT_MEASURE, diversity: Any = 'well', balance: str = 'none', synthetic_negatives: int | None = None, rejections: Mapping[Any, Any] | None = None) → RoundResult[source]

Fit a model on the labels so far, score every crop, close the loop.

This is the half of active learning that has been missing: the queue put the informative crops in front of the annotator, and then nothing happened. Annotating without retraining is not active learning — it is ordinary annotation in a clever order, and the order goes stale after the first few dozen labels because it still reflects a model that has not seen any of them.

One call does all five things the loop needs:

  1. fits a model on every label in annotation_column;

  2. scores it on a grouped held-out split, so the number is not an artefact of 190 labels coming from one well;

  3. writes per-class probabilities back into png_list as ROUND_PRED_PREFIX columns, which build_queue() prefers over the older pred — so the next queue is genuinely re-ranked;

  4. records the round, giving learning_curve() another point;

  5. returns the StoppingVerdict for the curve so far.

Parameters:
  • db_path – path to measurements.db.

  • annotation_column – the column holding the labels.

  • features – feature matrix indexed by the crop key. Omitted, it is read from the measurement tables with round_features().

  • model_type – 'logistic_regression' (default — it is the one that behaves at 20 labels), 'random_forest' or 'gradient_boosting'.

  • group_by – 'well' (default), 'plate', 'field' or 'none'. What the held-out split refuses to share. Matched exactly and in lower case; anything else raises, the way an unknown diversity= does in build_queue(). When the strategy is 'none', or the crop table has none of the columns it needs, the split is a stratified random one and split_rule says NOT grouped — it never claims a grouping it did not perform.

  • holdout – fraction held out.

  • seed – makes the split and the fit reproducible.

  • min_labels – refuse to fit below this many labels.

  • round_index – override the round number; defaults to next_round().

  • table – crop table.

  • key – crop key column.

  • image_type – substring filter on the crop key.

  • write_scores – write the new probabilities back into the database.

  • save_model – joblib-dump the fitted model beside the database.

  • model_dir – where to put it; defaults to <db dir>/active_learning.

  • write_card – write a model card beside the saved model.

  • label_window – passed to should_stop().

  • min_gain – passed to should_stop().

  • measure – recorded with the round, for the queue that follows.

  • diversity – likewise.

  • balance –

    'none' (default) leaves imbalance to the estimator’s own class_weight='balanced'; 'downsample' cuts every class to the size of the smallest BEFORE the grouped split.

    THE TWO ARE NOT THE SAME ANSWER. Reweighting and downsampling produce different probabilities from the same crops, and spacr.suggest.suggest_from_scores() sorts on those probabilities – so the round records which was in force, in notes and on the model card. The smaller class (“if there is class imbalance use the class with fewer”).

  • rejections – {crop key: rejected class} – suggestions the annotator REJECTED (spacr.suggest.rejected_suggestions()). In a two-class column a rejection of class 1 is an example of class 2, and it is fitted as one; a crop that has since been labelled is left to its label, and in a column with any other classes the rejection cannot be turned into a label and is counted in the notes instead. A rejection is information, not silence: without this the model that proposed the wrong class would be fitted on exactly the same evidence next round and propose it again.

  • synthetic_negatives –

    how many unannotated crops to draw at random and fit as the ABSENT class when only one class has been annotated. None (default) refuses instead, as before.

    THIS IS A DELIBERATE LIE AND THE ROUND SAYS SO. A random draw from the unannotated pool is mostly-negative, not negative, so what comes back is a ranking rather than a verdict; the count reaches the model card, because a card that does not say the negatives were invented describes a model that does not exist. Defined for the binary classes 1 and 2 only – see _absent_binary_class().

Returns:

a RoundResult.

Raises:

ValueError – below min_labels labels, with fewer than two classes annotated — neither is something to paper over with a model that will produce a confident-looking ranking out of nothing — or for an unrecognised group_by.

spacr.active_learning.round_features(db_path: str, table: str = PNG_TABLE, key: str = PNG_KEY, tables: Sequence[str] = ('cell', 'nucleus', 'pathogen', 'cytoplasm'), nuclei_limit: int = 10, pathogen_limit: int = 10) → pandas.DataFrame[source]

The measurement features for every crop, indexed by png_path.

The feature matrix an in-screen retrain fits on. Measurement features rather than pixels on purpose: the point of retraining from inside Annotate is to get a fresh ranking in seconds, on the machine the annotator is sitting at, without a GPU and without leaving the screen. A CNN retrain is the right thing to do at the end of the loop, not between two pages of crops.

Parameters:
  • db_path – path to measurements.db.

  • table – crop table carrying png_path and prcfo.

  • key – the crop key column.

  • tables – object tables to merge features from; missing ones are skipped.

  • nuclei_limit – passed through to spacr.io._read_and_merge_data().

  • pathogen_limit – likewise.

Returns:

numeric features indexed by png_path.

Raises:

ValueError – when no object table with features could be read.

spacr.active_learning.should_stop(curve: pandas.DataFrame, *, label_window: int = 50, min_gain: float = 0.003, min_rounds: int = 2) → StoppingVerdict[source]

Has the last label_window labels moved held-out accuracy at all?

The rule, in one line: look back over whole rounds until at least label_window new labels have accumulated, and compare held-out accuracy at the two ends. If it moved by less than min_gain, stop.

Why this rule and not another:

  • Labels, not rounds, are the unit of cost. A round is whatever size the annotator felt like; “no improvement for 3 rounds” says nothing when the rounds were 5, 300 and 8 labels. The thing being spent is human attention, one crop at a time, so the window is measured in crops.

  • It looks *back over* rounds, not *at* the last one. A single round that happened to land flat is noise; the question is whether the last fifty labels — however they were divided up — bought anything.

  • It refuses to answer early. Below label_window labels since the first recorded round there is no window to measure, and a rule that fired anyway would tell people to stop after their first twelve labels.

  • It distinguishes “flat” from “unmeasurable”. Held-out accuracy on 80 objects has a standard error near 0.05; a 0.3 % change is inside the noise, and StoppingVerdict.confident says so instead of dressing it up. Flat is still the recommendation — if more labels are not moving a number you can measure, they are not buying anything you can demonstrate — but the reason says the held-out set is too small to prove convergence, which is a different piece of work.

  • A falling curve stops too, and says so. Accuracy going down is not convergence; it usually means the newest labels disagree with the earlier ones, or that the held-out split moved. Either way, more of the same is the wrong next move.

Parameters:
  • curve – the frame from learning_curve().

  • label_window – how many labels the window must cover.

  • min_gain – accuracy change below which the window counts as flat.

  • min_rounds – rounds required before any verdict is given.

Returns:

a StoppingVerdict; bool(verdict) is the answer.

spacr.active_learning.uncertainty_scores(probs: Any, measure: Any = DEFAULT_MEASURE) → numpy.ndarray[source]

Score every row with measure.

Parameters:
  • probs – (N,) or (N, C) probabilities.

  • measure – name from UNCERTAINTY_MEASURES or a callable.

Returns:

(N,) uncertainty scores, larger = less certain.

Nested helpers

_segmentation_uncertainty._paint(values: np.ndarray) → np.ndarray
Parameters:

values – one value per reference object, in id order.

Returns:

each object’s pixels set to its value, 0 elsewhere.

spacr/active_learning.py:3910

crops_for_object_keys._register(target: Dict[str, Tuple[str, int | None]], composed: List[str], label: str, object_type: str | None, entry: Tuple[str, int | None]) → None

Register first-wins untyped and, when known, typed object keys.

spacr/active_learning.py:1637

crops_for_object_keys._resolve(name: str) → Tuple[str, int | None] | None

The escaped spelling first — it is the one a producer emits today.

spacr/active_learning.py:1700

predict_probabilities._run() → None

Score captured batches in order and append normalized matrices.

Loader tuples contribute their input element, and movable inputs are transferred to the selected device before the model is called.

spacr/active_learning.py:700