spacr.active_learning¶
Active-learning queue — order unannotated crops by model uncertainty.
Annotation is the bottleneck in every screen. A crop the model already
calls with 0.999 teaches the model nothing when a human labels it; the
crops worth a person’s afternoon are the ones sitting on the decision
boundary. This module turns an already-scored measurements.db into a
work queue ordered so the informative crops come first.
Public API¶
least_confidence(probs)/margin(probs)/entropy(probs)Per-row uncertainty scores. Pure numpy — no torch, no database.
disagreement(prob_sets)Spread across an ensemble or a set of MC-dropout passes.
rank_by_uncertainty(probs, measure=…)Row indices, most uncertain first, deterministically tie-broken.
build_queue(db_path, annotation_column, …)The queue itself, read straight out of
png_list.queue_rows(queue)The queue as
[(png_path, None), …]— the shape the Annotate screen already paginates (seespacr.qt.annotate_engine.fetch_page()).format_queue_summary(queue)The queue’s shape, class balance and caveats, as text.
probabilities_from_logits(logits)/as_probabilities(scores)Raw-head-output → probability matrix, both head shapes handled.
predict_probabilities(model, batches)Optional live-model bridge; the only function that touches torch.
Things this module refuses to get wrong¶
A softmax is not a probability. Everything here is called an
uncertainty score, never a confidence. Modern networks are badly
calibrated — typically over-confident, and more so the deeper they are
(Guo et al., 2017) — so a 0.87 from the head is not “87 % sure”. The
scores are used for one thing only: putting crops in an order. Nothing
in this module reports a calibrated probability, and neither should
anything built on it. See CALIBRATION_NOTE.
Already-annotated crops never enter the queue. NULL in the
annotation column is the abstention marker — the crop has not been
looked at. 0 is a real class, and a queue that re-serves it wastes
exactly the resource this module exists to save. The two are
distinguished by IS NULL, never by falsiness. (Same convention as
spacr.agreement.)
The two head shapes are handled separately. The classifier head
emits either a single logit (binary; needs a sigmoid) or C logits
(multiclass; needs a softmax) — see
spacr.deep_spacr.apply_model_to_tar(). Pushing a single-logit
column through a softmax yields a column of 1.0 and destroys the
ordering; pushing C logits through a sigmoid can invert it. Both are
silent failures that produce a confident-looking queue full of nonsense,
so probabilities_from_logits() branches on the shape and the tests
pin both directions.
margin and least_confidence are the same ranking on two
classes. For C = 2, 1 − (p₁ − p₂) = 2·min(p, 1−p) = 2·(1 − max p):
a linear transform, so identical order and identical ties. They are not
two independent choices on a binary screen — they differ only from three
classes up, where margin looks at the top two classes while
least_confidence looks only at the top one and entropy looks at
the whole distribution.
Pure uncertainty ranking collapses onto one region of feature space.
The 100 most uncertain crops on a real plate are routinely 100 near
copies from the same two wells — one ambiguity, labelled a hundred
times, for almost the information of labelling it once. So the queue is
diversified by default (round-robin across wells) and
format_queue_summary() prints how many wells the queue actually
covers. What that costs is stated plainly in build_queue().
Uncertainty sampling skews toward the majority class. On a screen that is 98 % negative, the decision boundary is mostly populated by negatives, so the queue will be too. That is not a bug to hide — it is reported as a class-balance table so the annotator can see it and cap or rebalance if they want to.
Determinism. Same inputs, same seed, same order — including ties.
Ties break on row index by default; seed swaps in a seeded
permutation, which is still reproducible. Nothing here consults an
unseeded RNG.
Nothing in this module imports torch at module scope; the ranking maths
is numpy only, so the Qt screen can build a queue without waking a
multi-second import chain. predict_probabilities() imports torch
lazily and is the only entry point that needs it.
Classes¶
What one retrain round produced. |
|
Whether the last stretch of annotation bought anything measurable. |
Functions¶
|
Summarize annotation coverage by class, well, plate and acquisition. |
|
Read stored scores as an |
|
Build the annotation queue: unlabelled crops, most uncertain first. |
|
Resolve object keys to crop rows, in the caller's order. |
|
How much several score sets disagree about each crop. |
|
Create |
|
Shannon entropy |
|
Render |
|
Render the round-by-round curve and the stopping verdict as text. |
|
Render a queue's shape, class balance and caveats as plain text. |
|
Held-out metrics, with the confusion matrix they were derived from. |
|
Per-label round provenance as a frame (empty when never recorded). |
|
Held-out accuracy per round, oldest first — the curve to watch flatten. |
|
|
|
|
|
The round number the next batch of labels belongs to. |
|
Run |
|
Convert raw classifier-head outputs to an |
|
The queue as |
|
Row indices ordered most-uncertain-first. |
|
Stamp each label with the round it was made in. |
|
Append (or replace) one row of the learning curve. |
|
Turn a measure name (or callable) into |
|
Fit a model on the labels so far, score every crop, close the loop. |
|
The measurement features for every crop, indexed by |
|
Has the last |
|
Score every row with |
Module Contents¶
- class spacr.active_learning.RoundResult(**fields: Any)[source]¶
What one retrain round produced.
- Parameters:
fields – named values for the round fields below; omitted list and mapping fields are normalized to empty containers.
round_index – the round number recorded.
n_labels – labels the model was fitted on.
n_new_labels – labels added since the previous round.
report – the
holdout_report()for this round.split_rule – how the held-out set was drawn, in words.
scored – how many crops were re-scored in the database.
score_columns – the columns written back.
model_path – where the fitted model was saved, if it was.
card_path – the model card beside it, if one was written.
verdict – the
StoppingVerdictafter this round.notes – anything the round wants the annotator to know.
Populate supported fields and normalize missing containers.
- class spacr.active_learning.StoppingVerdict(stop: bool, reason: str, *, gain: float | None = None, labels_in_window: int = 0, window_from: int | None = None, confident: bool = False, noise: float | None = None, trend: str = 'unknown')[source]¶
Whether the last stretch of annotation bought anything measurable.
- Parameters:
stop – the recommendation.
reason – one sentence, in the words the screen shows.
gain – held-out accuracy change over the window examined.
labels_in_window – how many labels that change is attributed to.
window_from – the round the window opened at.
confident – whether
gainis larger than one standard error of the held-out accuracy itself. When it is not, “flat” and “we cannot tell” look identical from the numbers, and this says which you have.noise – one standard error of the latest held-out accuracy,
sqrt(p(1-p)/n).trend –
'rising','flat','falling'or'unknown'.
Normalize and store the recommendation and its measured evidence.
- spacr.active_learning.annotation_coverage(db_path: str, annotation_column: str = 'annotate', table: str = PNG_TABLE, key: str = PNG_KEY, image_type: str | None = None) pandas.DataFrame[source]¶
Summarize annotation coverage by class, well, plate and acquisition.
The distribution of labels across experimental units determines whether a classifier can generalize beyond acquisition-specific staining, focus and confluency. A label set concentrated within one well can therefore yield optimistic performance under an object-level random split. This function exposes such concentration before model training or evaluation.
Reads
png_listread-only, plusROUND_TABLEwhen it is there, so labels can also be attributed to the active-learning round that surfaced them.- Parameters:
db_path – path to
measurements.db.annotation_column – the column the Annotate app writes into.
table – crop table (default
png_list).key – row key (default
png_path).image_type – substring filter on the key, matching the Annotate screen’s own filter. Every count below it,
n_rowsincluded, is over the crops that matched — a denominator taken from the whole table would put the numerator and the denominator on two different populations.n_rows_unfilteredkeeps the total, and a note says how many were excluded.
- Returns:
one row per
(plateID, rowID, columnID, class)that has at least one annotation, withnandshare— plus the whole breakdown inattrs['spacr_annotation_coverage']:by_class,by_plate,by_well,by_class_plate,by_class_well,by_round,by_source,concentration(per class) andnotes.- Raises:
ValueError – when the table has no such column — an empty result would read as “nothing annotated yet”, which is a different fact.
FileNotFoundError – when the database is not there.
- spacr.active_learning.as_probabilities(scores: Any) numpy.ndarray[source]¶
Read stored scores as an
(N, C)probability matrix.Unlike
probabilities_from_logits()this assumes the values are already probabilities — which is what thepredcolumn ofpng_listholds, becausespacr.deep_spacr.apply_model_to_tar()applied the sigmoid or softmax before writing it.a single column is read as the positive-class probability of a binary problem and expanded to
[1 − p, p];(N, C)rows are renormalised if they do not sum to 1;rows that cannot be a distribution (negatives, all-zero, values outside [0, 1] in the single-column case) become all-NaN, so they score NaN and get excluded rather than silently mis-ranked.
- Parameters:
scores – array-like of stored probabilities.
- Returns:
(N, C)float array.
- spacr.active_learning.build_queue(db_path: str, annotation_column: str = 'annotate', pred_column: Any = None, table: str = PNG_TABLE, key: str = PNG_KEY, measure: Any = DEFAULT_MEASURE, limit: int | None = None, diversity: Any = 'well', group_columns: Sequence[str] | None = None, seed: int | None = None, image_type: str | None = None) pandas.DataFrame[source]¶
Build the annotation queue: unlabelled crops, most uncertain first.
Reads
png_listread-only and returns one row per crop still waiting for a label, ordered so the crops that would teach the model most come first.Already-annotated crops are excluded.
NULLinannotation_columnmeans “not looked at”; anything else — including 0 — means a human committed to a class. If the column does not exist at all, nothing has been annotated yet and every crop is queued (with a note saying so). This is the same abstention convention asspacr.agreement.The queue is diversified by default, and that costs something. With
diversity='well'the ranked crops are dealt round-robin across wells: position 1 is still the single most uncertain crop, but position 2 is the most uncertain crop in a different well, which may be materially less uncertain than the runner-up overall. You give up some per-item uncertainty to stop the annotator labelling the same ambiguity a hundred times — the failure mode of pure uncertainty sampling, where the top 100 crops routinely come from two wells. Iflimitis smaller than the number of wells, the queue will contain roughly one crop from each oflimitwells and none from the rest. Passdiversity='none'for the pure order,'field'/'plate'for other strata, orgroup_columns=[…]for your own — including a cluster id you computed from features, which is the more thorough diversification this trades away for not needing a feature matrix.- Parameters:
db_path – path to
measurements.db.annotation_column – column the Annotate app writes into (default
'annotate').pred_column – name, or list of names, of the model-score column(s).
None(default) auto-detects:pred_0, pred_1, …style columns for multiclass, elsepred. A single column is read as the positive-class probability of a binary problem.table – table holding the crops (default
png_list).key – row key (default
png_path).measure –
'entropy'(default),'least_confidence','margin', or a callable. On two classes the last two give the same order.limit – keep at most this many crops.
diversity –
'well'(default),'plate','row','column','field', or'none'.group_columns – explicit columns to stratify over, overriding
diversity.seed – tie-breaking seed;
Nonebreaks ties on row order. Either way the result is reproducible.image_type – substring filter on
png_path— matches the Annotate screen’s ownimage_typefilter (e.g.'cell').
- Returns:
DataFrame with
rank(1-based), the key,uncertainty,predicted_class, the probability columns and whatever crop metadatapng_listhas. Empty (with the same columns) when there is nothing to annotate. Diagnostics live inqueue.attrs['spacr_active_learning']; render them withformat_queue_summary().- Raises:
ValueError – when the table or the prediction column is missing — both mean there is no queue to build, and guessing would produce a plausible-looking wrong order.
FileNotFoundError – when the database is not there.
Warning
uncertaintyis a ranking score, not a calibrated confidence. SeeCALIBRATION_NOTE.
- spacr.active_learning.crops_for_object_keys(db_path: str, keys: Sequence[str], *, table: str = PNG_TABLE, key: str = PNG_KEY, annotation_column: str | None = None, timelapse: bool = False, image_type: str | None = None) List[Tuple[str, int | None]][source]¶
Resolve object keys to crop rows, in the caller’s order.
The database half of the object-routing contract in
spacr.qt.linked_selection: a scatter plot or a confusion-matrix cell names objects byspacr.selection.OBJECT_KEY_COLUMNSkey, and the Annotate screen has to turn those into the crops it paginates. Kept here rather than in the Qt screen so it is testable without a display, and so a second consumer does not have to reimplement the'o5'-versus-5trap in_object_label().Input order is preserved so priority rankings such as
worst errors firstremain unchanged. Keys without a corresponding crop are omitted.Typed keys distinguish objects with the same numeric label in one field.
png_listidentifies the object type through the populated<type>_idcolumn (PNG_ID_COLUMN_TYPES), so nucleus 1 and pathogen 1 resolve independently. An untyped key selects the first matching crop. If the table does not expose object-type columns, typed lookup falls back to the corresponding untyped key.Escaped metadata components are resolved in both encoded and raw form. For example,
spacr.selection.object_keys()represents afieldIDof'f_1'as'f%5F1', while a key assembled from crop-table columns contains the raw underscore.- Parameters:
db_path – path to
measurements.db.keys – object keys, typed or not. A
png_path, aprcfoor afile_nameis also accepted, so a caller working from a crop table rather than a measurement table does not need a translation step.table – crop table.
key – crop key column.
annotation_column – read the existing label too, so an already annotated crop renders with its colour rather than blank.
timelapse – the keys carry a timepoint.
image_type – substring filter on the crop key.
- Returns:
[(png_path, annotation or None), …]in the keys’ order.
- spacr.active_learning.disagreement(prob_sets: Any, method: str = 'variance') numpy.ndarray[source]¶
How much several score sets disagree about each crop.
The measures above see one model’s opinion and call a 50/50 output “uncertain” whether the model is genuinely torn or merely mis-calibrated. An ensemble — several checkpoints, several folds, or several MC-dropout passes of one model — separates those: crops the members disagree about are where the model class itself is undecided (epistemic uncertainty), which is what a new label actually fixes.
- Parameters:
prob_sets – sequence of M score arrays, each
(N,)or(N, C)over the SAME N crops in the same order (a 3-D(M, N, C)array works too). Each member is coerced withas_probabilities().method –
'variance'(default) — mean across classes of the across-member variance (population variance, ddof=0). 0 when the members agree exactly.'bald'— mutual informationH(mean p) − mean H(p)(Houlsby et al., 2011). 0 when the members agree, regardless of how uncertain they jointly are — so unlikeentropyit does not fire on crops that are ambiguous.
- Returns:
(N,)scores; NaN where any member is unusable for that crop.- Raises:
ValueError – for fewer than one set, or sets of different shapes — a length mismatch means the members are not aligned to the same crops, and averaging them would be meaningless.
A single set returns all zeros: one opinion cannot disagree with itself. That is a real answer, not a failure, but it means the queue would be in row order, so check the count before using it.
- spacr.active_learning.ensure_round_tables(db_path: str) None[source]¶
Create
ROUND_TABLEandROUND_LOG_TABLEif absent.- Parameters:
db_path – existing
measurements.dbin which to create the tables.
Two tables, not one. Per-label provenance and per-round metrics have different cardinalities and different lifetimes: a label keeps its round forever, a round’s held-out accuracy is rewritten if the round is re-fit.
- spacr.active_learning.entropy(probs: Any, base: float | None = None, normalize: bool = False) numpy.ndarray[source]¶
Shannon entropy
−Σ p log pof each row.The only measure here that uses the whole distribution. Minimum 0 at a one-hot row; maximum
log Cat the uniform row (log 2 ≈ 0.6931for two classes in nats).0 · log 0is taken as 0.- Parameters:
probs –
(N,)or(N, C); coerced withas_probabilities().base – logarithm base.
None(default) means natural log, so the units are nats; pass2for bits.normalize – divide by
log Cso the maximum is 1. Monotone, so the ranking is unchanged.
- Returns:
(N,)uncertainty scores; NaN for unusable rows.
- spacr.active_learning.format_coverage_summary(coverage: pandas.DataFrame) str[source]¶
Render
annotation_coverage()as text, worst concentration first.- Parameters:
coverage – the frame from
annotation_coverage().- Returns:
multi-line text, no trailing newline.
- spacr.active_learning.format_learning_curve(curve: pandas.DataFrame, verdict: StoppingVerdict | None = None) str[source]¶
Render the round-by-round curve and the stopping verdict as text.
- Parameters:
curve – round-by-round metrics frame from
learning_curve().
- spacr.active_learning.format_queue_summary(queue: pandas.DataFrame) str[source]¶
Render a queue’s shape, class balance and caveats as plain text.
Reports the numbers that decide whether the queue is worth working through: how much of the screen is already labelled, how many crops could not be scored, the range of the scores, how many wells the queue actually spreads over (the diversity check), and the class balance of the queue next to the balance of the pool it came from — because uncertainty sampling on an imbalanced screen pulls hard toward the majority class’s boundary and the annotator should see that rather than discover it.
- Parameters:
queue – frame from
build_queue(). Works on a slice or a copy too, falling back to what can be recomputed from the rows whenattrsdid not survive.- Returns:
multi-line text, no trailing newline.
- spacr.active_learning.holdout_report(y_true: Any, probs: Any, classes: Sequence[Any] | None = None) Dict[str, Any][source]¶
Held-out metrics, with the confusion matrix they were derived from.
Torch-free, so both the classical-ML round here and
spacr.deep_spacr.held_out_report()can use one implementation and a model card written by either says the same thing in the same shape.Every derived figure is exactly the standard function of the matrix —
accuracy = trace / total,per_class[c] = M[c, c] / M[c, :].sum()— so a reader can recompute the card rather than trust it.nis the number of rows the matrix actually contains, and the supports sum to it. A row whose true class the head has no column for is counted as an error rather than dropped: the matrix grows to hold it, and the missing column stays empty because the head can never predict that class. Reportingnover one population andaccuracyover another is the one thing a model card must not do — a three-class held-out set scored by a binary head used to report a perfect score on a set the model got two of three right, and theaccuracy == trace / totalinvariant still checked out because the row was missing from both sides of it.- Parameters:
y_true – integer class ids, shape
(N,). Negative ids are not classes; they are excluded, counted inn_unscoredand explained innotesrather than silently folded into the total.probs –
(N,)positive-class probabilities, or(N, C)rows.classes – class names in head order.
- Returns:
n,n_unscored,num_classes,head_classes,classes,accuracy,f1_macro,per_class_accuracy,class_support,predicted_support,confusion_matrix,notes.- Raises:
ValueError – when there are not as many score rows as labels — the two are not aligned, and every figure below would be a comparison of one object’s label with another’s prediction.
- spacr.active_learning.label_rounds(db_path: str, annotation_column: str = 'annotate') pandas.DataFrame[source]¶
Per-label round provenance as a frame (empty when never recorded).
- Parameters:
db_path – path to the
measurements.dbto read.
- spacr.active_learning.learning_curve(db_path: str, annotation_column: str = 'annotate') pandas.DataFrame[source]¶
Held-out accuracy per round, oldest first — the curve to watch flatten.
- Parameters:
db_path – path to
measurements.db.annotation_column – the column the rounds labelled into.
- Returns:
a frame with one row per round:
round,finished_utc,n_labels,n_new_labels,n_holdout,holdout_accuracy,holdout_f1_macro,per_class(dict),split_rule,model_type,model_path,card_path,notes(list), and the derivedgain(accuracy change since the previous round). Empty (with those columns) when no round has been recorded.
- spacr.active_learning.least_confidence(probs: Any, normalize: bool = False) numpy.ndarray[source]¶
1 − max_c p_c— how much probability mass is not on the winner.Minimum 0 at a one-hot row; maximum
1 − 1/Cat the uniform row. It looks only at the top class, so on 3+ classes it cannot tell[0.5, 0.5, 0.0]from[0.5, 0.25, 0.25];entropy()can.- Parameters:
probs –
(N,)positive-class probabilities or(N, C)rows. Coerced withas_probabilities().normalize – rescale by
C / (C − 1)so the maximum is 1. A monotone rescaling — it never changes the ranking, only how the number reads.
- Returns:
(N,)uncertainty scores; NaN for unusable rows.
- spacr.active_learning.margin(probs: Any) numpy.ndarray[source]¶
1 − (p₁ − p₂)— closeness of the top two classes, as uncertainty.This returns an uncertainty score, not the margin: a small margin (the two leading classes neck and neck) is a large return value, so it is oriented like every other measure here. Minimum 0 at a one-hot row, maximum 1 at any row whose top two classes tie.
On two classes this is a linear transform of
least_confidence()—1 − (p₁ − p₂) = 2·(1 − max p)— so the two produce the identical order and the identical ties. They are one choice, not two, until C ≥ 3, where margin ignores everything below the runner-up and least-confidence ignores everything below the winner.- Parameters:
probs –
(N,)or(N, C); coerced withas_probabilities().- Returns:
(N,)uncertainty scores in [0, 1]; NaN for unusable rows.
- spacr.active_learning.next_round(db_path: str, annotation_column: str = 'annotate') int[source]¶
The round number the next batch of labels belongs to.
- Parameters:
db_path – path to the
measurements.dbwhose round log is queried.
Round 0 is “before any model was retrained from inside Annotate” — the labels that seeded the loop. The first retrain produces round 1.
- spacr.active_learning.predict_probabilities(model: Callable[[Any], Any], batches: Iterable[Any], device: Any = None, from_logits: bool = True) numpy.ndarray[source]¶
Run
modeloverbatchesand return an(N, C)probability matrix.A convenience for scoring crops that are not in the database yet. The queue does not need this —
build_queue()works from thepredcolumn thatspacr.deep_spacr.merge_predictions_into_db()already wrote, which is the normal path.torch is imported inside this function, and only to get
no_grad/devicehandling; if it is not importable the batches are iterated and the model is called directly. Nothing else in this module touches torch.- Parameters:
model – any callable mapping a batch to raw head outputs. A
torch.nn.Moduleis put ineval()mode first if it has one.batches – iterable of batches. A batch that is a
(inputs, …)tuple has its first element passed to the model, matching the loaders inspacr.deep_spacr.device – optional torch device to move inputs/model to.
from_logits – treat outputs as raw logits and apply
probabilities_from_logits()(the default — a model head emits logits). Set False if the model already outputs probabilities, which then go throughas_probabilities().
- Returns:
(N, C)probability matrix in batch order.
- spacr.active_learning.probabilities_from_logits(logits: Any) numpy.ndarray[source]¶
Convert raw classifier-head outputs to an
(N, C)probability matrix.The head shape decides the link function, and getting this wrong is silent:
(N,)or(N, 1)— a single-logit binary head. Sigmoid, then expanded to[1 − p, p]so every measure sees two classes. Pushing this through a softmax instead would return a column of 1.0 — every crop maximally certain, the whole ordering gone.(N, C)with C ≥ 2 — a C-logit head. Row-wise softmax. Sigmoiding these instead reads each logit in isolation and can invert the order:[5, 5]is a perfect 50/50 tie under softmax but looks like a confident 0.993 under a sigmoid of column 1.
Both branches are exactly what
spacr.deep_spacr.apply_model_to_tar()andspacr.deep_spacr.evaluate_model_performance()do at inference time.- Parameters:
logits – array-like or torch tensor of raw head outputs.
- Returns:
(N, C)float array whose rows sum to 1 (C ≥ 2).- Raises:
ValueError – for a scalar or a 3-D input.
Note
The output is a probability vector, not a calibrated probability. See
CALIBRATION_NOTE.
- spacr.active_learning.queue_rows(queue: pandas.DataFrame, key: str = PNG_KEY) List[Tuple[str, int | None]][source]¶
The queue as
[(png_path, None), …].Exactly the shape
spacr.qt.annotate_engine.fetch_page()andspacr.qt.annotate_engine.fetch_filtered_paths()return, so the Annotate screen can page through a queue with no other change. The annotation is alwaysNone— every crop in a queue is unlabelled by construction.- Parameters:
queue – frame from
build_queue().key – the path column (default
png_path).
- Returns:
list of
(path, None)tuples in queue order.
- spacr.active_learning.rank_by_uncertainty(probs: Any, measure: Any = DEFAULT_MEASURE, limit: int | None = None, seed: int | None = None, scores: Any | None = None) numpy.ndarray[source]¶
Row indices ordered most-uncertain-first.
Ordering is total and reproducible:
primary key — the uncertainty score, descending;
ties — row index ascending by default, or a permutation seeded with
seedwhen one is given. Both are deterministic: the same inputs and the same seed always give the same order. Ties are the normal case on a screen where thousands of crops score exactly 0.5, so an unseeded shuffle there would reshuffle the annotator’s queue on every refresh.NaN scores sort last, always, and never in front of a real one. They are ranked rather than dropped so the returned indices stay a permutation of
range(N);build_queue()drops them and says how many.
- Parameters:
probs –
(N,)or(N, C)probabilities.measure – name from
UNCERTAINTY_MEASURESor a callable.limit – keep only the first
limitindices.seed – seed for tie-breaking;
Nonebreaks ties on index.scores – pre-computed scores to rank instead of recomputing from
probs(used bybuild_queue(), and by anything ranking adisagreement()score).
- Returns:
(N,)(or(limit,)) int array of row indices.
- spacr.active_learning.record_labels(db_path: str, annotation_column: str, labels: Dict[str, Any], round_index: int, source: str = 'manual') int[source]¶
Stamp each label with the round it was made in.
Called by the Annotate screen every time it flushes a batch. The round a label came from is what makes early-round bias auditable: the first round’s labels are drawn from whatever ordering existed before any model had seen this screen, and if 80 % of a class’s labels carry round 0 then the “active learning” was mostly not active.
first_roundis preserved across re-labelling whileroundfollows the current value, so both “when was this crop first looked at” and “which round set the label it has now” survive a correction.- Parameters:
db_path – path to
measurements.db.annotation_column – the column the labels were written into.
labels –
{png_path: class or None}.Noneis a cleared label and is recorded as such rather than dropped — a crop that was looked at and deliberately left blank is not the same as one never seen.round_index – the round these labels belong to.
source – how the crop reached the annotator —
'manual','queue', or a caller’s own tag.
- Returns:
number of rows written.
- spacr.active_learning.record_round(db_path: str, annotation_column: str, round_index: int, **fields: Any) None[source]¶
Append (or replace) one row of the learning curve.
- Parameters:
db_path – path to
measurements.db.annotation_column – the column this round labelled into.
round_index – the round number.
fields – any of
n_labels,n_new_labels,n_holdout,holdout_accuracy,holdout_f1_macro,per_class(dict),split_rule,model_type,model_path,card_path,measure,diversity,notes(list).
- spacr.active_learning.resolve_measure(measure: Any) Tuple[str, Callable[..., numpy.ndarray]][source]¶
Turn a measure name (or callable) into
(name, function).- Parameters:
measure – a key of
UNCERTAINTY_MEASURES, or any callablef(probs) -> (N,).- Raises:
ValueError – for an unknown name, listing the valid ones.
- spacr.active_learning.retrain_round(db_path: str, annotation_column: str = 'annotate', *, features: pandas.DataFrame | None = None, model_type: str = 'logistic_regression', group_by: str = 'well', holdout: float = 0.25, seed: int = 0, min_labels: int = 8, round_index: int | None = None, table: str = PNG_TABLE, key: str = PNG_KEY, image_type: str | None = None, write_scores: bool = True, save_model: bool = True, model_dir: str | None = None, write_card: bool = True, label_window: int = 50, min_gain: float = 0.003, measure: Any = DEFAULT_MEASURE, diversity: Any = 'well', balance: str = 'none', synthetic_negatives: int | None = None, rejections: Mapping[Any, Any] | None = None) RoundResult[source]¶
Fit a model on the labels so far, score every crop, close the loop.
This is the half of active learning that has been missing: the queue put the informative crops in front of the annotator, and then nothing happened. Annotating without retraining is not active learning — it is ordinary annotation in a clever order, and the order goes stale after the first few dozen labels because it still reflects a model that has not seen any of them.
One call does all five things the loop needs:
fits a model on every label in
annotation_column;scores it on a grouped held-out split, so the number is not an artefact of 190 labels coming from one well;
writes per-class probabilities back into
png_listasROUND_PRED_PREFIXcolumns, whichbuild_queue()prefers over the olderpred— so the next queue is genuinely re-ranked;records the round, giving
learning_curve()another point;returns the
StoppingVerdictfor the curve so far.
- Parameters:
db_path – path to
measurements.db.annotation_column – the column holding the labels.
features – feature matrix indexed by the crop key. Omitted, it is read from the measurement tables with
round_features().model_type –
'logistic_regression'(default — it is the one that behaves at 20 labels),'random_forest'or'gradient_boosting'.group_by –
'well'(default),'plate','field'or'none'. What the held-out split refuses to share. Matched exactly and in lower case; anything else raises, the way an unknowndiversity=does inbuild_queue(). When the strategy is'none', or the crop table has none of the columns it needs, the split is a stratified random one andsplit_rulesaysNOT grouped— it never claims a grouping it did not perform.holdout – fraction held out.
seed – makes the split and the fit reproducible.
min_labels – refuse to fit below this many labels.
round_index – override the round number; defaults to
next_round().table – crop table.
key – crop key column.
image_type – substring filter on the crop key.
write_scores – write the new probabilities back into the database.
save_model – joblib-dump the fitted model beside the database.
model_dir – where to put it; defaults to
<db dir>/active_learning.write_card – write a model card beside the saved model.
label_window – passed to
should_stop().min_gain – passed to
should_stop().measure – recorded with the round, for the queue that follows.
diversity – likewise.
balance –
'none'(default) leaves imbalance to the estimator’s ownclass_weight='balanced';'downsample'cuts every class to the size of the smallest BEFORE the grouped split.THE TWO ARE NOT THE SAME ANSWER. Reweighting and downsampling produce different probabilities from the same crops, and
spacr.suggest.suggest_from_scores()sorts on those probabilities – so the round records which was in force, innotesand on the model card. The smaller class (“if there is class imbalance use the class with fewer”).rejections –
{crop key: rejected class}– suggestions the annotator REJECTED (spacr.suggest.rejected_suggestions()). In a two-class column a rejection of class 1 is an example of class 2, and it is fitted as one; a crop that has since been labelled is left to its label, and in a column with any other classes the rejection cannot be turned into a label and is counted in the notes instead. A rejection is information, not silence: without this the model that proposed the wrong class would be fitted on exactly the same evidence next round and propose it again.synthetic_negatives –
how many unannotated crops to draw at random and fit as the ABSENT class when only one class has been annotated.
None(default) refuses instead, as before.THIS IS A DELIBERATE LIE AND THE ROUND SAYS SO. A random draw from the unannotated pool is mostly-negative, not negative, so what comes back is a ranking rather than a verdict; the count reaches the model card, because a card that does not say the negatives were invented describes a model that does not exist. Defined for the binary classes 1 and 2 only – see
_absent_binary_class().
- Returns:
a
RoundResult.- Raises:
ValueError – below
min_labelslabels, with fewer than two classes annotated — neither is something to paper over with a model that will produce a confident-looking ranking out of nothing — or for an unrecognisedgroup_by.
- spacr.active_learning.round_features(db_path: str, table: str = PNG_TABLE, key: str = PNG_KEY, tables: Sequence[str] = ('cell', 'nucleus', 'pathogen', 'cytoplasm'), nuclei_limit: int = 10, pathogen_limit: int = 10) pandas.DataFrame[source]¶
The measurement features for every crop, indexed by
png_path.The feature matrix an in-screen retrain fits on. Measurement features rather than pixels on purpose: the point of retraining from inside Annotate is to get a fresh ranking in seconds, on the machine the annotator is sitting at, without a GPU and without leaving the screen. A CNN retrain is the right thing to do at the end of the loop, not between two pages of crops.
- Parameters:
db_path – path to
measurements.db.table – crop table carrying
png_pathandprcfo.key – the crop key column.
tables – object tables to merge features from; missing ones are skipped.
nuclei_limit – passed through to
spacr.io._read_and_merge_data().pathogen_limit – likewise.
- Returns:
numeric features indexed by
png_path.- Raises:
ValueError – when no object table with features could be read.
- spacr.active_learning.should_stop(curve: pandas.DataFrame, *, label_window: int = 50, min_gain: float = 0.003, min_rounds: int = 2) StoppingVerdict[source]¶
Has the last
label_windowlabels moved held-out accuracy at all?The rule, in one line: look back over whole rounds until at least
label_windownew labels have accumulated, and compare held-out accuracy at the two ends. If it moved by less thanmin_gain, stop.Why this rule and not another:
Labels, not rounds, are the unit of cost. A round is whatever size the annotator felt like; “no improvement for 3 rounds” says nothing when the rounds were 5, 300 and 8 labels. The thing being spent is human attention, one crop at a time, so the window is measured in crops.
It looks *back over* rounds, not *at* the last one. A single round that happened to land flat is noise; the question is whether the last fifty labels — however they were divided up — bought anything.
It refuses to answer early. Below
label_windowlabels since the first recorded round there is no window to measure, and a rule that fired anyway would tell people to stop after their first twelve labels.It distinguishes “flat” from “unmeasurable”. Held-out accuracy on 80 objects has a standard error near 0.05; a 0.3 % change is inside the noise, and
StoppingVerdict.confidentsays so instead of dressing it up. Flat is still the recommendation — if more labels are not moving a number you can measure, they are not buying anything you can demonstrate — but the reason says the held-out set is too small to prove convergence, which is a different piece of work.A falling curve stops too, and says so. Accuracy going down is not convergence; it usually means the newest labels disagree with the earlier ones, or that the held-out split moved. Either way, more of the same is the wrong next move.
- Parameters:
curve – the frame from
learning_curve().label_window – how many labels the window must cover.
min_gain – accuracy change below which the window counts as flat.
min_rounds – rounds required before any verdict is given.
- Returns:
a
StoppingVerdict;bool(verdict)is the answer.
- spacr.active_learning.uncertainty_scores(probs: Any, measure: Any = DEFAULT_MEASURE) numpy.ndarray[source]¶
Score every row with
measure.- Parameters:
probs –
(N,)or(N, C)probabilities.measure – name from
UNCERTAINTY_MEASURESor a callable.
- Returns:
(N,)uncertainty scores, larger = less certain.
Nested helpers¶
- _segmentation_uncertainty._paint(values: np.ndarray) np.ndarray¶
- Parameters:
values – one value per reference object, in id order.
- Returns:
each object’s pixels set to its value, 0 elsewhere.
spacr/active_learning.py:3910
- crops_for_object_keys._register(target: Dict[str, Tuple[str, int | None]], composed: List[str], label: str, object_type: str | None, entry: Tuple[str, int | None]) None¶
Register first-wins untyped and, when known, typed object keys.
spacr/active_learning.py:1637