spacr.suggest¶
Turn a retrained model’s scores into labels you can accept or throw away.
Annotating is the slow part of every screen, and the first two hundred crops already contain most of what separates the classes. Annotate can already fit a model on those and re-rank the queue by uncertainty – what it could not do is WRITE the model’s opinion down as a proposed label that a reviewer accepts or rejects in bulk. That is all this module adds.
IT DELEGATES THE MODEL, DELIBERATELY. spacr.active_learning.retrain_round()
already fits on every label, scores on a GROUPED held-out split so the number
is not an artefact of 190 labels coming from one well, and writes per-class
probabilities into png_list. Re-fitting here would be a second model with
a second answer, drifting from the one the queue is ranked by – and a first
draft of this file did exactly that before the existing one was found. What is
new is the three rules below, not the classifier.
THREE RULES THAT PROTECT THE ANNOTATIONS¶
A SUGGESTION IS NEVER CONFUSABLE WITH A DECISION. Suggestions are stored as their own values – a suggested 1 becomes 11 – so nothing that reads the annotation column can mistake one for a human’s answer, and a run that goes wrong is undone by deleting a value rather than by remembering which rows were touched.
A SUGGESTION NEVER OVERWRITES AN ANNOTATION. Writes carry
IS NULL. Not when the model is confident, not on a re-run. A human’s labels are the ground truth the model was fitted on; losing one silently would cost hours and would not be noticed until a run came out wrong.NO CONFIDENCE FLOOR, SO THE ORDER CARRIES THE DOUBT. Every unannotated crop gets a suggestion, sorted most-confident first, and the reviewer stops where they stop agreeing. A threshold would make that decision for them, with a number nobody chose.
A JUDGEMENT IS RECORDED, NOT INFERRED. Confirming a suggestion turns it into an ordinary label and rejecting one clears it – and both are written down beside the column, in
<column>_verdict, as the class that was confirmed (+c) or rejected (-c). A rejection is information the next round can train on – “not class 1” is an example of class 2 in a two-class column – where a NULL would have been silence. The verdict column is added the first time a source is opened, so tables made before it existed gain it without a migration step.
Classes¶
What a suggestion run proposes, and how far to trust it. |
Functions¶
|
Add the verdict column to |
|
The recorded judgement of each of |
|
How the column's suggestions stand: judged, and still to judge. |
|
How many suggestions are waiting to be accepted or thrown away. |
|
The crops whose suggestion was rejected, and the class that was refused. |
|
Accept suggestions as annotations, or clear them away. |
|
Read the last retrain's probabilities and propose a label for each crop. |
|
The column beside |
|
Store suggestions, and ONLY where nothing has been annotated. |
Module Contents¶
- class spacr.suggest.Suggestions[source]¶
What a suggestion run proposes, and how far to trust it.
- Parameters:
frame – one row per unannotated crop with
png_path,suggested,storedandconfidence, most confident first.note – what a reader must be told before accepting in bulk.
scored – how many crops carried usable scores.
classes – the class values the scores describe, in column order.
- spacr.suggest.ensure_verdict_column(db_path: str, annotation_column: str, *, png_table: str = 'png_list') bool[source]¶
Add the verdict column to
png_tableif it is missing.Called when a source is opened, which is how a table made before the column existed gains it: an
ALTER TABLE ... ADD COLUMNwith no default is a metadata change in SQLite and rewrites no rows.- Parameters:
db_path – path to a
measurements.db.annotation_column – the column whose judgements it will hold.
png_table – the crop table.
- Returns:
True when the column exists afterwards.
- spacr.suggest.fetch_verdicts(db_path: str, annotation_column: str, paths: Sequence[str], *, png_table: str = 'png_list') Dict[str, int][source]¶
The recorded judgement of each of
pathsthat has one.- Parameters:
db_path – path to a
measurements.db.annotation_column – the column the judgements belong to.
paths – the crops on the page.
png_table – the crop table.
- Returns:
{png_path: verdict}for the crops that carry one; empty when the column does not exist yet.
- spacr.suggest.judgement_counts(db_path: str, annotation_column: str, *, png_table: str = 'png_list') Dict[str, int][source]¶
How the column’s suggestions stand: judged, and still to judge.
- Parameters:
db_path – path to a
measurements.db.annotation_column – the column the suggestions were written into.
png_table – the crop table.
- Returns:
{"left": n, "confirmed": n, "rejected": n}–leftis the suggestions nobody has judged yet, the other two are every judgement recorded in the column so far.
- spacr.suggest.pending_suggestions(db_path: str, annotation_column: str, *, png_table: str = 'png_list') int[source]¶
How many suggestions are waiting to be accepted or thrown away.
- Parameters:
db_path – path to a
measurements.db.annotation_column – the column holding them.
png_table – the crop table.
- Returns:
the count, or 0 when the column does not exist.
- spacr.suggest.rejected_suggestions(db_path: str, annotation_column: str, *, png_table: str = 'png_list') Dict[str, int][source]¶
The crops whose suggestion was rejected, and the class that was refused.
Only crops the annotator has NOT since labelled are returned: a label made after a rejection is the stronger statement and is what the fit reads from the column itself, so handing the rejection over as well would count the crop twice.
- Parameters:
db_path – path to a
measurements.db.annotation_column – the column the suggestions were written into.
png_table – the crop table.
- Returns:
{png_path: rejected class}; empty when nothing was rejected or the verdict column does not exist.
- spacr.suggest.resolve_suggestions(db_path: str, annotation_column: str, *, keep: bool, png_table: str = 'png_list', paths: Sequence[str] | None = None) int[source]¶
Accept suggestions as annotations, or clear them away.
Accepting rewrites the offset value to the real class; rejecting sets it back to NULL. Both act ONLY on suggestion values, so a human annotation caught by the same query is untouched either way.
- Parameters:
db_path – path to a
measurements.db.annotation_column – the column holding both.
keep – True to accept, False to discard.
png_table – the crop table.
paths – restrict to these crops; None means every suggestion.
- Returns:
how many rows changed.
A bulk KEEP is a confirmation of every suggestion it keeps, so it is recorded in the verdict column too (when the column exists): the crops then wear the same mark as ones confirmed one at a time, and the count the screen shows agrees with what happened. A bulk THROW is not a judgement – “I do not want to review these” is not “these are wrong” – so it records nothing.
- spacr.suggest.suggest_from_scores(db_path: str, annotation_column: str, *, png_table: str = 'png_list', classes: Sequence[int] | None = None, withhold_rejected: bool = True) Suggestions[source]¶
Read the last retrain’s probabilities and propose a label for each crop.
Reads rather than re-fits, so the suggestion a reviewer sees and the ranking the queue uses come from ONE model. A second fit here would drift from it and there would be no way to tell which was right.
- Parameters:
db_path – path to a
measurements.db.annotation_column – the column holding the labels.
png_table – the crop table.
classes – the class value each score column stands for. Defaults to the values already present in the annotation column, in order, which is what the retrain encoded them from.
withhold_rejected – leave out every crop whose suggestion the annotator REJECTED (a negative
<column>_verdict), so a rejected crop is not suggested again in a later round. It still trains as an example of the other class (rejected_suggestions()); it is only not proposed.
- Returns:
a
Suggestions; its frame is empty when nothing has been scored, andnotesays why.
- spacr.suggest.verdict_column(annotation_column: str) str[source]¶
The column beside
annotation_columnthat records judgements.One value per crop:
+cwhen a suggested classcwas confirmed,-cwhen it was rejected, NULL when nothing was judged. It lives in the crop table rather than in the screen so a judgement survives a restart and reaches the next round of training.- Parameters:
annotation_column – the column the suggestions were written into.
- Returns:
the verdict column’s name.
- spacr.suggest.write_suggestions(db_path: str, annotation_column: str, suggestions: pd, *, png_table: str = 'png_list') int[source]¶
Store suggestions, and ONLY where nothing has been annotated.
The
IS NULLis rule 2 and is not an optimisation.- Parameters:
db_path – path to a
measurements.db.annotation_column – the column to write into.
suggestions – the frame from
suggest_from_scores().png_table – the crop table.
- Returns:
how many rows were written.
Shared connections are in autocommit mode, so the connection context alone would not make these updates one atomic batch.
Nested helpers¶
- _score_columns.index(name)¶
The integer suffix of a score column, for ordering.
spacr/suggest.py:143