spacr.confusion¶
C8 — the confusion matrix as a set of live queries rather than a picture.
A confusion matrix is the most-looked-at and least-acted-on artefact a classifier run produces. It says “43 uninfected objects were called infected” and then stops: the 43 are anonymous, so the only thing anyone can do with the number is feel bad about it. This module is the part that turns each cell back into the objects it counted, so the number becomes a question you can open.
Three ideas, and they are separate on purpose¶
A cell is a set of objects. cell_rows() is the whole primitive: the
rows of the out-of-fold prediction table whose true class is this and whose
predicted class is that. Everything else here is a way of ordering,
splitting or counting that set.
A confident error and an unsure error have different causes. They are therefore two lists, never one sorted list with a gradient in it:
high confidence, and wrong — the model was sure. When a model that is right 95% of the time is certain about an object and disagrees with the annotation, the likeliest explanation is that the annotation is wrong. These are the crops to re-label.
low confidence, and wrong — the model was unsure, and fell the wrong side. The label is probably fine; the boundary is where the work is — more examples near it, a better feature, or an admission that the two classes are not separable on this stain.
Handing back one list sorted by confidence buries that distinction in the
middle of a scroll. split_by_confidence() returns the two lists and
guarantees they partition the cell, so nothing is silently dropped between
them.
A cell is not a homogeneous population. 43 errors spread evenly over 20
wells is a model problem. 43 errors all from well A01 is a staining problem,
and re-labelling any of them is wasted work — the fix is upstream, at the
bench. breakdown_by() and describe_breakdown() are that check,
made before anyone opens a single crop.
Where the frame comes from¶
spacr.classifier_evaluation.evaluate_predictions() writes the table this
module reads: one row per held-out object, with true_class,
predicted_class, confidence (the calibrated probability of the class
the model chose) and the identity columns
sample_identity() parsed out of the crop path. Nothing here imports Qt, so the
same analysis runs in a notebook, and the expensive half never touches the
event loop.
Exceptions¶
A prediction table that cannot answer the question being asked of it. |
Classes¶
One off-diagonal cell, ranked against the others. |
|
One clicked cell, already split and already counted. |
Functions¶
|
Where one cell's objects came from, most concentrated first. |
|
The rows one confusion-matrix cell counted, in the table's own order. |
|
Where "the model was sure" starts, for a |
|
Counts per (true, predicted) pair, as a square frame. |
|
One cell's origin, in words, with the verdict spelled out. |
|
The off-diagonal mass in words, worst first. |
|
How many rows of |
|
Which column of |
|
The object keys of |
|
Every off-diagonal cell, worst first. |
|
Split one cell into suspect the label and suspect the boundary. |
Module Contents¶
- exception spacr.confusion.ConfusionError[source]¶
Bases:
ValueErrorA prediction table that cannot answer the question being asked of it.
Raised rather than returning an empty result. “No rows in that cell” and “this frame has no
true_classcolumn” look identical to a caller that only sees a count, and the second one has produced a confusion matrix of zeros that everybody read as a perfect classifier.Initialize self. See help(type(self)) for accurate signature.
- class spacr.confusion.Confusion[source]¶
One off-diagonal cell, ranked against the others.
- Parameters:
true_class – annotated class naming the confusion-matrix row.
predicted_class – model-assigned class naming the confusion-matrix column.
count – number of objects in this off-diagonal cell.
share_of_errors – this cell’s fraction of all off-diagonal errors.
rate_within_true – this cell’s fraction of the complete row for
true_class.
- class spacr.confusion.ConfusionCell[source]¶
One clicked cell, already split and already counted.
Built by
build()so that the screen does one call and gets everything it draws — the two lists, the keys to route, and the sentences — rather than orchestrating five functions in a mouse handler.- Parameters:
true_class – annotated class naming the confusion-matrix row.
predicted_class – model class naming the confusion-matrix column.
threshold – confidence boundary separating the two review queues.
rows – every object in this matrix cell, retained in source-table order.
high – rows with confidence at least
threshold, ordered most confident first.low – rows below
threshold—including missing confidence—ordered least confident first.
- classmethod build(predictions: pandas.DataFrame, true_class: Any, predicted_class: Any, *, threshold: float | None = None, n_classes: int | None = None) ConfusionCell[source]¶
Resolve a cell and split it.
- Parameters:
predictions – evaluated-object table containing the true and predicted class columns. A non-empty resolved cell must also carry confidence; identity columns are retained when present but are not required.
true_class – annotated class naming the matrix row to resolve.
predicted_class – model class naming the matrix column to resolve.
threshold – where “sure” starts. Defaults to
confidence_threshold()ofn_classes, or of the number of distinct classes inpredictionswhen that is not given.
- keys(which: str = 'all', *, column: str | None = None) pandas.Index[source]¶
Object keys for
"high","low"or"all", in list order."all"is high-confidence first and then low-confidence — not table order — so a caller that opens the whole cell still gets the objects most likely to be mislabelled at the front of the grid.
- reason(which: str = 'all') str[source]¶
The line the receiving view puts above the crops.
Required by
spacr.selection.ObjectRequest, and load-bearing: a grid of twelve crops that does not say why reads as the whole dataset. It names the hypothesis, not just the cell, so the person looking at the crops knows what they are being asked to decide.
- spacr.confusion.breakdown_by(rows: pandas.DataFrame, level: str) pandas.DataFrame[source]¶
Where one cell’s objects came from, most concentrated first.
- Parameters:
rows – a cell, from
cell_rows().level – an identity column —
"plate","well"or"field". Any column ofrowsis accepted, so a bundle carrying extra metadata can be broken down by it too.
- Returns:
a frame of
level,countandshare, sorted by count descending then by name, so the answer is stable across runs.- Raises:
ConfusionError – for a level the frame does not carry — silently returning an empty breakdown would read as “the errors are spread evenly”, which is the opposite of what an absent column means.
- spacr.confusion.cell_rows(predictions: pandas.DataFrame, true_class: Any, predicted_class: Any) pandas.DataFrame[source]¶
The rows one confusion-matrix cell counted, in the table’s own order.
- Parameters:
predictions – evaluated-object table carrying true and predicted class columns.
true_class – annotated class naming the matrix row.
predicted_class – model class naming the matrix column.
Compared as text, deliberately. A confusion matrix read back from CSV has string class names in its index and header, while the prediction table may have kept an integer or a categorical — and a cell that matched nothing because
1 != "1"renders as an empty grid with no error, which reads as “no mistakes here”.- Returns:
a copy, so a caller sorting it cannot reorder the bundle.
- spacr.confusion.confidence_threshold(n_classes: int) float[source]¶
Where “the model was sure” starts, for a
n_classes-way problem.- Parameters:
n_classes – number of mutually exclusive classifier classes.
confidenceis the probability of the class the model chose, so it can never fall below1 / n_classes— a two-class model is at 0.5 when it is maximally undecided, and a ten-class model is at 0.1. A fixed 0.5 would therefore call every ten-class error “high confidence” and every two-class error nothing at all, which is the sort of default that makes a feature look broken on somebody else’s data.The midpoint between chance and certainty is the honest default: 0.75 for two classes, 0.55 for ten. It is a default, not a law — every function here takes an explicit
threshold, and the screen exposes it, because where “sure” starts is a property of the assay and not of arithmetic.- Raises:
ConfusionError – for fewer than two classes.
- spacr.confusion.confusion_counts(predictions: pandas.DataFrame, classes: Sequence[Any] | None = None) pandas.DataFrame[source]¶
Counts per (true, predicted) pair, as a square frame.
Recomputed from the prediction table rather than read from the bundle’s
confusion_counts.csv, so that the matrix on screen and the objects a cell opens can never disagree — the two used to be separate artefacts and a filtered prediction table silently produced a matrix nobody could reproduce.- Parameters:
predictions – one row per evaluated object, including
TRUE_COLUMNandPREDICTED_COLUMN.classes – the class order. Defaults to every class appearing as a true label or a prediction, sorted, so a class the model never chose still gets its column rather than vanishing from the matrix.
- spacr.confusion.describe_breakdown(rows: pandas.DataFrame, level: str) str[source]¶
One cell’s origin, in words, with the verdict spelled out.
- Parameters:
rows – evaluated-object rows belonging to one confusion cell.
level – identity column, such as
wellorplate, used to group the rows.
The point of this line is to stop wasted work. If all 43 errors come from well A01, re-labelling any of them corrects nothing that will recur — the fix is a staining or a focus problem at the bench, and the crops are evidence for that conversation rather than a re-annotation queue.
- Returns:
one or two lines, no trailing newline.
- spacr.confusion.describe_confusions(counts: pandas.DataFrame, *, limit: int = 3) str[source]¶
The off-diagonal mass in words, worst first.
A matrix of numbers makes the reader do the ranking, and the ranking is the only part of it anybody acts on. This says it: “Your worst confusion is uninfected → infected: 43 object(s), 45% of all errors…”
- Parameters:
counts – square true-by-predicted count table accepted by
rank_confusions().limit – how many confusions to name before summarising the rest.
- Returns:
one or more lines, no trailing newline. Never empty — a perfect classifier gets a sentence saying so, because a blank panel reads as a panel that failed to load.
- spacr.confusion.key_collisions(frame: pandas.DataFrame, *, column: str | None = None) int[source]¶
How many rows of
frameshare a key with an earlier row.- Parameters:
frame – evaluated-object rows whose crop keys are checked.
Zero for a healthy prediction table. Non-zero means the crop grid will hold fewer objects than the confusion cell counted, and the difference is this number — worth saying on screen rather than leaving as an unexplained discrepancy between a matrix and a grid.
A bundle whose
object_keycolumn carries the object type (spacr.selection.OBJECT_TYPE_COLUMN) no longer counts a cell’s nucleus and its pathogen as one object here, because they are no longer one key.
- spacr.confusion.object_key_column(frame: pandas.DataFrame) str[source]¶
Which column of
framenames the objects, best first.- Parameters:
frame – evaluated-object table to route back to crops.
- Raises:
ConfusionError – when none of
KEY_COLUMNSis present, which means the rows cannot be routed anywhere and a “show me these crops” button would be a button that raises on click.
- spacr.confusion.object_keys_for(frame: pandas.DataFrame, *, column: str | None = None) pandas.Index[source]¶
The object keys of
frame’s rows, in frame order, de-duplicated.- Parameters:
frame – evaluated-object rows whose crop keys are requested.
Order is load-bearing — it is what carries “worst error first” through
spacr.qt.linked_selection.open_objects()— so this preserves it rather than sorting or using a set.Duplicates are dropped. They are not assumed away: two rows can legitimately share a key — augmented copies of one crop collapse onto the same
objectstem — sokey_collisions()says how many rows the grid will be short of the count the matrix showed, rather than leaving that to be discovered by counting tiles.Two sources of duplication have since been removed from underneath this, and both were silent:
spacr.selection.object_keys()was not injective when an identity component contained the key separator (it now escapes one), and it carried no object type, so a nucleus 1 and a pathogen 1 in one field were one key. What is left here is the augmentation case, which is real and belongs to the training set rather than to the key.
- spacr.confusion.rank_confusions(counts: pandas.DataFrame) List[Confusion][source]¶
Every off-diagonal cell, worst first.
“Worst” is share of total errors, because that is what answers “what should I fix next”. Ties break on the within-class rate and then on the class names, so the ranking is stable rather than dependent on dict order.
- Parameters:
counts – a square frame from
confusion_counts()or from a bundle’sconfusion_counts.csv.
- spacr.confusion.split_by_confidence(rows: pandas.DataFrame, threshold: float) Tuple[pandas.DataFrame, pandas.DataFrame][source]¶
Split one cell into suspect the label and suspect the boundary.
- Parameters:
rows – a cell, from
cell_rows().threshold – confidence at or above which the model counts as sure.
- Returns:
(high, low).highis confidence>= threshold, most confident first — the model’s flattest contradictions of the annotator, which is the order to re-label in.lowis confidence< threshold, least confident first — the objects nearest the decision boundary, which is the order to look at when asking whether the boundary is in the right place.
The two always partition
rowsexactly: every row is in one and no row is in both, including rows whose confidence is missing. A NaN confidence goes tolow, because “we do not know how sure the model was” is not evidence that the annotation is wrong, and quietly dropping those rows would make the two lists sum to less than the cell they came from — a discrepancy nobody notices until they try to reconcile the totals.Sorting is stable, so rows tied on confidence keep the table’s order and the split is reproducible run to run.