spacr.agreement

Workflow inputs and outputs

Annotator Agreement

Compare independent annotation columns and inspect discordant objects; agreement is separate from classification accuracy.

Open: Annotate → Annotator Agreement.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Training annotations — A chosen annotation column in measurements/measurements.db, table png_list; labels belong to object identities. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: prcfo.

  • Object crops — data/**/*_png when save_png is enabled; png_list indexes saved crops. Supported workflows can instead stream crops from merged arrays and masks. Relevant tables, depending on the route: png_list. Relevant columns, depending on the route: png_path, prcfo.

Outputs

  • Quality-control results — Stored project checks and QC reports; a missing check is not a passing result.

  • Figures and table exports — The output location chosen by the tool; exports describe the selected data and filters.

Before this module

  • Annotate: Choose independent annotation columns for the same objects.

API reference.

Module tutorial.

Multi-annotator agreement for spaCR annotation columns.

Two people label the same 3 000 crops. How much do they actually agree, and where do they disagree? This module answers both, reading nothing but the png_list table of a measurements.db.

Annotation columns are the ones the Annotate app writes: an INTEGER column added to png_list with ALTER TABLE … ADD COLUMN (see spacr.qt.annotate_engine.ensure_annotation_column()), holding the class the annotator pressed — 1 for a left click, 2 for a right click — and NULL for every crop they have not looked at yet. Two annotators means two such columns over the same png_path rows.

Public API

cohens_kappa(a, b)

κ for one pair of label vectors.

fleiss_kappa(matrix)

κ for three or more annotators, from a subjects × categories count matrix.

agreement_report(db_path, columns)

the whole picture for a database: per-pair κ, overall κ, per-class κ, confusion matrices, and the counts that make them interpretable.

disagreements(db_path, columns)

exactly the rows the annotators labelled differently, for review.

format_agreement(report)

the report as text.

Statistics that this module refuses to get wrong

An unlabelled cell is an abstention, not a disagreement. Every pair is scored on the rows both annotators labelled (pairwise complete cases). Rows where only one of them committed are counted and reported separately as n_abstained. Treating them as disagreements would silently deflate κ in proportion to how far behind the slower annotator is — which is a property of the calendar, not of the annotation.

κ has no value when there is no variance. κ = (pₒ − pₑ)/(1 − pₑ). If both annotators put every compared row in the same single class, pₑ = 1 and the denominator vanishes: the answer is undefined, not 1.0. If just one annotator used a single class, pₑ collapses onto the other’s marginal and κ is identically 0 no matter how well they agree. Both cases return nan with an explanation attached, because on a screen where 98 % of cells are negative they are the normal case, and a returned 0.0 or 1.0 there is a lie with a number on it.

Raw agreement is reported next to κ, always. 95 % agreement with κ ≈ 0 is the prevalence paradox (Feinstein & Cicchetti, 1990), not a broken annotator: when one class dominates, chance agreement is already almost as high as the observed agreement, so κ has almost no room left. Hiding either number hides half the story.

The interpretation bands are a convention. Landis & Koch (1977) “slight/fair/moderate/substantial/almost perfect” is a rule of thumb with no distributional basis. It is reported as a label, and labelled as a convention.

Nothing here imports torch, cellpose or any GPU stack — pandas, numpy and the standard library only — so the Qt screen can compute agreement without waking a 4-second import chain.

Classes

AgreementReport

Everything agreement_report() worked out, in one object.

PairAgreement

Cohen's κ for one pair of annotators, with everything needed to read it.

Functions

agreement_report(, labels)

Score how well two or more annotation columns agree.

annotation_columns(→ List[str])

Guess which columns of table hold human annotations.

cohens_kappa() → float)

Cohen's κ between two annotators' labels.

confusion_matrix() → pandas.DataFrame)

Return the a × b contingency table over rows both labelled.

disagreements(, complete_only, limit)

Return the rows the annotators labelled differently.

fleiss_kappa(→ float)

Fleiss' κ from a subjects × categories count matrix.

format_agreement(→ str)

Render an AgreementReport as a plain-text block.

interpret_kappa(→ str)

Return the Landis & Koch band for kappa (a convention).

kappa_detail(, name_a, name_b)

Compute Cohen's κ for a vs b and everything around it.

load_annotations() → pandas.DataFrame)

Read key + columns from table into a normalised frame.

table_columns(→ List[str])

Return the column names of table, in declaration order.

Module Contents

class spacr.agreement.AgreementReport[source]

Everything agreement_report() worked out, in one object.

Parameters:
  • db_path – path of the source annotation database.

  • table – source table that holds the annotation columns.

  • key – column that identifies each annotated row.

  • columns – annotator columns, in report order.

  • pairs – one PairAgreement per unordered column pair.

  • overall_kappa – Cohen’s κ for two annotators, Fleiss’ κ for three or more (computed on rows every annotator labelled).

  • overall_method – name of the κ statistic used for the overall value.

  • overall_note – interpretive caveat or reason the overall value is undefined; empty when no caveat applies.

  • interpretation – Landis–Koch convention label for overall_kappa.

  • labels – ordered class universe used throughout the report.

  • per_class – one row per class — its one-vs-rest κ, how often the annotators were unanimous on it, and its prevalence. This is where “we agree on the negatives, we argue about the positives” shows up.

  • n_rows – total annotation-table rows examined.

  • n_complete – rows every annotator labelled.

  • n_partial – rows some but not all labelled — abstentions, not disagreements.

  • n_unlabelled – rows none of the annotators labelled.

  • n_disagreements – rows where two annotators who both committed chose differently. This is the review queue’s length.

  • percent_agreement – fraction of complete rows with unanimous labels.

  • convention – named interpretation scale applied to κ values.

  • warnings – report-level caveats that must be shown to the reader.

kappa_table() → pandas.DataFrame[source]

The per-pair numbers as a DataFrame, ready to render.

pair(a: str, b: str) → PairAgreement | None[source]

Return the PairAgreement for two columns, either order.

Parameters:
  • a – name of either annotator column in the pair.

  • b – name of the other annotator column in the pair.

property defined: bool[source]

False when the overall κ is nan.

property n_annotators: int[source]

Return the number of annotation columns included in the report.

class spacr.agreement.PairAgreement[source]

Cohen’s κ for one pair of annotators, with everything needed to read it.

Parameters:
  • column_a – name of the first annotator column (the confusion-matrix row axis).

  • column_b – name of the second annotator column (the confusion-matrix column axis).

  • kappa – Cohen’s κ, or nan when it is undefined/degenerate — check defined before quoting it.

  • percent_agreement – raw pₒ, the fraction of compared rows where the two labels are identical. Always meaningful, even when κ is not.

  • expected_agreement – pₑ, agreement expected from the marginals alone.

  • n_compared – rows both annotators labelled — κ’s denominator.

  • n_agree – compared rows on which the two annotators agreed.

  • n_disagree – compared rows on which the two annotators disagreed.

  • n_abstained – rows exactly one of them labelled. Excluded from κ (an abstention is not a disagreement) and reported here instead.

  • n_neither – rows neither of them has reached yet.

  • labels – ordered class universe used for the confusion matrix.

  • confusion – a labels down the rows, b across the columns.

  • note – why κ is nan, or what to watch out for when it is not.

  • interpretation – convention label that explains the κ magnitude.

__str__() → str[source]

Return a readable one-line summary of this annotator pair.

The line includes both column names, signed three-decimal κ (or undefined), its interpretation, raw agreement, and compared-row count.

property defined: bool[source]

False when κ is nan — i.e. the data cannot support a κ.

spacr.agreement.agreement_report(db_path: str, columns: Sequence[str], table: str = PNG_TABLE, key: str = PNG_KEY, missing_values: Sequence[Any] = (), labels: Sequence[Any] | None = None) → AgreementReport[source]

Score how well two or more annotation columns agree.

Each pair is scored on the rows both of its annotators labelled. The overall κ for three or more annotators is Fleiss’ κ over the rows every annotator labelled, which is stricter — the report carries n_complete and n_partial so the difference is visible rather than silent.

Parameters:
  • db_path – path to measurements.db; opened read-only.

  • columns – two or more annotation columns of table.

  • table – table holding them (default png_list).

  • key – row key (default png_path).

  • missing_values – extra values to treat as abstentions, e.g. (0,) for legacy databases where 0 meant “not looked at”.

  • labels – optional fixed label universe.

Returns:

an AgreementReport.

Raises:

ValueError – for fewer than two distinct columns, or an unknown table/column.

spacr.agreement.annotation_columns(db_path: str, table: str = PNG_TABLE, key: str = PNG_KEY, max_classes: int = 20, min_labelled: int = 1, include_model_columns: bool = False) → List[str][source]

Guess which columns of table hold human annotations.

The Annotate app adds a plain INTEGER column per annotation pass, so an annotation column is one that is not part of the crop metadata, not written by a model, holds few distinct values, and has at least one non-NULL.

The model exclusion is the point. A classifier writes into this same table — pred/cv_predictions from the CV stage, predictions/ml_pred from the ML one — and its class column is indistinguishable by shape from an annotation pass. Offering it made agreement_report score the classifier as a third annotator, which is a different question with the same units: on a real database, four “annotators” gave κ = -0.004 where the two humans agree at 0.471. See _MODEL_COLUMNS.

Parameters:
  • db_path – path to measurements.db.

  • table – table to inspect (default png_list).

  • key – row key, always excluded.

  • max_classes – reject columns with more distinct values than this — a continuous measurement is not an annotation.

  • min_labelled – reject columns with fewer labelled rows.

  • include_model_columns – offer the model’s own columns too. For the deliberate question “how well does the classifier agree with the annotators?”, which is model validation, not inter-annotator agreement. Off by default because it is never the question somebody means when they ask for agreement between annotators.

Returns:

candidate column names, in table order.

spacr.agreement.cohens_kappa(a: Sequence[Any], b: Sequence[Any], labels: Sequence[Any] | None = None, missing_values: Sequence[Any] = ()) → float[source]

Cohen’s κ between two annotators’ labels.

Rows either annotator left unlabelled are dropped (abstention, not disagreement). Returns nan — never a flattering 0.0 or 1.0 — when the compared rows carry no variance; kappa_detail() gives the same number plus the reason.

Parameters:
  • a – first annotator’s labels.

  • b – second annotator’s labels, row-aligned with a.

  • labels – optional label universe.

  • missing_values – extra values that count as abstentions.

Returns:

κ in [-1, 1], or nan when undefined.

Raises:

ValueError – when the sequences differ in length.

spacr.agreement.confusion_matrix(a: Sequence[Any], b: Sequence[Any], labels: Sequence[Any] | None = None, missing_values: Sequence[Any] = ()) → pandas.DataFrame[source]

Return the a × b contingency table over rows both labelled.

Parameters:
  • a – first annotator’s labels.

  • b – second annotator’s labels, row-aligned with a.

  • labels – label universe; inferred from the data when omitted.

  • missing_values – extra values that count as abstentions.

Returns:

integer DataFrame, a labels on the index.

spacr.agreement.disagreements(db_path: str, columns: Sequence[str], table: str = PNG_TABLE, key: str = PNG_KEY, missing_values: Sequence[Any] = (), complete_only: bool = False, limit: int | None = None) → pandas.DataFrame[source]

Return the rows the annotators labelled differently.

A row is a disagreement when at least two annotators committed to a label and those labels are not all the same. A row where one annotator abstained is not a disagreement — by default it is scored from the available labels, and dropped entirely if that leaves fewer than two labels.

Parameters:
  • db_path – path to measurements.db; opened read-only.

  • columns – two or more annotation columns.

  • table – table holding them (default png_list).

  • key – row key (default png_path) — the crop to look at.

  • missing_values – extra values to treat as abstentions.

  • complete_only – when True, only consider rows every annotator labelled.

  • limit – cap the number of returned rows (the review list can be long); None returns all of them.

Returns:

DataFrame with key, one column per annotator, plus n_labelled and n_classes, in table order.

Raises:

ValueError – for fewer than two distinct columns, or an unknown table/column.

spacr.agreement.fleiss_kappa(matrix: Any) → float[source]

Fleiss’ κ from a subjects × categories count matrix.

matrix[i][j] is how many annotators put subject i in category j. Every row must sum to the same number of annotators n ≥ 2.

Fleiss’ κ generalises Scott’s π rather than Cohen’s κ: it pools the annotators into one marginal distribution instead of giving each its own, so on two annotators it agrees with Scott’s π and will usually differ slightly from cohens_kappa().

Parameters:

matrix – 2-D array-like of non-negative counts.

Returns:

κ, or nan when every rating fell in one category (no variance — pₑ=1, so κ has no denominator).

Raises:

ValueError – for a ragged/empty matrix, negative counts, rows that disagree on the number of annotators, or fewer than 2 annotators per subject.

spacr.agreement.format_agreement(report: AgreementReport) → str[source]

Render an AgreementReport as a plain-text block.

Every κ is printed next to its raw percent agreement and the number of rows behind it, so an undefined or paradoxical κ reads as information rather than as a missing value.

Parameters:

report – the report to render.

Returns:

multi-line text, no trailing newline.

spacr.agreement.interpret_kappa(kappa: float) → str[source]

Return the Landis & Koch band for kappa (a convention).

Parameters:

kappa – a κ value, possibly nan.

Returns:

band name, or "undefined" for nan.

spacr.agreement.kappa_detail(a: Sequence[Any], b: Sequence[Any], labels: Sequence[Any] | None = None, missing_values: Sequence[Any] = (), name_a: str = 'a', name_b: str = 'b') → PairAgreement[source]

Compute Cohen’s κ for a vs b and everything around it.

Only rows both annotators labelled enter the calculation. Rows where exactly one of them abstained are counted into PairAgreement.n_abstained; they are not disagreements.

Parameters:
  • a – first annotator’s labels (None/NaN/”” = abstention).

  • b – second annotator’s labels, row-aligned with a.

  • labels – label universe; inferred from the compared rows when omitted. Passing it keeps confusion matrices comparable across pairs.

  • missing_values – extra values that count as abstentions, e.g. (0,) for legacy databases where 0 meant “not looked at”.

  • name_a – name to record for the first column.

  • name_b – name to record for the second column.

Returns:

a PairAgreement.

Raises:

ValueError – when the two sequences differ in length.

spacr.agreement.load_annotations(db_path: str, columns: Sequence[str], table: str = PNG_TABLE, key: str = PNG_KEY, missing_values: Sequence[Any] = ()) → pandas.DataFrame[source]

Read key + columns from table into a normalised frame.

Labels come back normalised: integers where the value is integral, None for every abstention. Row order is the table’s own.

Parameters:
  • db_path – path to measurements.db (opened read-only).

  • columns – annotation columns to read.

  • table – table holding them (default png_list).

  • key – row-identifying column (default png_path).

  • missing_values – extra values to treat as abstentions.

Returns:

DataFrame with key first, then one column per annotator.

Raises:

ValueError – for an unknown table or column.

spacr.agreement.table_columns(db_path: str, table: str = PNG_TABLE) → List[str][source]

Return the column names of table, in declaration order.

Parameters:

db_path – path to the SQLite database, opened read-only.

Raises:

ValueError – when the database has no such table.

Nested helpers

kappa_detail._pack(kappa: float, p_o: float, p_e: float, note: str) → PairAgreement

Build one complete agreement result from the captured counts.

Parameters:
  • kappa – computed Cohen’s kappa, or NaN when it is undefined.

  • p_o – observed agreement fraction.

  • p_e – agreement fraction expected from the marginals.

  • note – explanation of the ordinary or degenerate result.

Returns:

a result containing the supplied statistics together with the captured annotator names, counts, labels, confusion matrix, and interpretation.

spacr/agreement.py:441