spacr.predictions

Write per-object classification scores back into a spaCR measurements.db.

Two classifiers in this package score every object of a whole database:

Both produce one row per crop, and both used to leave those rows in a CSV next to the model. A CSV is not where the rest of spaCR looks: the Annotate app, the active-learning queue, generate_image_umap, the plate heatmaps and every GUI table read the png_list table of measurements.db. So the scores have to land there, on the row they belong to.

Why this is its own module

The merge used to live in spacr.deep_spacr, which is an odd home for something spacr.ml needs just as much — importing deep_spacr pulls in torch and torchvision, which the classical-ML path has no use for. The obvious alternative, spacr.utils, already owns the database write helpers (_append_to_measurements_db, rename_columns_in_db), but it is also 8000 lines and imports most of the package; adding a third database concern there would make it harder, not easier, to see that CV and ML share one code path. This module is small, imports nothing from spaCR at module scope except spacr.utils helpers pulled in lazily, and is what both callers import.

The join key

prcfo — plate_row_column_field_object — is the canonical per-object identity in this codebase. spacr.utils.filepaths_to_database() writes it onto every png_list row, spacr.io._read_and_merge_data() indexes the merged feature frame by it, and spacr.ml.generate_ml_scores() already joins annotations onto features with it. Keying the merge on it means the CV and the ML stage land on the same row, which is the whole point of letting them coexist.

The previous implementation matched on os.path.basename(png_path). A basename is not an identity:

  • the tar handed to the model is built with arcname=os.path.basename(img_path) (spacr.utils.add_images_to_tar()), so a run over two source folders whose plates are both called plate1 — which is what happens whenever the plate name comes from the source folder name, see spacr.io _rename_and_organize_image_files — puts two different crops in the archive under one member name; and

  • the old lookup was a plain dict assignment, so the second of those two crops silently overwrote the first and one of the two plates was scored with the other plate’s predictions.

The second half of that is the real defect, and it is not fixed by changing the key: two crops that share a basename share a prcfo too. So the fix is to detect the collision. A key that arrives twice with two different values is recorded as ambiguous, written nowhere, and counted in the report. A wrong score is worse than a missing one.

Key selection is measured, not assumed. Candidate keys are built for both sides (prcfo, then the full png_path, then file_name — never a basename computed behind the caller’s back when a real column exists) and the one that actually matches the most rows wins, ties going to the earliest in KEY_PRIORITY. The chosen key and every count are printed, because a merge that matched 3 of 40000 rows used to look exactly like one that matched all of them.

Atomicity

Python’s sqlite3 opens an implicit transaction for DML only, so an ALTER TABLE runs in autocommit and lands immediately — the same trap spacr.utils.rename_columns_in_db() was fixed for. The transaction here is opened explicitly and rolled back on any error, so an interrupted merge leaves the table exactly as it was: no half-added column, no half-scored rows.

Classes

MergeReport

What one merge did, in numbers.

Functions

attach_predictions(objects, results, *[, ...])

Join prediction columns onto an object frame IN MEMORY.

crop_name_metadata(→ pandas.DataFrame)

Parse spaCR crop file names into the metadata prcfo is built from.

first_present(→ Optional[str])

The first of names that frame actually has, else None.

merge_cv_predictions(→ Optional[MergeReport])

Merge spacr.deep_spacr.apply_model_to_tar() results into table.

merge_ml_predictions(→ Optional[MergeReport])

Merge spacr.ml.ml_analysis() results into table.

merge_prediction_results(→ Optional[MergeReport])

Write a classifier's per-object results onto the rows of table.

migrate_prediction_columns(→ List[Tuple[str, str, int]])

Put a legacy prediction column back into the encoding it claims to be in.

Module Contents

class spacr.predictions.MergeReport[source]

What one merge did, in numbers.

Parameters:
  • table – database table into which prediction results were merged.

  • key – join-key strategy selected for the merge.

  • columns – prediction columns requested for insertion or update.

  • db_rows – target-table rows considered by the merge.

  • result_rows – incoming prediction rows considered by the merge.

  • matched_rows – target rows that received at least one prediction.

  • matched_keys – distinct incoming identities found in the target table.

  • unmatched_db_rows – target rows left unchanged because no result carried their identity.

  • unmatched_result_rows – parseable result rows whose identity was not present in the target table.

  • unparsed_result_rows – result rows from which no join identity could be constructed.

  • ambiguous_keys – identities repeated with conflicting prediction values and therefore deliberately not written.

  • ambiguous_result_rows – incoming rows involved in those conflicts.

  • fanout_rows – additional target rows sharing a matched identity and receiving the same value, such as alternate crops of one object.

  • repaired – legacy prediction columns repaired before this merge, as (table, column, rows_repaired) records.

  • added_columns – prediction columns newly created in the target table.

Returned by merge_prediction_results() and printed by it. Every count is here because a merge that matched three rows of forty thousand used to be indistinguishable from one that matched all of them.

__str__() → str[source]

Return the same human-readable report as summary().

Returns:

Multi-line prediction-merge summary.

summary() → str[source]

Return the human-readable multi-line report.

spacr.predictions.attach_predictions(objects, results, *, score_source: str = 'pred', class_source: str = 'cv_predictions', score_col: str = CV_SCORE_COLUMN, class_col: str = CV_CLASS_COLUMN, timelapse: bool = False)[source]

Join prediction columns onto an object frame IN MEMORY.

THE SAME JOIN AS merge_prediction_results(), AND NOTHING WRITTEN. _choose_key picks the key by MEASURING which one lands on the most rows, so a montage reading scores out of a score CSV and a database that had the same CSV merged into it cannot disagree about which object got which number – which two separate join implementations eventually would.

This supports projects whose png_list table has no prediction column but whose regression inputs already contain one score row per cell. The join uses those scores without changing the database.

Parameters:
  • objects – the per-object frame, e.g. png_list read back.

  • results – the score table – path, pred, cv_predictions, as process_vision_results and the regression module’s score CSVs carry.

Returns:

(frame, matched) – a COPY of objects with the score and class columns added where they joined, and how many rows matched. matched is 0 when nothing lined up, and the frame comes back without the columns, so a caller can refuse with a real number.

spacr.predictions.crop_name_metadata(names, timelapse: bool = False) → pandas.DataFrame[source]

Parse spaCR crop file names into the metadata prcfo is built from.

Every consumer of a classifier’s results has to answer the same question – which object is this crop? – and there is exactly one right way to answer it: spacr.utils._map_wells_png(), the parser spacr.utils.filepaths_to_database() used to write these very columns onto png_list. Re-deriving them with the writer’s own parser is what guarantees the two sides cannot drift apart.

It also recovers what a positional guess gets wrong. spacr.utils.process_vision_results() takes the object id as path.split('_')[3], which on a timelapse crop (plate_well_field_time_object) is the timepoint, not the object.

Parameters:
  • names – crop names or paths; only the basename is parsed.

  • timelapse – whether the names carry a timepoint component.

Returns:

a DataFrame aligned to names with plateID, rowID, columnID, fieldID, timeID (only when timelapse), object_label (the bare integer id, no 'o' prefix) and prcfo. A name that cannot be parsed gives None throughout.

spacr.predictions.first_present(frame, names) → str | None[source]

The first of names that frame actually has, else None.

spacr.predictions.merge_cv_predictions(df, db_path, table: str = PNG_TABLE, score_col: str = CV_SCORE_COLUMN, class_col: str = CV_CLASS_COLUMN, score_source: str = 'pred', class_source: str = 'cv_predictions', verbose: bool = True) → MergeReport | None[source]

Merge spacr.deep_spacr.apply_model_to_tar() results into table.

Parameters:
  • df – frame from apply_model_to_tar -> process_vision_results (path, pred, cv_predictions).

  • db_path – SQLite database to write into.

  • table – target table. Default 'png_list'.

  • score_col – database column for the probability.

  • class_col – database column for the thresholded class.

  • score_source – column of df the probability comes from.

  • class_source – column of df the class comes from.

  • verbose – print the report.

Returns:

a MergeReport, or None if the database is missing.

spacr.predictions.merge_ml_predictions(df, db_path, table: str = PNG_TABLE, score_col: str = ML_SCORE_COLUMN, class_col: str = ML_CLASS_COLUMN, verbose: bool = True) → MergeReport | None[source]

Merge spacr.ml.ml_analysis() results into table.

ml_analysis returns the predicted class in predictions and the per-class probabilities in prediction_probability_class_<i>; the positive-class probability is taken from class 1 when the model produced one, and skipped when it did not (a one-class fit) rather than inventing a column.

Parameters:
  • df – the scored frame – ml_analysis output [0].

  • db_path – SQLite database to write into.

  • table – target table. Default 'png_list'.

  • score_col – database column for the probability.

  • class_col – database column for the predicted class.

  • verbose – print the report.

Returns:

a MergeReport, or None if the database is missing, or if the frame carries no prediction column at all.

spacr.predictions.merge_prediction_results(results, db_path, columns, table: str = PNG_TABLE, key: str = 'auto', timelapse: bool | None = None, verbose: bool = True) → MergeReport | None[source]

Write a classifier’s per-object results onto the rows of table.

Shared by Classify (CV) and Classify (ML) — see the module docstring for why the key is prcfo, why an ambiguous key is refused rather than guessed at, and why the whole thing is one transaction.

Parameters:
  • results – DataFrame of per-object results. Must carry the source columns named in columns, plus something to key on: a prcfo column, an index named prcfo, or a path/png_path/ file_name column holding spaCR crop names.

  • db_path – SQLite or DuckDB file, Parquet store directory, or PostgreSQL connection string. A missing local store is reported and skipped.

  • columns – mapping of database column name to (results column, 'REAL' | 'INTEGER').

  • table – target table. Default 'png_list'.

  • key – 'auto' (measure every candidate and use the best) or one of KEY_PRIORITY to force it.

  • timelapse – whether crop names carry a timepoint. None detects it from the presence of a time column on the table.

  • verbose – print the report.

Returns:

a MergeReport, or None if a local store is missing.

Raises:
  • KeyError – if results lacks one of the source columns.

  • ValueError – if a Parquet table is missing or prediction output would replace a join-key column in an alternate store.

spacr.predictions.migrate_prediction_columns(db_path, table: str = PNG_TABLE, verbose: bool = True) → List[Tuple[str, str, int]][source]

Put a legacy prediction column back into the encoding it claims to be in.

spacr.utils.add_column_to_database() — the ML stage’s old write path — replaced every 0 with a 2 on the way into the database, because the Annotate app labels classes 1 and 2. Nothing else did that, so png_list.predictions said 2 where results.csv from the very same run said 0: one number, two meanings, depending on which file you opened. This repairs it, so a database written before the change reads correctly with no manual action — the same repair-on-read contract spacr.utils.rename_columns_in_db() has, and the same properties:

  • Idempotent. After the repair the column holds 0s and 1s, which is not the encoding this looks for, so a second pass does nothing.

  • Never destructive. The substitution is only reversed when the column holds nothing but 1s and 2s. A three-class model’s genuine class 2 is indistinguishable from a mangled 0, and guessing would be worse than leaving it exactly as it is.

  • All or nothing. SQLite, DuckDB and PostgreSQL repair in a transaction; Parquet publishes a complete replacement part snapshot. A half-repaired column would be a column in two encodings at once.

Parameters:
  • db_path – SQLite or DuckDB file, Parquet store directory, or PostgreSQL connection string. A missing local store is a no-op.

  • table – table to migrate. Default 'png_list'.

  • verbose – print each repair.

Returns:

list of (table, column, rows_repaired).

Nested helpers

crop_name_metadata.bare_object_label(value)

Normalize a parsed object label and remove one leading o.

spacr/predictions.py:455

crop_name_metadata.convert(value)

Parse and cache one crop basename, or return all-missing metadata.

spacr/predictions.py:441