spacr.infection

Infection metrics, assembled from a finished Measure run.

A REPORT, NOT A PIPELINE, and that was a deliberate choice made after the derivability was measured: thirteen of the sixteen candidate infection metrics fall out of tables Measure has already written, with no new image processing. The three that do not – intracellular fraction, vacuoles per cell, distance to the host boundary – belong to the invasion module and to segmentation, and are deliberately absent here rather than approximated.

THE DENOMINATOR IS THE WHOLE PROBLEM, and it is why every number this module produces carries its own denominator beside it. Almost every way of getting an infection metric wrong is a denominator that does not match its numerator:

  • The relationship table has NO ROW for a cell with no parasites. measure.get_components explodes the per-cell child list and drops the empty ones, so counting rows there counts INFECTED cells and calls it the cell count. The denominator has to come from the cell table, which holds every segmented cell whether or not anything is inside it.

  • io.py drops pathogens with no cell_id and can filter rows by pathogen_prcfo_count. Neither touches the cell table. A numerator built from a filtered pathogen table over an unfiltered cell table counts two different populations, and the ratio looks entirely plausible.

So infection_report() returns the denominator, and its size, on every row. A reader who disagrees with a number can see immediately which of the two halves they disagree with.

IT DOES NOT INVENT A FIFTH VOCABULARY. uninfected, pathogen_count and pathogen_prcfo_count already exist across measure, io and timelapse. This module reuses them and adds no synonym.

Functions

border_rules_agree(→ Optional[bool])

Whether cells and parasites were border-filtered the same way.

host_contrast(→ pandas.DataFrame)

One host measurement, infected against uninfected, IN THE SAME WELL.

infection_report(→ pandas.DataFrame)

Infection metrics per well, each with the denominator it was computed over.

multiplicity_distribution(→ pandas.DataFrame)

The histogram behind the means, which is what the means hide.

parasites_per_cell(→ pandas.DataFrame)

Every host cell, with how many parasites it contains -- ZERO INCLUDED.

uninfected_cells_were_measured(→ Optional[bool])

Whether the Measure run that wrote db_path kept uninfected cells.

write_infection_report(→ Optional[str])

Write the infection report beside the database it came from.

Module Contents

spacr.infection.border_rules_agree(db_path: str) → bool | None[source]

Whether cells and parasites were border-filtered the same way.

THE NUMERATOR AND THE DENOMINATOR MUST OBEY THE SAME RULE, which is what 377’s PART 1 asks for in as many words: “the infection denominator has to use the same rule or the rate is computed against a different cell count than the numerator”.

remove_border_objects is applied at SEGMENTATION time, per object type, so it decides what is in each table rather than how the table is counted. The two settings are independent and default to False together, so they agree unless somebody changed one:

  • pathogens cleared and cells kept – a host cell whose only parasite touched the edge is still in the cell table, now with a count of zero. It reads as UNINFECTED and deflates the rate.

  • cells cleared and pathogens kept – a parasite can name a cell that is no longer in the cell table, so it is counted in neither numerator nor denominator and the infection index drifts.

THIS IS THE SAME FAILURE AS include_uninfected, WHICH ALREADY SHIPPED: a setting that silently changes what a table contains, and therefore what a ratio over it means, with nothing in the output saying so. That one turned every well into 1.000000 on a real plate. This one is quieter, which is worse – a deflated rate looks like biology.

Parameters:

db_path – path to a measurements.db.

Returns:

True when both rules match, False when they differ, None when the run recorded neither – an older database, where the honest answer is that we cannot tell.

spacr.infection.host_contrast(db_path: str, columns: Sequence[str], *, statistic: str = 'mean') → pandas.DataFrame[source]

One host measurement, infected against uninfected, IN THE SAME WELL.

The comparison is made within a well on purpose: comparing an infected well to an uninfected one confounds infection with everything else that differs between two wells, and the whole reason this is worth computing is that both populations sit in the same one.

Parameters:
  • db_path – path to a measurements.db.

  • columns – the host-cell measurements to contrast, e.g. ("area", "eccentricity").

  • statistic – any name DataFrameGroupBy.agg accepts.

Returns:

one row per (well, measurement) with the infected and uninfected values and both group sizes.

spacr.infection.infection_report(db_path: str, *, by_field: bool = False, monolayer_filter: bool = False) → pandas.DataFrame[source]

Infection metrics per well, each with the denominator it was computed over.

EVERY ROW CARRIES ITS DENOMINATOR because that is where these numbers go wrong. infection rate over “cells” and parasites per infected over “infected cells” are different populations, and a table that reports only the ratios cannot be checked.

Parameters:
  • db_path – path to a measurements.db.

  • by_field – group by field as well as well, which is what shows a settling gradient a well mean hides.

Returns:

tidy frame – identity columns, then metric, value, denominator and n_denominator. Empty when there is no cell table to count.

spacr.infection.multiplicity_distribution(db_path: str) → pandas.DataFrame[source]

The histogram behind the means, which is what the means hide.

Two conditions can share an infection index and differ completely: a few heavily infected cells against many lightly infected ones is different biology with the same average. The distribution is the number that distinguishes them, so it is reported rather than summarised.

Parameters:

db_path – path to a measurements.db.

Returns:

one row per (well, parasite count) with cells and the fraction of the well’s cells at that count.

spacr.infection.parasites_per_cell(db_path: str) → pandas.DataFrame[source]

Every host cell, with how many parasites it contains – ZERO INCLUDED.

THE ZEROES ARE THE POINT. The relationship between a cell and its parasites is stored on the parasite, so a cell containing none appears nowhere in the pathogen table. Any count taken from that table alone is a count of infected cells wearing the name of a cell count.

Parameters:

db_path – path to a measurements.db.

Returns:

one row per cell, with the identity columns, object_label and pathogen_count. Empty when there is no cell table.

spacr.infection.uninfected_cells_were_measured(db_path: str) → bool | None[source]

Whether the Measure run that wrote db_path kept uninfected cells.

THE ANSWER CHANGES WHAT THE CELL TABLE IS. include_uninfected=False – the default for a screen that only crops infected cells, and what the TSG101 plates were measured with – means Measure wrote no row for a cell with no parasite. The cell table is then a selection of infected cells rather than the segmented population, and infected / cells is 1.0 for every well no matter what the biology did.

That is not a rounding problem to note in a docstring. A reader handed a column of 1.000 reads 100% infection as a RESULT, and nothing in the table says otherwise, so this module refuses the rate instead of printing it.

Parameters:

db_path – path to a measurements.db.

Returns:

True or False when the run recorded the setting, None when it did not – an older database, where the honest answer is that we cannot tell.

spacr.infection.write_infection_report(db_path: str, *, by_field: bool = False, destination: str | None = None) → str | None[source]

Write the infection report beside the database it came from.

A MEASURE RUN EMITS THE REPORT, so anyone who has measured a plate already has it. There is no button and no screen, and nothing has to be asked for – which is the point, because the runs that most need these numbers are the ones that would never have thought to ask.

NOTHING IS WRITTEN WHEN THERE IS NOTHING TO SAY. A plate with no cell table, or none of the columns the metrics need, gives an empty report, and an empty CSV beside a database is a file that invites somebody to wonder what went wrong. The answer is None instead.

Written to a dot-name in the same folder and renamed over the target, so a reader never sees half a file.

Parameters:
  • db_path – a measurements.db.

  • by_field – group by field as well as well.

  • destination – where to write it; the default is REPORT_NAME beside db_path.

Returns:

the path written, or None when the report is empty.

Nested helpers

infection_report.add(metric, value, denominator, n)

Record one metric together with the population it is over.

The denominator travels with the value rather than being implied by the metric’s name, because two of these are ratios over different populations and a reader cannot check a ratio whose denominator they have to guess.

Parameters:
  • metric – the metric’s name.

  • value – its value for this group.

  • denominator – what the value was computed over, in words.

  • n – how many things that denominator contained.

spacr/infection.py:418