spacr.gene_measurement_sweep

Every gene against every measurement, corrected, in one matrix operation.

The question is “which genes move which measurements”, asked of the whole screen at once rather than one gene at a time.

IT IS ONE MATMUL. A representative screen with 1,376 guides x 785 measurements over 1,366 wells, computed in 0.1 seconds. A loop over a million regressions answers the same question in a morning; anything here that looks like one is a bug.

THE THREE CORRECTIONS ARE THE POINT OF THE MODULE. Each was found by hand while sweeping ONE gene, in about ten minutes, and each would be made again by anyone doing this themselves:

  • IDENTIFIERS ARE NOT MEASUREMENTS. pathogen_object_label came out at p=2.5e-07 for EAF1 and it is a LABEL – the cells picked out sit lower in the segmentation’s numbering, which is position in a list. pathogen_pathogen is the same column under another name (spearman 0.9979, identical in 141,626 of 226,467 rows). Both were nearly reported.

  • 785 TESTS IS NOT ONE TEST. pathogen_solidity at p=1.8e-03 looked interesting and does not survive Benjamini-Hochberg.

  • CIRCULARITY IS A COLUMN. The classification score is a function of the image, so a measurement it already tracks cannot corroborate anything derived from it. On this screen spearman(pred, pathogen_channel_1_mean_intensity) = -0.389, and the “strongest result” for GRA14 was exactly that measurement, in exactly the direction the correlation predicts.

Classes

HOUSE

Define spaCR's scientific-figure palette and sizing constants.

SweepResult

The grid, and the tidy table a reader actually looks at.

Functions

gene_fractions(→ pandas.DataFrame)

A gene's fraction in each well: the SUM of its guides' fractions.

gene_of_guide(→ Optional[str])

The gene a GUIDE NAME belongs to, for both spellings in use.

is_measurement(→ bool)

Whether name measures the object rather than naming it.

measurement_columns(→ List[str])

Every numeric column of frame that is a measurement.

measurement_family(→ str)

Which family name belongs to. See MEASUREMENT_FAMILIES.

plot_calibration(result[, path, title, level])

#10 -- is this screen calibrated, or is everything significant?

plot_circularity(result[, path, alpha, title, level])

Plot whether each surviving measurement adds independent evidence.

plot_effect_against_representation(result[, path, ...])

Plot effect counts against each gene's effective representation.

plot_gene_profile(result, gene[, path, alpha, top, title])

#6 -- ONE gene's fingerprint: every measurement it moves, in order.

plot_gene_similarity(result[, path, alpha, top, ...])

#7 -- which genes behave ALIKE, by correlating their whole profiles.

plot_grid_volcano(result[, path, alpha, title, level])

#5 -- every gene x measurement pair at once: effect against evidence.

plot_guide_concordance(result[, path, alpha, top, title])

Do a gene's own guides agree with each other?

plot_measurement_families(result[, path, alpha, top, ...])

What KIND of thing each gene moves, as a stacked bar per gene.

plot_measurement_hits(result[, path, alpha, top, ...])

#8 -- which MEASUREMENTS are informative, and which everything moves.

plot_sweep(result[, path, alpha, max_circularity, ...])

A heatmap of what SURVIVED, clustered so related things sit together.

sweep(→ SweepResult)

Associate every guide with every measurement, blocked and corrected.

Module Contents

class spacr.gene_measurement_sweep.HOUSE[source]

Define spaCR’s scientific-figure palette and sizing constants.

The palette follows recurring colors in published apicomplexan-genomics figures, including Waldman et al. (Cell, 2020; Figures 1 and 3) and Giuliano et al. (Nature Microbiology, 2024; Figure 1). Grey represents background data, while accent colors identify the smaller subset being interpreted. Family colors remain stable across panels so readers do not have to relearn category mappings.

Text, spines, and ticks use theme-aware colors selected by _readable() rather than this palette, which keeps figures legible in spaCR’s dark interface while retaining consistent scientific data colors.

class spacr.gene_measurement_sweep.SweepResult[source]

The grid, and the tidy table a reader actually looks at.

Parameters:
  • table – tidy result frame with one row per fitted guide or gene-measurement pair and its effect, significance, support, and control evidence.

  • effects – dense blocked association matrix indexed by fitted guide or gene and columned by measurement; these effects are repeated in table.

  • n_wells – number of well identities shared by wells and fractions before guide-specific filtering; unlike table['n_wells'], this is screen-wide.

  • n_blocks – number of distinct block labels removed from a fitted design, or zero when the sweep returned before fitting.

  • dropped – numeric wells columns omitted from the effect grid, including automatically rejected identifier or duplicate columns and explicitly dropped measurements.

  • circularity_known – whether at least three joined score values produced a finite measurement-to-score circularity; when false, circularity is NaN and survivors() refuses a circularity cutoff.

describe() → str[source]

Summarize grid size, significance, circularity, and omitted inputs.

survivors(alpha: float = 0.05, max_circularity: float = 1.0) → pandas.DataFrame[source]

Rows past the correction, optionally past a circularity bar too.

Raises:

ValueError – a circularity bar was asked for and the score never joined to the wells. Filtering on a column of NaN returns nothing and looks like a result.

spacr.gene_measurement_sweep.gene_fractions(fractions: pandas.DataFrame, gene_of: Any | None = None) → pandas.DataFrame[source]

A gene’s fraction in each well: the SUM of its guides’ fractions.

The same rule the regression applies – wells_for_coefficient with guide_aggregation='sum' – because “does this GENE move this measurement” must not be a different arithmetic from the fit that found the gene in the first place.

Parameters:
  • fractions – well-by-guide fraction matrix. Columns assigned to the same gene are summed row-wise.

  • gene_of – guide name -> gene id. Defaults to spacr.hits.gene_of(), which is the key the metadata join uses, so the two cannot disagree about which gene a guide belongs to.

Returns:

one column per gene. Guides that name no gene are left out rather than pooled into an “unknown” gene that no experiment ran.

spacr.gene_measurement_sweep.gene_of_guide(guide: Any, prefix: str | None = None) → str | None[source]

The gene a GUIDE NAME belongs to, for both spellings in use.

spacr.hits.gene_of reads a DESIGN TERM – fraction:grna[225160_1] – and truncates at the first underscore. Handed the bare TGGT1_225160_2 that a count table actually carries, that rule returns TGGT1: the organism, for every guide in the screen, which pools the entire library into one “gene”. So a bracketed term still goes to hits.gene_of.

A BARE NAME IS READ BY ITS SHAPE. process_reads splits <organism>_ on exactly three components, so a three-part name gives up its gene as the middle one and a two-part name as the first. That works for PF3D7_0100100_1 and for a human library without either being named anywhere.

A PREFIX THAT WAS REMOVED IS NOT REMOVED TWICE. Once an organism has been taken off the front – the measured prefix, or one of GUIDE_PREFIXES – the remainder is <gene>_, so the gene is everything but the guide number. TGGT1_ROP18_kinase_1 is ROP18_kinase, not kinase.

Parameters:
  • guide – bare guide name or bracketed regression design term to parse.

  • prefix – an organism prefix MEASURED from the library, as spacr.control_names.common_prefix() returns. Given, it is removed first, and the remainder read as <gene>_. This is how a caller that has the whole library in hand beats a rule that only has one name.

spacr.gene_measurement_sweep.is_measurement(name: str) → bool[source]

Whether name measures the object rather than naming it.

Parameters:

name – candidate column name to classify.

spacr.gene_measurement_sweep.measurement_columns(frame: pandas.DataFrame) → List[str][source]

Every numeric column of frame that is a measurement.

Parameters:

frame – object-measurement table whose numeric columns are screened.

A COLUMN THAT DUPLICATES AN IDENTIFIER IS ALSO OUT, whatever it is called. pathogen_pathogen passes the name test and is the object label to four decimal places; the only way to catch it is to look.

spacr.gene_measurement_sweep.measurement_family(name: Any) → str[source]

Which family name belongs to. See MEASUREMENT_FAMILIES.

spacr.gene_measurement_sweep.plot_calibration(result: SweepResult, path: str | None = None, *, title: str = '', level: str | None = None)[source]

#10 – is this screen calibrated, or is everything significant?

The observed P values against the uniform they would follow if nothing were real. A grid that hugs the diagonal has no signal; one that lifts off it at the left has some; one that lifts off everywhere has a systematic effect – a plate term that did not get blocked out, or an aggregation that correlated every well with itself.

THE FIRST PLOT TO LOOK AT and the last one anybody builds. It says whether the other nine are worth reading at all.

spacr.gene_measurement_sweep.plot_circularity(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, title: str = '', level: str | None = None)[source]

Plot whether each surviving measurement adds independent evidence.

A measurement the classifier already tracks cannot corroborate a result derived from that classifier. The plot exposes that dependence rather than presenting a correlated measurement as separate confirmation.

Every point is drawn because the correlation cutoff is a judgment. This lets users see where their hits fall and choose a threshold appropriate to their screen.

spacr.gene_measurement_sweep.plot_effect_against_representation(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, title: str = '', level: str | None = None)[source]

Plot effect counts against each gene’s effective representation.

The plot exposes representation as a possible confound without modifying the underlying statistic. It shows whether a gene’s number of significant measurements follows the overall representation trend or departs from it.

x is the gene’s EFFECTIVE WELL COUNT – the participation ratio, which is literally the sample size each p-value was computed on – and y is how many measurements it moved past the correction. A gene high on the trend line is doing what its statistical weight predicts; a gene ABOVE the line is the interesting one, and a gene at the far right with a huge count is exactly the one to be suspicious of.

EFFECTIVE WELLS AND NOT share, WHICH MEASURES SOMETHING ELSE. share is the median fraction a gene takes of the wells it is IN – how concentrated it is when present – and a rare gene can score high on it precisely by being rare. Measured on this module’s own fixture: a gene in 18 of 120 wells has share 0.44 and one in all 120 has 0.20, which is the opposite of the ordering the reader is asking about. What drives the ranking is POWER, and the participation ratio is the number the test actually used.

Controls are drawn with a distinct marker because their relationship to the trend provides a useful assay calibration.

Returns:

the matplotlib Figure, or None when nothing survived.

spacr.gene_measurement_sweep.plot_gene_profile(result: SweepResult, gene: Any, path: str | None = None, *, alpha: float = 0.05, top: int = 24, title: str = '')[source]

#6 – ONE gene’s fingerprint: every measurement it moves, in order.

The per-gene readout the grid cannot give. “GRA14 moves pathogen intensity up and pathogen area down, and nothing else” is a sentence somebody can take to a bench; a row of a heatmap is not.

Bars are coloured by measurement family, so a profile that is all one family reads as one finding rather than as twenty.

spacr.gene_measurement_sweep.plot_gene_similarity(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, top: int = 30, title: str = '', level: str | None = None)[source]

#7 – which genes behave ALIKE, by correlating their whole profiles.

Two genes in one pathway should move the same measurements the same way, and this is the only view here that can say so: every other one reads a gene on its own. It is also the honest way to ask “is my hit list one finding or twelve”.

Correlated across the WHOLE effect row, not just the significant part – a shared sub-threshold pattern is exactly the evidence that two genes belong together, and thresholding first would throw it away.

spacr.gene_measurement_sweep.plot_grid_volcano(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, title: str = '', level: str | None = None)[source]

#5 – every gene x measurement pair at once: effect against evidence.

THE SHAPE OF THE WHOLE GRID, which the heatmap cannot show because it draws only survivors. A screen where everything is significant looks different here from one with a handful of real effects, and that difference is the first thing to check before reading any single row.

Colour is CIRCULARITY where it is known – a hit the classifier already tracks is a restatement, not a corroboration – and grey where it is not. Grey is not “clean”: the sweep says so in the legend rather than letting an uncomputed number read as zero.

spacr.gene_measurement_sweep.plot_guide_concordance(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, top: int = 20, title: str = '')[source]

Do a gene’s own guides agree with each other?

THE ONLY INTERNAL CONTROL THIS DESIGN HAS. Two guides against the same gene are two independent perturbations of it, so a gene whose guides agree on the sign of an effect is saying something the gene’s own biology explains, and a gene whose guides disagree is saying something about the guides.

Needs a table with GUIDE rows: sweep(..., level='guide') or 'both'. A gene-level table has nothing to compare, which is a reason to say so rather than to draw an empty axis.

Returns:

the matplotlib Figure, or None when there is nothing to compare.

spacr.gene_measurement_sweep.plot_measurement_families(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, top: int = 14, title: str = '', level: str | None = None)[source]

What KIND of thing each gene moves, as a stacked bar per gene.

767 measurements is not a list a reader can hold, but six families is. “this gene moves pathogen intensity and nothing else” is a sentence about biology; “this gene has 41 significant measurements” is not.

The families are coarse on purpose – see MEASUREMENT_FAMILIES.

Parameters:
  • result – completed guide/measurement sweep whose significant rows are grouped into measurement families.

  • top – how many genes to draw, most-hits first.

Returns:

the matplotlib Figure, or None when nothing survived.

spacr.gene_measurement_sweep.plot_measurement_hits(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, top: int = 26, title: str = '', level: str | None = None)[source]

#8 – which MEASUREMENTS are informative, and which everything moves.

The grid read down its other axis. A measurement moved by half the library is not a discriminating readout: it is a plate effect, a focus drift or a confluence artefact wearing a measurement’s name, and it will put a hit on every gene in the screen. Ranking measurements by how many genes move them is how you find those before trusting any of them.

spacr.gene_measurement_sweep.plot_sweep(result: SweepResult, path: str | None = None, *, alpha: float = 0.05, max_circularity: float = 1.0, top: int = 40, title: str = '', level: str | None = None)[source]

A heatmap of what SURVIVED, clustered so related things sit together.

THE WHOLE GRID IS NOT A PICTURE. 1,240 guides x 767 measurements is 951,080 cells; drawn, it is a texture, and every one of them is coloured whether or not it means anything. So the default view is the survivors – the guides and measurements with at least one entry past the correction – and everything else is a filter away.

Parameters:
  • result – what sweep() returned.

  • top – the most guides and measurements to draw. A screen with hundreds of survivors is a table, not a picture, and saying so beats drawing something illegible.

Returns:

the matplotlib Figure, or None when nothing survived.

spacr.gene_measurement_sweep.sweep(wells: pandas.DataFrame, fractions: pandas.DataFrame, *, blocks: Sequence | None = None, scores: Sequence[float] | None = None, alpha: float = 0.05, controls: Sequence[str] | None = None, min_wells: int = 5, measurements: Iterable[str] | None = None, drop_measurements: Iterable[str] | None = None, drop_guides: Iterable[str] | None = None, max_share: float | None = None, max_wells_fraction: float | None = None, level: str = 'guide') → SweepResult[source]

Associate every guide with every measurement, blocked and corrected.

Parameters:
  • wells – one row per well, the measurements in its columns.

  • fractions – one row per well, one column per guide, holding that guide’s fraction. Aligned to wells by index.

  • blocks – the plate of each well. Absent means one block, which is honest but weaker: a plate difference can then look like a gene.

  • scores – the per-well classification score, used ONLY to compute each measurement’s circularity. Absent leaves that column at 0 and the caller must not read it as “not circular”.

  • controls – guide or gene names to MARK as controls. They are not removed: the regression drops the control COLUMNS of the plate because they are not part of the contrast it fits, but this asks a different question – whether a gene moves a measurement – and a control is exactly the thing whose answer you want to see. Marked so a reader can find them, never filtered away behind their back.

  • drop_measurements – columns to leave out, by name. The complement of measurements: naming the two or three that are wrong is easier than listing the seven hundred that are not.

  • drop_guides – guides or genes to leave out, by name. Matched at BOTH levels – a screen swept at gene level names genes and one at guide level names guides, and a user typing a gene id should not have to know which they are looking at.

  • max_share – drop a guide whose median well fraction, where it is present, is above this.

  • max_wells_fraction – drop a guide present in more than this share of wells. This prevalence filter is separate from max_share, which measures median abundance where the guide is present. A guide can be common at low abundance or rare at high abundance, so the two filters identify different forms of over-representation.

  • min_wells – guides present in fewer wells than this are dropped – a correlation over three wells is not an effect.

  • level – 'guide', 'gene', or 'both'. A gene’s fraction in a well is the SUM of its guides’ – the same rule the regression applies, because “does this GENE move this measurement” must not be a different arithmetic from the fit that found the gene.

Returns:

a SweepResult. With 'both' the table carries a level column and the guide rows stay reachable beside the gene ones.

Nested helpers

_order_like_neighbours.order(matrix)

Order one matrix axis by correlation-cluster neighbours.

Parameters:

matrix – numeric two-dimensional rows to arrange.

Returns:

identity indices below three rows; otherwise average- linkage leaf indices from correlation distances. Non-finite matrix values are normalized for distance calculation and undefined distances become 1.0. Any clustering failure is handled by the parent’s mean-order fallback.

spacr/gene_measurement_sweep.py:781