spacr.qt.widgets.power_design

The design arithmetic behind the Power / Design screen — no Qt in here.

spacr.power_simulate and spacr.power_model answer the question “given these sixteen parameters, how well does the model recover the hits?”. The question a screener actually asks is “how many cells per well, and how many wells, do I need to detect an effect of size X?”. This module is the translation between the two, and it is deliberately separate from the widget so the translation can be tested without a QApplication — the same split pca_model / pca and pivot_spec / tabulate already use.

What lives here:

  • DesignSpec — the experiment as a screener describes it (genes, guides, wells, plates, effect size, prevalence, depth), plus the real-screen values for everything the screener does not want to name.

  • simulator_kwargs() — that description as the keyword arguments spacr.power_model.scan_parameters() takes. This is the only place the translation happens. The screen never builds simulator arguments of its own and never computes a metric of its own; it calls the library and renders what comes back.

  • cells_grid() / wells_grid() — the two sweep axes, bracketing whatever the user typed rather than a fixed grid, so the curve always has the user’s own design on it.

  • power_curve() — replicates collapsed into a detection probability.

  • plain_sentence() — the one sentence the whole screen exists to print.

  • CAVEATS — the port’s documented departures from spaCRPower, as data, so the screen can render the ones that change the number next to the number instead of leaving them in a docstring nobody opens.

Where the defaults come from

Every default in DesignSpec is the fitted value for the real T. gondii screen this simulator was built against, as recorded in proposals/SIM_PORT_PLAN.md §1 and §10 and in the port’s own docstrings: 452 genes at ~4 guides each, spotted into 4 x 384 wells at ~4.6 constructs per well, ~123 cells imaged per well (var ~8000), a MaxViT classifier at 0.80 / 0.12, ~2.5 % of genes true hits, ~3e4 reads per well. That makes the screen open on a working example rather than on filler, and it means the first number a user sees is one they can check against a screen that has actually been run.

What “detection” means here, exactly

A power analysis needs a yes/no per simulated screen before it can report a percentage. The model does not emit a yes/no — it emits a ranking of genes, scored by AUROC and average precision. So the screen asks the user for the bar (detection_auroc, default 0.80) and reports the fraction of simulated screens that clear it. Two things about that fraction are load-bearing:

  1. A replicate whose fit failed or did not converge counts as a non-detection, not as a missing value. Dropping it would raise the reported power by removing exactly the runs where the design was too thin to fit — which is the failure mode the analysis is supposed to find.

  2. The mean AUROC is reported beside the power, over the replicates that did converge, with the count of those that did not. A power of 0/5 with five non-converged fits and a power of 0/5 with five converged fits at AUROC 0.52 are different findings.

Classes

Caveat

One documented departure or limitation, ready to render.

DesignSpec

A screen as its designer describes it, with real-screen defaults.

Functions

cells_grid(→ List[float])

The cells-per-well sweep for spec.

changes_the_number(→ Tuple[Caveat, ...])

The caveats that belong next to the headline number.

estimate_runtime_s(→ float)

Very rough seconds for the two sweeps, for a "this will take…" line.

plain_sentence(→ str)

The one sentence the screen exists to print.

power_curve(→ pandas.DataFrame)

Collapse a scan_parameters() frame into a curve.

simulator_kwargs(→ Dict[str, Any])

The design as keyword arguments for the simulator, all held fixed.

wells_grid(→ List[float])

The wells sweep for spec.

Module Contents

class spacr.qt.widgets.power_design.Caveat[source]

One documented departure or limitation, ready to render.

Variables:
  • key – stable identifier, for tests and for the export.

  • headline – the one line that goes on screen next to the number.

  • detail – the paragraph behind it.

  • changes_the_number – whether believing the opposite would change the power the screen reports. Only these are rendered next to the headline sentence; the rest go in the list below it. The distinction is the whole point of the panel — a caveat list where everything is equally urgent is read as a disclaimer and skipped.

class spacr.qt.widgets.power_design.DesignSpec[source]

A screen as its designer describes it, with real-screen defaults.

The first block is what the screen puts on the form. The second block is everything the simulator needs that a screener does not want to name; they are fields rather than constants so a test — or a later advanced panel — can vary them, and so the exported record of a run carries every number that went into it.

Variables:
  • n_genes – genes in the library.

  • n_grnas_per_gene – guides per gene. Only reaches the simulation when score_per is "guide"; see the no_guide_level_model caveat.

  • score_per – "gene" (guides pooled, what the real analysis and spaCRPower do) or "guide" (each construct is its own library unit).

  • cells_per_well – mean cells imaged per well — the microscope time.

  • wells_per_plate – plate format.

  • n_plates – plates in the screen.

  • constructs_per_well – mean library units spotted into each well. This is well_abundance_factor_mu, and it means the same number of constructs per well at any library size because gene abundances sum to 1.

  • background_positive_rate – probability a non-hit cell is called positive — the classifier’s false-positive rate.

  • effect_fold – how many times more often a hit-genotype cell is called positive. The effect size. 0.80 / 0.12 = 6.667 in the real screen.

  • hit_rate – fraction of library units that are true hits.

  • reads_per_well – mean sequencing reads per well.

  • gene_abundance_alpha – Dirichlet concentration on library abundance; small is skewed.

  • cells_per_well_var – variance of the imaged cell count per well.

  • class_pos_var – variance of the hit-cell positive rate.

  • class_neg_var – variance of the background positive rate.

  • well_abundance_var – variance of the per-well abundance factor.

  • sequencing_cells_per_well – cells per gene per well contributing DNA — far more than are imaged, because sequencing sees the whole well.

  • pcr_factor_mu – log-scale mean of the per-well amplification.

  • pcr_factor_var – log-scale variance of it.

  • read_depth_cv – coefficient of variation of depth between wells.

  • sequencing_error_rate – probability a barcode read is credited to the wrong gene. 0.0 reproduces spaCRPower and every power figure this screen has ever printed; see the sequencing_error_hides_untested_genes caveat for why the number that moves is not the one you would expect.

  • min_cells_per_well – wells with fewer imaged cells than this are dropped before the fit. 0 keeps every well, which is what both packages did; see the thin_wells_count_the_same_as_full_ones caveat.

  • imaging_split – "abundance" or "uniform"; see the even_split_overstates_power caveat.

  • n_replicates – simulated screens per grid point.

  • detection_auroc – the AUROC a replicate must reach to count as a detection.

  • seed – master seed used to derive each grid-point replicate seed. Reproduction also requires the complete design, the same sweep grid and order, resolved backend, and software stack.

  • backend – inference backend, passed straight to spacr.power_model.scan_parameters().

validate() → List[str][source]

Return validation messages for parameters that prevent simulation.

Returns:

list of str – User-facing problems. An empty list means the design is valid.

with_values(**changes: Any) → DesignSpec[source]

Return a copy with changes applied. Thin wrapper on replace.

property expected_hits: float[source]

Expected number of true hits in the library.

property hit_positive_rate: float[source]

Probability a hit-genotype cell is called positive.

The effect size applied to the background rate. Capped just below 1 because a rate of exactly 1 is only a realisable beta at zero variance, and silently swapping the requested spread for zero is the kind of substitution spacr.power_simulate.rbeta_mean_variance() refuses to make.

property n_library_units: int[source]

Rows of the library the model actually estimates a coefficient for.

Genes when guides are pooled, constructs when they are not. This is what reaches n_genes_in_library, and it is the p in the p >> n the horseshoe prior exists to handle.

property n_wells: int[source]

Total wells in the screen.

spacr.qt.widgets.power_design.cells_grid(spec: DesignSpec) → List[float][source]

The cells-per-well sweep for spec.

Parameters:

spec – the design.

Returns:

sorted cells-per-well values, including the design’s own.

spacr.qt.widgets.power_design.changes_the_number() → Tuple[Caveat, ...][source]

The caveats that belong next to the headline number.

Returns:

the subset of CAVEATS whose changes_the_number is set, in declaration order.

spacr.qt.widgets.power_design.estimate_runtime_s(spec: DesignSpec) → float[source]

Very rough seconds for the two sweeps, for a “this will take…” line.

Scales SECONDS_PER_FIT_REFERENCE by design size relative to the real screen it was measured on, then multiplies by the number of fits. Deliberately crude: it exists so a user is not surprised by a five-minute wait, and the screen labels it as an estimate.

Parameters:

spec – the design.

Returns:

estimated wall-clock seconds.

spacr.qt.widgets.power_design.plain_sentence(spec: DesignSpec, cells: pandas.DataFrame | None, wells: pandas.DataFrame | None = None) → str[source]

The one sentence the screen exists to print.

Reads the user’s OWN design off the curve — the point at spec.cells_per_well — rather than the best point on it, because the question was “is my design enough”, not “what is the best design in this grid”. If the run did not reach that point, says so instead of interpolating.

Parameters:
  • spec – the design the sweep was run for.

  • cells – the cells-per-well curve from power_curve().

  • wells – the wells curve, used only for the follow-on clause about what a different well count would buy.

Returns:

one or two sentences of plain English.

spacr.qt.widgets.power_design.power_curve(scan: pandas.DataFrame, sweep_column: str, detection_auroc: float) → pandas.DataFrame[source]

Collapse a scan_parameters() frame into a curve.

One row per swept value. power is the fraction of that point’s replicates whose fit both succeeded and reached detection_auroc:

  • the denominator is EVERY replicate at the point, including the ones that failed or did not converge;

  • mean_auroc and mean_ap average only the replicates that did converge, so they answer “how well did it do when it worked” — and n_not_converged / n_failed sit beside them so that question is never mistaken for the first one.

Parameters:
  • scan – the frame scan_parameters returned. Needs the swept column plus status, model_auroc, model_ap and ap_baseline.

  • sweep_column – the simulator parameter that was swept.

  • detection_auroc – the bar a replicate must clear to count.

Returns:

pandas.DataFrame with POWER_CURVE_COLUMNS, sorted by value. Empty input gives an empty frame with those columns rather than a KeyError.

Raises:

KeyError – if sweep_column or any of status, model_auroc, model_ap, ap_baseline is missing. Named rather than defaulted: a frame without status would otherwise be summarised as if every replicate had succeeded, which turns a mis-wired call into a design that looks better than it is.

spacr.qt.widgets.power_design.simulator_kwargs(spec: DesignSpec) → Dict[str, Any][source]

The design as keyword arguments for the simulator, all held fixed.

The single translation point between the form and spacr.power_simulate.simulate_screen(). A sweep is produced by overwriting ONE of these keys with a list — see cells_grid() and wells_grid() — and handing the whole dict to spacr.power_model.scan_parameters(), which is what makes a run from this screen identical to the same run typed at a Python prompt.

Parameters:

spec – the design.

Returns:

keyword arguments for the simulator. Keys are the simulator’s own parameter names, spelling included.

spacr.qt.widgets.power_design.wells_grid(spec: DesignSpec) → List[float][source]

The wells sweep for spec.

Floored at 2 wells rather than 1: a one-well fit is refused by spacr.power_model.fit_model(), and a grid point that can only ever be an error is not a data point.

Parameters:

spec – the design.

Returns:

sorted well counts, including the design’s own total.