spacr.regression_diagnostics

Diagnostics that say whether a screen regression can be believed.

A volcano plot shows what the model concluded. It cannot show whether the model was entitled to conclude it. These are the checks that can, and they are the ones that would have caught the failure this module was written after: a design with 824 guides in 587 wells, fitted simultaneously, returning a confident coefficient and P value for every guide from a rank-deficient matrix.

The suite is deliberately split in two:

design_report / plot_design_diagnostics

Properties of the design, computable before any model is fitted, and the ones that decide whether the fit means anything: how many wells per parameter, the rank of the design matrix, its condition number, how many wells each guide appears in, and which guides are so collinear that no method can separate them.

residual_report / plot_residual_diagnostics

Properties of a fitted model: the residual-versus-fitted, scale-location, QQ and leverage panels, plus Cook’s distance.

plot_inference_diagnostics

Properties of the test: the P-value histogram, whose shape tells you whether the null is calibrated, and the observed-versus-expected quantile plot with its genomic inflation factor.

Every function takes plain arrays or frames and returns plain numbers, so they can be asserted on in tests rather than eyeballed.

Functions

collinear_guide_pairs(→ pandas.DataFrame)

Guide pairs whose well patterns are so alike they cannot be separated.

design_report(→ dict)

Measure whether a simultaneous fit on this design is identifiable.

plot_design_diagnostics(fractions, *[, block, ...])

Six panels describing the design, before any model is fitted.

plot_design_identifiability(fractions, *[, block, ...])

One compact panel reporting wells, predictors, rank and guide support.

plot_inference_diagnostics(p_values, *[, adjusted, ...])

Is the null calibrated, and how much of the family is non-null?

plot_residual_diagnostics(observed, fitted, *[, ...])

The four classical residual panels, plus Cook's distance.

residual_report(→ dict)

Summarise the residuals of a fitted model.

score_design(→ object)

Is a simultaneous fit on this design entitled to its coefficients?

score_inference(→ object)

Is the null calibrated, or is the whole family shifted?

score_residuals(→ object)

Are these residuals behaved enough for the inference drawn from them?

variance_inflation_factors(→ pandas.DataFrame)

Compute per-guide variance inflation factors for a full-rank design.

write_diagnostic_suite(→ Mapping[str, str])

Write every diagnostic the supplied inputs can support.

Module Contents

spacr.regression_diagnostics.collinear_guide_pairs(fractions: pandas.DataFrame, *, threshold: float = 0.95, limit: int = 500) → pandas.DataFrame[source]

Guide pairs whose well patterns are so alike they cannot be separated.

This is the mechanism behind the screen’s false positives: a guide that appears in nearly the same wells as a true hit inherits its signal, and no amount of correction distinguishes them, because the data contain no contrast between them. Reported as a table so the offenders can be named.

Parameters:
  • fractions – well-by-guide fraction matrix whose nonconstant guide columns are compared pairwise.

  • threshold – absolute Pearson correlation at or above which a pair is listed.

  • limit – stop after this many pairs; a wide screen has millions and the first few hundred are the ones worth reading.

spacr.regression_diagnostics.design_report(fractions: pandas.DataFrame, *, block: pandas.Series | None = None, presence_threshold: float = 0.0) → dict[source]

Measure whether a simultaneous fit on this design is identifiable.

Parameters:
  • fractions – well-by-guide matrix, one row per analysed well.

  • block – optional per-well block labels (normally the plate), which cost one parameter each beyond the first.

  • presence_threshold – a guide counts as present in a well when its value exceeds this.

Returns:

a dict of scalars. identifiable is the one that matters: it is False when the design has fewer wells than parameters, which is the state in which per-guide coefficients are not unique.

spacr.regression_diagnostics.plot_design_diagnostics(fractions: pandas.DataFrame, *, block: pandas.Series | None = None, save_path=None, save_format=None, presence_threshold: float = 0.0)[source]

Six panels describing the design, before any model is fitted.

Parameters:

fractions – well-by-guide fraction matrix, one row per analysed well, as taken by design_report().

spacr.regression_diagnostics.plot_design_identifiability(fractions: pandas.DataFrame, *, block: pandas.Series | None = None, save_path=None, save_format=None, presence_threshold: float = 0.0)[source]

One compact panel reporting wells, predictors, rank and guide support.

The six-panel plot_design_diagnostics() remains the detailed audit. This compact version is the publication-facing answer to a narrower question: can one coefficient per guide be identified from these wells?

Parameters:

fractions – well-by-guide fraction matrix, one row per analysed well, as taken by design_report().

Returns:

(written_path, report) in the same form as the other plotters.

spacr.regression_diagnostics.plot_inference_diagnostics(p_values, *, adjusted=None, alpha: float = 0.05, save_path=None, save_format=None, label: str = '')[source]

Is the null calibrated, and how much of the family is non-null?

A well-behaved screen gives a flat P-value histogram with a spike at zero. A histogram that slopes or humps in the middle means the test is mis-calibrated, and no correction repairs that.

Parameters:

p_values – raw P values of the tested family; non-finite values are dropped before plotting.

spacr.regression_diagnostics.plot_residual_diagnostics(observed, fitted, *, design: numpy.ndarray | None = None, save_path=None, save_format=None, label: str = '')[source]

The four classical residual panels, plus Cook’s distance.

Parameters:
  • observed – observed response per observation.

  • fitted – fitted value per observation, aligned with observed; the residuals are observed - fitted.

spacr.regression_diagnostics.residual_report(observed, fitted, *, design: numpy.ndarray | None = None) → dict[source]

Summarise the residuals of a fitted model.

Parameters:
  • observed – observed response values.

  • fitted – model predictions aligned one-to-one with observed.

  • design – the model matrix. When given, leverage and Cook’s distance are computed from its hat matrix; without it those keys are omitted rather than guessed.

spacr.regression_diagnostics.score_design(report: Mapping) → object[source]

Is a simultaneous fit on this design entitled to its coefficients?

Parameters:

report – the dict design_report() returns.

Returns:

a PanelVerdict.

Rank deficiency is a failure because a design with fewer independent directions than parameters has no unique coefficient vector. A pseudo-inverse can still return numbers, but individual effects are not identifiable.

The other two rules are the conventional ones and are cited in detail: a design needs more wells than parameters to have any residual degrees of freedom at all, and a guide seen in one well or none has no contrast to be estimated from.

spacr.regression_diagnostics.score_inference(report: Mapping) → object[source]

Is the null calibrated, or is the whole family shifted?

Read off the genomic inflation factor: the median observed chi-square over its null median, which is 1.0 when the null behaves. The bands are the ones the GWAS literature uses – 1.1 is where a report starts explaining itself and 1.2 is where the p-values are not believed.

A SPIKE OF SMALL p IS NOT A FAULT. It is what a screen with real hits looks like, and lambda is deliberately a MEDIAN so a handful of true hits barely move it. Scoring the spike itself would flag every successful screen.

Parameters:

report – inference summary with tests, genomic_inflation and optionally pi0 and estimated_non_null, as built by plot_inference_diagnostics(). No tests, or a non-finite inflation factor, scores unknown.

spacr.regression_diagnostics.score_residuals(report: Mapping) → object[source]

Are these residuals behaved enough for the inference drawn from them?

Parameters:

report – the dict residual_report() returns.

A DIAGNOSTIC TEST’S p IS BACKWARDS AND THAT IS THE COMMONEST WAY TO READ ONE WRONG: the null is “the assumption holds”, so a LARGE p is the good outcome. Every rule below is written in that direction and says so.

Thresholds are the conventional ones: p >= 0.05 for a diagnostic test, Cook’s distance of 0.5 and 1.0. A threshold invented here would be a number a reviewer cannot check.

spacr.regression_diagnostics.variance_inflation_factors(fractions: pandas.DataFrame, *, max_guides: int = 200) → pandas.DataFrame[source]

Compute per-guide variance inflation factors for a full-rank design.

All nonconstant guide columns participate in the calculation. The output is then limited to the guides with the widest well support, so max_guides controls reporting rather than the fitted design.

Parameters:
  • fractions (pandas.DataFrame) – Well-by-guide design matrix.

  • max_guides (int, default=200) – Maximum number of guides to return.

Returns:

pandas.DataFrame – Columns guide, vif, and wells_with_guide, ordered by decreasing VIF.

Raises:

ValueError – If the number of wells is not greater than the number of nonconstant guide columns. Use design_report() and collinear_guide_pairs() to inspect that rank-deficient design.

spacr.regression_diagnostics.write_diagnostic_suite(destination, *, fractions=None, block=None, observed=None, fitted=None, design=None, p_values=None, adjusted=None, alpha: float = 0.05, label: str = '', presence_threshold: float = 0.0, formats: Sequence[str] | None = None) → Mapping[str, str][source]

Write every diagnostic the supplied inputs can support.

Nothing is required: pass what a given analysis mode has. The permutation test has a design and P values but no fitted values; a simultaneous fit has all of them. Each block is skipped silently when its inputs are absent, and a block that raises is recorded as an error rather than aborting the run – a diagnostic that fails must never take the analysis down with it.

Every sheet also SCORES itself – see score_design(), score_residuals() and score_inference() – and wears its own badge. The verdicts are rows in diagnostic_summary.csv, per sheet and for the suite (its worst), rather than entries in the returned mapping: that mapping’s contract is “key -> a file that exists”. A reader is told whether the numbers are fine rather than expected to know.

Parameters:
  • destination – output directory created as needed for all diagnostic figures, CSV tables, and the optional suite summary.

  • formats –

    force particular formats, one file per format. None, the default, writes ONE file per panel in the format the user’s figure preference asks for.

    IT USED TO DEFAULT TO ("pdf", "png"), and that was wrong twice over. Each entry re-runs the whole panel – the same numbers computed and the same picture drawn a second time – and the pair ignored the figure-format preference completely, so a user who had chosen PNG got a PDF anyway. It also broke the rule the design sets, that a saved figure and a visible one are the same event: two files for one panel is two tiles in the gallery for one picture.

Nested helpers

write_diagnostic_suite._emit(name, function, **kwargs)

Write one diagnostic in every captured requested format.

Parameters:
  • name – stable panel name used for paths and manifest keys.

  • function – writer returning (written_path, report).

  • kwargs – diagnostic inputs forwarded to the writer.

Returns:

None. Successful artifacts are keyed by their actual output extension and their report is retained; advisory failures become per-format error entries without stopping the remaining writers.

spacr/regression_diagnostics.py:1012