spacr.trial_metrics¶
Everything a sweep row needs to be worth reading.
A row that reports only a hit count is insufficient for choosing a sweep configuration: a larger count can coincide with losing a known positive control. Each row therefore records model quality, design diagnostics, and control recovery alongside its hit count.
So a row carries four kinds of thing, and they answer different questions:
fit quality is the model describing the data at all residuals are its standard errors, and therefore its p-values, real design is the thing even identifiable, and how much data reached it controls does it recover what is known to be true
The controls matter most. When the user has named a positive control, it is the yardstick a configuration should be read against – a setting that buries a gene known to be real has told you something no goodness-of-fit statistic will.
Every function here returns a flat dict of scalars, because a sweep row is a row: nothing nested, nothing that needs unpacking to sort on.
EVERYTHING HERE IS CHEAP, AND THAT IS THE DESIGN CONSTRAINT. A sweep runs
hundreds of trials, so a per-trial diagnostic that costs seconds costs hours
across the run. The full 23-panel suite in spacr.regression_qc was
measured at ~5.8 s per fit – ten minutes per hundred trials, for pictures
nobody opens while a sweep is running. These scalars were measured at 73-230 ms
on screen-shaped fits (606-2400 rows, 200-400 parameters), which is under half
a percent of the ~60 s a real trial takes. So a sweep gets the NUMBERS on every
row and the PICTURES only for the rows worth reopening.
Nothing here refits anything. Every statistic is read off the model the trial already fitted, which is what keeps it cheap.
Functions¶
|
Is the null flat? Inflation above 1 says the hits are partly artefact. |
|
Summarize where configured controls rank in a result table. |
|
Rank, identifiability and collinearity -- read off the existing fit. |
|
How much data reached the fit, and whether it was identifiable. |
|
R-squared and the information criteria, when the model reports them. |
|
How many of the hits rest on a single guide. |
|
How many things the trial called, raw and corrected. |
|
Summarize design and inference diagnostics for one sweep trial. |
|
Homoscedasticity, autocorrelation and normality of the residuals. |
|
Every metric, flat, for one trial's row. |
Module Contents¶
- spacr.trial_metrics.calibration(results: pandas.DataFrame) dict[source]¶
Is the null flat? Inflation above 1 says the hits are partly artefact.
- Parameters:
results – result table carrying raw
p_valuevalues.
- spacr.trial_metrics.control_recovery(results: pandas.DataFrame, settings: Mapping[str, Any]) dict[source]¶
Summarize where configured controls rank in a result table.
Both rank and percentile are returned because result tables can contain different numbers of coefficients. Missing controls are reported with a false
*_control_foundvalue and no invented rank.- Parameters:
results (pandas.DataFrame) – Coefficient table produced by a regression trial.
settings (mapping) – Trial settings containing optional positive and negative controls.
- Returns:
dict – Control presence, rank, percentile, effect, and significance fields.
- spacr.trial_metrics.design_diagnostics(model) dict[source]¶
Rank, identifiability and collinearity – read off the existing fit.
- Parameters:
model – fitted model result object, or
None.
NOTHING IS RECOMPUTED HERE THAT THE FIT ALREADY KNOWS.
spacr.regression_diagnostics.design_report()answers the same questions, but it takes the well-by-guide matrix and pays for a freshmatrix_rankand a full SVD. A sweep does not have that matrix – it has a fitted model – and statsmodels computed the rank at fit time and kept it onmodel.model.rank. The field names below deliberately matchdesign_report’s so the two read the same in a trial folder.identifiableis the one that decides whether the rest of the row means anything. A rank-deficient design does not have unique coefficients, so its per-guide effects are one of infinitely many solutions and its p-values rank an arbitrary choice among them. spaCR fits these designs happily – statsmodels falls back to a pseudo-inverse rather than refusing – so nothing else in the output says it happened.
- spacr.trial_metrics.design_summary(output: Mapping[str, Any]) dict[source]¶
How much data reached the fit, and whether it was identifiable.
- Parameters:
output – trial output mapping optionally carrying prepared
model_dataand measured per-levelfit_designs. Fitted counts are reported per level; an unqualified count is emitted only when every fit records the same value. Legacy outputs without fit records retain their table-based counts.
- spacr.trial_metrics.fit_quality(model) dict[source]¶
R-squared and the information criteria, when the model reports them.
- Parameters:
model – fitted model result object, or
None.
The supported model families do not agree on which statistics exist. An RLM fit has no R-squared, a penalised model’s value is not comparable to OLS’s, and a permutation test has no model object at all. Reporting NaN for an absent statistic is honest; inventing one is not.
- spacr.trial_metrics.guide_support_summary(results: pandas.DataFrame, alpha: float = 0.05) dict[source]¶
How many of the hits rest on a single guide.
- Parameters:
results – guide-level result table used to compute gene support.
A gene with one surviving guide has a gene-level p identical to that guide’s, so it is not independent evidence – and on this screen the top of the list is exactly that. A count of them belongs in the row.
- spacr.trial_metrics.hit_counts(output: Mapping[str, Any], alpha: float = 0.05) dict[source]¶
How many things the trial called, raw and corrected.
- Parameters:
output – trial output carrying result and significant-hit frames.
- spacr.trial_metrics.qc_verdicts(row: Mapping[str, Any]) dict[source]¶
Summarize design and inference diagnostics for one sweep trial.
- Parameters:
row (mapping) – Trial metrics produced by
summarise_trial(). Design scoring uses well, parameter, rank, and condition-number fields. Inference scoring uses the result count and genomic-inflation estimate.- Returns:
dict – Available
qc_designandqc_inferencelevels, plusqc_verdictcontaining the more severe available level. A diagnostic is omitted when its required metrics are unavailable.
Notes
The overall verdict is intentionally the worst available diagnostic rather than a pass rate. This prevents a severe identifiability or calibration problem from being hidden by otherwise acceptable checks.
- spacr.trial_metrics.residual_diagnostics(model) dict[source]¶
Homoscedasticity, autocorrelation and normality of the residuals.
- Parameters:
model – fitted model result object, or
None.
These are what decide whether the p-values mean anything. A funnel in the residuals inflates or deflates every standard error in the fit, so a heteroscedastic model can rank genes plausibly and still be wrong about all of them.
- spacr.trial_metrics.summarise_trial(output: Mapping[str, Any], settings: Mapping[str, Any]) dict[source]¶
Every metric, flat, for one trial’s row.
- Parameters:
output – completed trial output with results, model, and fitted data.
settings – trial settings used for thresholds and control recovery.
Each block is guarded on its own: a family with no R-squared must still contribute its control recovery, and a design too wide for White’s test must still contribute its residual trend. One missing statistic is a NaN in one column, never a lost row.