spacr.nonparametric_fits¶
The seven nonparametric methods, sorted by what they can honestly answer.
WHY THIS IS NOT SEVEN MORE ENTRIES IN THE regression_type MENU. The fits spaCR already offers all answer in one currency – a coefficient per guide, with a standard error and a P value – and the rest of the screen is built on it: the volcano plots effect against significance, the hit list ranks genes by effect with a q over the genes tested, the attribution and the model card consume the same table.
FOUR OF THESE SEVEN PRODUCE NO SUCH NUMBER. LOWESS is descriptive. Kernel regression and KNN give a fitted surface, not a slope. A forest gives importances, which are not coefficients and are not comparable across features on the same scale. Choosing one of them as “the regression” would hand the volcano nothing to draw, so they are offered as what they are:
A FIT THAT ANSWERS IN THE SAME CURRENCY – joins
regression_type.A DIAGNOSTIC LAID OVER THE DATA – belongs on a plot, never decides hits.
AN AGREEMENT CHECK – reports a comparison against the linear ranking, and names the guides the two disagree about.
AND WHAT EACH IS FITTING IS NOT ALWAYS THE GUIDE DESIGN. spaCR’s fit is guide -> phenotype at WELL level with one column per guide: high dimensional, sparse and categorical, which is the worst case for most of these. Against a CONTINUOUS covariate – guide abundance in the well, cell count, plate position – they are on home ground, and that is where the smoothers earn their place: showing whether the phenotype moves smoothly with a nuisance variable the linear model is assuming away.
GROUP BY WELL. Cells in one well are not independent, so any split these methods need goes through the well, never through the cell.
Classes¶
Functions¶
|
Rank guides a second way and compare it with the linear ranking. |
|
One sentence naming what a method is for and what it costs. |
|
A monotone fit of |
|
Every method belonging to |
|
Why |
|
Run the agreement check on a finished fit and return what to print. |
|
Fit one of the diagnostic smoothers to |
|
Replace each named covariate with its spline basis. Returns a frame. |
Module Contents¶
- class spacr.nonparametric_fits.Agreement[source]¶
Two rankings of the same guides, and where they disagree.
THE OUTPUT IS A COMPARISON, NOT A COEFFICIENT TABLE. It asks whether an effect is supported by the data or induced by the model: agreement strengthens the result, while disagreement is itself a finding to inspect.
- Parameters:
method – alternative ranking method compared with the linear fit.
guides – guide names shared by the design and linear-effect mapping, in design-column order.
linear_rank – one-based guide ranks ordered by descending absolute linear effect.
other_rank – one-based guide ranks ordered by descending permutation importance from
method.correlation – Spearman correlation between the two rankings, or
nanwhen fewer than three guides are shared.disagreements –
(guide, linear_rank, other_rank)tuples whose ranks differ by the reportable threshold, ordered by largest movement.note – caveat needed to interpret the alternative ranking.
- class spacr.nonparametric_fits.Curve[source]¶
A fitted curve to draw over a scatter, and what it is.
NEVER A HIT LIST.
p_valuesdoes not exist on this object on purpose: a diagnostic that could be mistaken for an inferential test would be more misleading than omitting the method.- Parameters:
method – registered diagnostic method that produced the curve.
x – ordered predictor coordinates at which the curve is evaluated.
y – fitted response values aligned one-to-one with
x.lower – optional lower uncertainty-band coordinates.
upper – optional upper uncertainty-band coordinates aligned with
lower.note – preprocessing or interpretation detail that belongs beside the curve.
- spacr.nonparametric_fits.agreement(design, response, linear_effect: Dict[str, float], *, method: str = 'random_forest', groups=None, moved_by: int = 10, seed: int = 0) Agreement[source]¶
Rank guides a second way and compare it with the linear ranking.
- Parameters:
design – wells x guides. One row per WELL, never per cell.
response – phenotype value for every row of
design.linear_effect – the fit’s own per-guide effect, ranked by magnitude to give the ranking this is compared against.
groups – the well each row belongs to. Passed to the splitter so one well’s rows never straddle a split – cells in one well share that well’s phenotype, and a split that crossed one would score a model on its own training data.
moved_by – how many places a guide must move to be worth naming.
- spacr.nonparametric_fits.describe(name: str) str[source]¶
One sentence naming what a method is for and what it costs.
- Parameters:
name – registered method name to describe.
Said WHERE IT IS CHOSEN. The whole point of the three-way split is that a reader knows before picking, not after running.
- spacr.nonparametric_fits.isotonic_fit(x, y, *, increasing: bool = True)[source]¶
A monotone fit of
yon one orderedx. Returns (grid, fitted).- Parameters:
x – ordered-predictor values, one per observation.
y – response values aligned one-to-one with
x.
ONE DIMENSION AND ONE DIRECTION, which is the whole of what isotonic regression claims.
refuse('isotonic', ordered=False)is what says so before a caller points it at the guide design.
- spacr.nonparametric_fits.methods_in(category: str) Tuple[str, ...][source]¶
Every method belonging to
category, in declaration order.- Parameters:
category – one of the
fit,diagnostic, oragreementmethod categories.
- spacr.nonparametric_fits.refuse(name: str, *, rows: int = 0, ordered: bool = True, predictors: int = 1) str | None[source]¶
Why
namecannot be run on data of this shape, or None.- Parameters:
name – registered method whose applicability is being checked.
CHOOSING A METHOD ON DATA IT CANNOT FIT REFUSES WITH THE REASON, rather than returning a fit nobody should read. That is the rule this function exists for.
- spacr.nonparametric_fits.report_agreement(coefficients, design, response, *, method: str = 'random_forest', seed: int = 0) str[source]¶
Run the agreement check on a finished fit and return what to print.
Takes the coefficient table a run already produced, so nothing is refitted and the comparison is against the ranking the run actually reported.
- Parameters:
coefficients – the run’s table, with
featureandcoefficient.design – the completed fit’s design matrix. Guide-design columns are matched back to coefficient feature names.
response – phenotype vector aligned to the rows of
design.
- Returns:
the sentence to print, or “” when there is nothing to compare – too few shared guides, or a table with no coefficients in it.
- spacr.nonparametric_fits.smooth(x, y, *, method: str = 'lowess', points: int = 200, scaled: bool = True) Curve[source]¶
Fit one of the diagnostic smoothers to
yagainstx.- Parameters:
x – predictor values, one per observation. They may arrive unsorted;
xandyare sorted together before fitting.y – response values aligned one-to-one with
x.scaled – standardise
xbefore fitting for the methods that need it, and say so in the note. KNN and the Gaussian process are distance-based, so an unscaled covariate silently makes one unit of it mean whatever its range happens to be.
- Raises:
ValueError – when the method cannot be run on this shape – with the reason, rather than a fit nobody should read.
- spacr.nonparametric_fits.spline_design(frame, covariates: Sequence[str], *, knots: int = SPLINE_KNOTS, degree: int = SPLINE_DEGREE)[source]¶
Replace each named covariate with its spline basis. Returns a frame.
- Parameters:
frame – design frame containing guide and nuisance columns.
covariates – nuisance columns to replace with spline bases when their values support the requested degree.
THE GUIDE COLUMNS ARE NOT TOUCHED, and that is what keeps this in category A. Each guide keeps exactly one column, so the fit still produces one coefficient and one P value per guide and the volcano and the hit list draw it with no special-casing. What becomes nonlinear is the NUISANCE – the covariate the straight line was assuming away.
A basis column is named
<covariate>_spline, which carries nogrnaand nogene, so every filter that already dropsrowID[T.r2]drops these too.