spacr.profiler¶
Evaluate a fitted model by varying one input and plotting its predictions.
A coefficient table answers “which terms matter”. It does not answer the question anyone actually has in front of a fitted screen model, which is what would this model predict for a well like mine — and then, having seen that, what happens if this one gRNA’s fraction doubles and everything else stays where it is. That is a profile: one input swept across its range, every other input pinned at a value the user chose, and the prediction plotted against the input that moved.
Three things make this harder than it sounds, and this module exists to settle all three in one place:
Seventeen backends predict differently. spacr.ml.REGRESSION_TYPES
spans statsmodels results (OLS, WLS, RLM, four GLM families,
QuantReg, MixedLM, BetaModel), scikit-learn estimators
(Lasso, Ridge, ElasticNet, LinearSVC) and the horseshoe
fitter. predict() is the one call that works on all of them, and
response_scale() names what came back — because a GLM-binomial returns
a probability, a Poisson GLM returns a rate, and LinearSVC.predict
returns a CLASS LABEL, which would draw a step function and look like a
finding. The hinge backend is profiled through decision_function for
exactly that reason, and the scale says so.
Nothing here re-fits. The model handed in is the model profiled. A
profiler that quietly re-fits is showing the user a second model and
labelling it the first, and on a penalised backend with alpha='auto'
the second one is not even the same model. Where a caller has only the
written-out coefficients and no live object, from_coefficients()
builds a FittedLinear around them — that is reading the fit, not
repeating it, and the class says which link it is applying.
A held value is a choice, not a default. reference_row() picks
the median of each column and is explicit that it did; the at argument
overrides any of them; and every Profile carries the full held
vector, so a curve can always be traced back to the assumptions that
produced it.
Public API:
from spacr.profiler import profile, reference_row, sensitivity
curve = profile(model, design, "fraction:grna[233460_1]", n=41)
curve.predictions # what the model says as that input sweeps
curve.held # where everything else was pinned
sensitivity(model, design) # which input moves the prediction most
Classes¶
A fitted linear predictor rebuilt from coefficients already written. |
|
One input swept, everything else pinned, and what the model said. |
|
How much one input moves the prediction, with the others pinned. |
Functions¶
|
Read a coefficient table from a path, DataFrame or fitted object. |
|
Build a |
|
Predict from any of the fitted objects |
|
Sweep one input; hold the rest; return what the model predicts. |
|
One profile of |
|
The row every profile holds the non-moving inputs at. |
|
Name what |
|
Rank the inputs by how far each one moves the prediction. |
Module Contents¶
- class spacr.profiler.FittedLinear[source]¶
A fitted linear predictor rebuilt from coefficients already written.
Not a re-fit. A regression run writes
results.csv, and those numbers ARE the fit; this wraps them so the profiler can be pointed at a run that finished last week without the original object being alive. It quacks like a statsmodels result —paramsandpredict— which is exactly the surfacepredict()andprofile()need.- Parameters:
params – coefficients indexed by design-matrix column name; an
Interceptentry is applied to every row.link – which inverse link to apply; a key of
LINKS.label – what to call this model in a plot.
- __post_init__() None[source]¶
Validate the link and normalize coefficients to floating point.
- Returns:
None.- Raises:
ValueError –
linkdoes not name a supported inverse link.
- predict(exog: Any) numpy.ndarray[source]¶
Predict on the response scale, applying
link.- Parameters:
exog – design row or rows aligned with the fitted coefficients.
- class spacr.profiler.Profile[source]¶
One input swept, everything else pinned, and what the model said.
- Parameters:
variable – the input that moved.
values – the values it took.
predictions – the model’s output at each of them.
held – every other input and the value it was held at.
baseline – the prediction with EVERY input at its held value — the point the curve is a departure from.
scale – what the predictions mean; see
response_scale().model_label – what to call the model.
reference_method – how the held values were chosen.
- at(value: float) float[source]¶
The prediction at the swept point nearest
value.- Parameters:
value – swept-variable value whose nearest prediction is requested.
- to_frame() pandas.DataFrame[source]¶
The curve as a two-column frame, for a plot or a CSV.
- class spacr.profiler.Sensitivity[source]¶
How much one input moves the prediction, with the others pinned.
- Parameters:
variable – the input.
low – the value swept from.
high – the value swept to.
prediction_low – what the model said at
low.prediction_high – what it said at
high.span –
prediction_high - prediction_low; the ranking key is its magnitude.coefficient – the fitted coefficient, when the model exposes one.
- spacr.profiler.coefficient_frame(source: Any) pandas.DataFrame[source]¶
Read a coefficient table from a path, DataFrame or fitted object.
- Parameters:
source – a CSV path, a DataFrame with
featureandcoefficientcolumns, or a fitted object carryingparams.- Returns:
a frame with
featureandcoefficient.- Raises:
ValueError – when the columns are not there.
- spacr.profiler.from_coefficients(source: Any, *, link: str = 'identity', label: str = '', drop_zero: bool = False) FittedLinear[source]¶
Build a
FittedLinearfrom a written-out coefficient table.- Parameters:
source – anything
coefficient_frame()accepts.link – the inverse link the original fit used;
"identity"for the least-squares and robust backends,"logit"/"probit"for the binomial ones,"log"for Poisson and horseshoe.label – what to call the model in a plot.
drop_zero – leave out coefficients that are exactly zero. A penalised fit sets most of them there, and a profiler with three thousand flat inputs is unusable — but this is OFF by default, because “this gRNA does nothing” is a real answer a user may want to see the flat line for.
- Returns:
a fitted linear predictor over those coefficients.
- spacr.profiler.predict(model: Any, exog: pandas.DataFrame, *, offset: Sequence[float] | None = None) numpy.ndarray[source]¶
Predict from any of the fitted objects
spacr.mlproduces.The order the branches are tried in is load-bearing:
decision_function—LinearSVC(thehingebackend) has both that andpredict, and itspredictreturns a 0/1 CLASS. A profile of a class label is a step function, which reads as a finding and is an artefact of asking the wrong method.a statsmodels results object —
predict(exog)applies the inverse link, so a GLM comes back on the response scale.offsetis passed when the object accepts one.anything else with
predict— the scikit-learn regressors.the linear predictor from
params/coef_, for a fitted object that carries coefficients and nothing else (the horseshoe fitter).
- Parameters:
model – a fitted object.
exog – design rows to predict for; columns must match the fit’s.
offset – per-row offset for a model fitted with one (Poisson and horseshoe use
log(cell count)). Omitted means zero offset, i.e. the prediction for a well of unit exposure.
- Returns:
one float per row of
exog.- Raises:
TypeError – when the object carries neither a predict method nor coefficients — there is nothing to profile, and guessing would draw a curve that means nothing.
- spacr.profiler.profile(model: Any, design: pandas.DataFrame, variable: str, *, values: Sequence[float] | None = None, at: Mapping[str, float] | None = None, n: int = 25, method: str = 'median', offset: float | None = None, label: str = '') Profile[source]¶
Sweep one input; hold the rest; return what the model predicts.
- Parameters:
model – a fitted object — anything
predict()handles.design – the design matrix, used for the sweep range and for the held values. It is READ, never fitted on.
variable – the column to move.
values – sweep these exact values instead of the column’s range.
at – hold named inputs at these values instead of the reference.
n – how many points to sweep when
valuesis not given.method – how unspecified inputs are held; see
reference_row().offset – per-row offset for a model fitted with one.
label – what to call the model in a plot; defaults to its class.
- Returns:
a
Profile.- Raises:
KeyError – when
variableis not a column of the design.
- spacr.profiler.profile_by(model: Any, design: pandas.DataFrame, variable: str, *, by: str, levels: Sequence[float], **kwargs: Any) List[Profile][source]¶
One profile of
variableper level of a second input.The “with the other inputs held at chosen values” question asked several times at once: does this gRNA’s effect look different in a well with a high control fraction? Each returned profile carries the level it was drawn at in
Profile.held.- Parameters:
model – a fitted object.
design – the design matrix.
variable – the input that moves.
by – the input held at each of
levelsin turn.levels – the values to hold
byat.kwargs – passed through to
profile().
- Returns:
one
Profileper level, in the order given.- Raises:
ValueError – when
levelsis empty.
- spacr.profiler.reference_row(design: pandas.DataFrame, *, method: str = 'median', at: Mapping[str, float] | None = None) pandas.Series[source]¶
The row every profile holds the non-moving inputs at.
- Parameters:
design – the design matrix (or any frame with the fit’s columns).
method –
"median"(the default — robust to the long right tail a per-gRNA fraction column has),"mean","zero"or"min".at – explicit values that override the chosen method, column by column. This is the “held at chosen values” half of the profiler.
- Returns:
one row, indexed by the design’s columns.
- Raises:
ValueError – on an unknown method, or a column named in
atthat the design does not have — a typo there would silently hold nothing and produce a curve with no explanation.
- spacr.profiler.response_scale(model: Any) str[source]¶
Name what
predict()returns for this model, for the axis label.- Parameters:
model – fitted model or result whose prediction scale is identified.
Not cosmetic. The same curve means “probability that a well is positive”, “positive objects per cell” or “distance from a decision boundary” depending on the backend, and a plot that does not say which invites the wrong reading of all three.
- spacr.profiler.sensitivity(model: Any, design: pandas.DataFrame, *, variables: Iterable[str] | None = None, at: Mapping[str, float] | None = None, method: str = 'median', quantiles: Tuple[float, float] = (0.05, 0.95), offset: float | None = None, limit: int | None = None) List[Sensitivity][source]¶
Rank the inputs by how far each one moves the prediction.
Each input is swept from its low quantile to its high quantile with everything else held at the reference, and the inputs are ranked by the magnitude of the resulting change. Quantiles rather than min/max because one outlier well would otherwise decide the ranking.
This is what turns a three-thousand-column design into something a user can open a profiler on: the list is the answer to “which input should I move first”.
- Parameters:
model – a fitted object.
design – the design matrix.
variables – restrict to these columns; default is every numeric column that is not constant and not the intercept.
at – hold named inputs at these values.
method – how the rest are held; see
reference_row().quantiles – the low and high sweep points.
offset – per-row offset for a model fitted with one.
limit – keep only the top this many.
- Returns:
Sensitivityrecords, largest absolute span first.