spacr.profiler

Evaluate a fitted model by varying one input and plotting its predictions.

A coefficient table answers “which terms matter”. It does not answer the question anyone actually has in front of a fitted screen model, which is what would this model predict for a well like mine — and then, having seen that, what happens if this one gRNA’s fraction doubles and everything else stays where it is. That is a profile: one input swept across its range, every other input pinned at a value the user chose, and the prediction plotted against the input that moved.

Three things make this harder than it sounds, and this module exists to settle all three in one place:

Seventeen backends predict differently. spacr.ml.REGRESSION_TYPES spans statsmodels results (OLS, WLS, RLM, four GLM families, QuantReg, MixedLM, BetaModel), scikit-learn estimators (Lasso, Ridge, ElasticNet, LinearSVC) and the horseshoe fitter. predict() is the one call that works on all of them, and response_scale() names what came back — because a GLM-binomial returns a probability, a Poisson GLM returns a rate, and LinearSVC.predict returns a CLASS LABEL, which would draw a step function and look like a finding. The hinge backend is profiled through decision_function for exactly that reason, and the scale says so.

Nothing here re-fits. The model handed in is the model profiled. A profiler that quietly re-fits is showing the user a second model and labelling it the first, and on a penalised backend with alpha='auto' the second one is not even the same model. Where a caller has only the written-out coefficients and no live object, from_coefficients() builds a FittedLinear around them — that is reading the fit, not repeating it, and the class says which link it is applying.

A held value is a choice, not a default. reference_row() picks the median of each column and is explicit that it did; the at argument overrides any of them; and every Profile carries the full held vector, so a curve can always be traced back to the assumptions that produced it.

Public API:

from spacr.profiler import profile, reference_row, sensitivity

curve = profile(model, design, "fraction:grna[233460_1]", n=41)
curve.predictions          # what the model says as that input sweeps
curve.held                 # where everything else was pinned
sensitivity(model, design) # which input moves the prediction most

Classes

FittedLinear

A fitted linear predictor rebuilt from coefficients already written.

Profile

One input swept, everything else pinned, and what the model said.

Sensitivity

How much one input moves the prediction, with the others pinned.

Functions

coefficient_frame(→ pandas.DataFrame)

Read a coefficient table from a path, DataFrame or fitted object.

from_coefficients(→ FittedLinear)

Build a FittedLinear from a written-out coefficient table.

predict(→ numpy.ndarray)

Predict from any of the fitted objects spacr.ml produces.

profile(→ Profile)

Sweep one input; hold the rest; return what the model predicts.

profile_by(→ List[Profile])

One profile of variable per level of a second input.

reference_row(→ pandas.Series)

The row every profile holds the non-moving inputs at.

response_scale(→ str)

Name what predict() returns for this model, for the axis label.

sensitivity(, offset, limit)

Rank the inputs by how far each one moves the prediction.

Module Contents

class spacr.profiler.FittedLinear[source]

A fitted linear predictor rebuilt from coefficients already written.

Not a re-fit. A regression run writes results.csv, and those numbers ARE the fit; this wraps them so the profiler can be pointed at a run that finished last week without the original object being alive. It quacks like a statsmodels result — params and predict — which is exactly the surface predict() and profile() need.

Parameters:
  • params – coefficients indexed by design-matrix column name; an Intercept entry is applied to every row.

  • link – which inverse link to apply; a key of LINKS.

  • label – what to call this model in a plot.

__post_init__() → None[source]

Validate the link and normalize coefficients to floating point.

Returns:

None.

Raises:

ValueError – link does not name a supported inverse link.

predict(exog: Any) → numpy.ndarray[source]

Predict on the response scale, applying link.

Parameters:

exog – design row or rows aligned with the fitted coefficients.

property feature_names: Tuple[str, ...][source]

Every coefficient name except the intercept, in fitted order.

class spacr.profiler.Profile[source]

One input swept, everything else pinned, and what the model said.

Parameters:
  • variable – the input that moved.

  • values – the values it took.

  • predictions – the model’s output at each of them.

  • held – every other input and the value it was held at.

  • baseline – the prediction with EVERY input at its held value — the point the curve is a departure from.

  • scale – what the predictions mean; see response_scale().

  • model_label – what to call the model.

  • reference_method – how the held values were chosen.

__len__() → int[source]

How many points the curve has.

at(value: float) → float[source]

The prediction at the swept point nearest value.

Parameters:

value – swept-variable value whose nearest prediction is requested.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy of the profile.

to_frame() → pandas.DataFrame[source]

The curve as a two-column frame, for a plot or a CSV.

property slope: float[source]

Change in prediction per unit of the input, end to end.

End to end rather than fitted: for a linear model they are the same number, and for a GLM the end-to-end value is the honest summary of a curve whose local slope changes.

property span: float[source]

How far the prediction moved across the whole sweep.

class spacr.profiler.Sensitivity[source]

How much one input moves the prediction, with the others pinned.

Parameters:
  • variable – the input.

  • low – the value swept from.

  • high – the value swept to.

  • prediction_low – what the model said at low.

  • prediction_high – what it said at high.

  • span – prediction_high - prediction_low; the ranking key is its magnitude.

  • coefficient – the fitted coefficient, when the model exposes one.

to_dict() → Dict[str, Any][source]

A JSON-serializable copy.

spacr.profiler.coefficient_frame(source: Any) → pandas.DataFrame[source]

Read a coefficient table from a path, DataFrame or fitted object.

Parameters:

source – a CSV path, a DataFrame with feature and coefficient columns, or a fitted object carrying params.

Returns:

a frame with feature and coefficient.

Raises:

ValueError – when the columns are not there.

spacr.profiler.from_coefficients(source: Any, *, link: str = 'identity', label: str = '', drop_zero: bool = False) → FittedLinear[source]

Build a FittedLinear from a written-out coefficient table.

Parameters:
  • source – anything coefficient_frame() accepts.

  • link – the inverse link the original fit used; "identity" for the least-squares and robust backends, "logit" / "probit" for the binomial ones, "log" for Poisson and horseshoe.

  • label – what to call the model in a plot.

  • drop_zero – leave out coefficients that are exactly zero. A penalised fit sets most of them there, and a profiler with three thousand flat inputs is unusable — but this is OFF by default, because “this gRNA does nothing” is a real answer a user may want to see the flat line for.

Returns:

a fitted linear predictor over those coefficients.

spacr.profiler.predict(model: Any, exog: pandas.DataFrame, *, offset: Sequence[float] | None = None) → numpy.ndarray[source]

Predict from any of the fitted objects spacr.ml produces.

The order the branches are tried in is load-bearing:

  1. decision_function — LinearSVC (the hinge backend) has both that and predict, and its predict returns a 0/1 CLASS. A profile of a class label is a step function, which reads as a finding and is an artefact of asking the wrong method.

  2. a statsmodels results object — predict(exog) applies the inverse link, so a GLM comes back on the response scale. offset is passed when the object accepts one.

  3. anything else with predict — the scikit-learn regressors.

  4. the linear predictor from params / coef_, for a fitted object that carries coefficients and nothing else (the horseshoe fitter).

Parameters:
  • model – a fitted object.

  • exog – design rows to predict for; columns must match the fit’s.

  • offset – per-row offset for a model fitted with one (Poisson and horseshoe use log(cell count)). Omitted means zero offset, i.e. the prediction for a well of unit exposure.

Returns:

one float per row of exog.

Raises:

TypeError – when the object carries neither a predict method nor coefficients — there is nothing to profile, and guessing would draw a curve that means nothing.

spacr.profiler.profile(model: Any, design: pandas.DataFrame, variable: str, *, values: Sequence[float] | None = None, at: Mapping[str, float] | None = None, n: int = 25, method: str = 'median', offset: float | None = None, label: str = '') → Profile[source]

Sweep one input; hold the rest; return what the model predicts.

Parameters:
  • model – a fitted object — anything predict() handles.

  • design – the design matrix, used for the sweep range and for the held values. It is READ, never fitted on.

  • variable – the column to move.

  • values – sweep these exact values instead of the column’s range.

  • at – hold named inputs at these values instead of the reference.

  • n – how many points to sweep when values is not given.

  • method – how unspecified inputs are held; see reference_row().

  • offset – per-row offset for a model fitted with one.

  • label – what to call the model in a plot; defaults to its class.

Returns:

a Profile.

Raises:

KeyError – when variable is not a column of the design.

spacr.profiler.profile_by(model: Any, design: pandas.DataFrame, variable: str, *, by: str, levels: Sequence[float], **kwargs: Any) → List[Profile][source]

One profile of variable per level of a second input.

The “with the other inputs held at chosen values” question asked several times at once: does this gRNA’s effect look different in a well with a high control fraction? Each returned profile carries the level it was drawn at in Profile.held.

Parameters:
  • model – a fitted object.

  • design – the design matrix.

  • variable – the input that moves.

  • by – the input held at each of levels in turn.

  • levels – the values to hold by at.

  • kwargs – passed through to profile().

Returns:

one Profile per level, in the order given.

Raises:

ValueError – when levels is empty.

spacr.profiler.reference_row(design: pandas.DataFrame, *, method: str = 'median', at: Mapping[str, float] | None = None) → pandas.Series[source]

The row every profile holds the non-moving inputs at.

Parameters:
  • design – the design matrix (or any frame with the fit’s columns).

  • method – "median" (the default — robust to the long right tail a per-gRNA fraction column has), "mean", "zero" or "min".

  • at – explicit values that override the chosen method, column by column. This is the “held at chosen values” half of the profiler.

Returns:

one row, indexed by the design’s columns.

Raises:

ValueError – on an unknown method, or a column named in at that the design does not have — a typo there would silently hold nothing and produce a curve with no explanation.

spacr.profiler.response_scale(model: Any) → str[source]

Name what predict() returns for this model, for the axis label.

Parameters:

model – fitted model or result whose prediction scale is identified.

Not cosmetic. The same curve means “probability that a well is positive”, “positive objects per cell” or “distance from a decision boundary” depending on the backend, and a plot that does not say which invites the wrong reading of all three.

spacr.profiler.sensitivity(model: Any, design: pandas.DataFrame, *, variables: Iterable[str] | None = None, at: Mapping[str, float] | None = None, method: str = 'median', quantiles: Tuple[float, float] = (0.05, 0.95), offset: float | None = None, limit: int | None = None) → List[Sensitivity][source]

Rank the inputs by how far each one moves the prediction.

Each input is swept from its low quantile to its high quantile with everything else held at the reference, and the inputs are ranked by the magnitude of the resulting change. Quantiles rather than min/max because one outlier well would otherwise decide the ranking.

This is what turns a three-thousand-column design into something a user can open a profiler on: the list is the answer to “which input should I move first”.

Parameters:
  • model – a fitted object.

  • design – the design matrix.

  • variables – restrict to these columns; default is every numeric column that is not constant and not the intercept.

  • at – hold named inputs at these values.

  • method – how the rest are held; see reference_row().

  • quantiles – the low and high sweep points.

  • offset – per-row offset for a model fitted with one.

  • limit – keep only the top this many.

Returns:

Sensitivity records, largest absolute span first.