spacr.group_lasso¶
Guides grouped by gene, selected or dropped as a set.
This model estimates which guides and genes are associated with a response while respecting the grouping of guides within genes.
WHAT IT IS. Ordinary lasso penalises each guide on its own, so from a gene’s four correlated guides it keeps whichever one happens to fit best and drops the rest – which reads as “one guide works and three do not” when the truth is “the gene matters and the four guides are four measurements of it”. Group lasso penalises the gene’s guides as a BLOCK:
minimise ||y - Xb||^2 / (2n) + lambda * sum_g sqrt(p_g) * ||b_g||_2
The L2 norm inside the sum is not squared, so its subgradient at zero is a ball rather than a point, and the whole block goes to exactly zero or none of it does. sqrt(p_g) is the standard correction that stops a gene with six guides being penalised more than one with three purely for having more.
WHY IT IS RECOMMENDED FOR A CRISPR SCREEN. It is the penalised analogue of
what the mixed model says with random effects – “which GENES are involved,
where each gene is measured by several guides” – and it reaches the same
question with no random effects, no REML fit, and no gene_fraction column.
That last part matters: gene_fraction is the sum of a gene’s guide
fractions, so a design carrying both blocks is singular by construction.
Here the gene enters as a GROUPING of the guide columns,
never as a column of its own, so the design stays full rank.
It also fits where OLS is undefined. This screen is 610 wells and 823 guides; the penalty is what makes the answer well posed at p > n, and it is honest about that – a penalised fit is a statement about the assumption that most guides do nothing, not new information.
WHAT IT DOES NOT GIVE YOU IS A P VALUE. A penalised coefficient’s sampling
distribution is not the OLS one and a naive P value from it is wrong in a
direction that flatters the fit. stability_selection() is the answer:
re-fit on many subsamples and report the FRACTION of them in which a gene
survives, which is an error-controlled statement (Meinshausen & Buhlmann)
that needs no distributional assumption at all.
THE SOLVER is block coordinate descent with the exact group-soft-threshold proximal step, in numpy. No new dependency: scikit-learn has no group lasso, and the two packages that do are unmaintained.
Functions¶
|
Select a group-lasso penalty by held-out prediction error. |
|
The formula and what is modelled, for the model tab's text box. |
|
Fit the group lasso and return |
|
One number per gene: the L2 norm of its guide block. |
|
The smallest penalty that zeroes every group. |
|
Penalties from |
|
How often each gene survives a re-fit on half the wells. |
Module Contents¶
- spacr.group_lasso.choose_lambda(X, y, labels, *, folds: int = PATH_FOLDS, points: int = PATH_POINTS, depth: float = PATH_DEPTH, required=None, seed: int = 0, **kwargs) float[source]¶
Select a group-lasso penalty by held-out prediction error.
A candidate is eligible only when its fit on the complete dataset selects at least one column allowed by
requiredand its held-out error is finite. If a penalty path contains no eligible candidate, progressively smaller penalties are evaluated for up toPATH_EXTENSIONSadditional ranges. If no eligible candidate is found, the smallest penalty evaluated is returned.- Parameters:
X – two-dimensional design matrix with one row per well and one column per guide.
y – response values aligned to the rows of
X.labels – group label for every column of
X; guides sharing a label are selected or dropped as one block.folds – how many held-out splits. Wells, not guides.
required – optional boolean mask over COLUMNS. A penalty counts as selecting only when one of these columns is non-zero. Defaults to every column.
seed – fixes the split, so two runs of one screen agree.
- Returns:
the chosen penalty, as a float.
- spacr.group_lasso.describe(lam: float = 0.05) str[source]¶
The formula and what is modelled, for the model tab’s text box.
- spacr.group_lasso.fit(X, y, labels, *, lam: float = 0.05, max_iterations: int = MAX_ITERATIONS, tolerance: float = TOLERANCE, fit_intercept: bool = True)[source]¶
Fit the group lasso and return
(coefficients, intercept, converged).- Parameters:
X – the design, one column per guide.
y – the response, one entry per well.
labels – the gene each COLUMN belongs to, length
X.shape[1].lam – the penalty. Zero is ordinary least squares and is allowed – it is how a caller checks the solver against a known answer.
fit_intercept – centre both sides, so the intercept is never penalised. Penalising it would shrink the response’s mean toward zero, which is not a claim anybody wants to make.
- Returns:
coefficients(one per column),intercept, and whether the sweep converged insidemax_iterations.- Raises:
ValueError – the shapes disagree, or
lamis negative.
- spacr.group_lasso.gene_effects(X, y, labels, **kwargs)[source]¶
One number per gene: the L2 norm of its guide block.
- Parameters:
X – design matrix with one column per guide.
y – response vector with one value per design row.
labels – gene label for each design column.
The block is zero or it is not, so the norm is the natural per-gene effect size – and unlike a coefficient it is never negative, which is honest: group lasso says a gene’s guides move the response TOGETHER, not in which direction each one does.
- spacr.group_lasso.max_lambda(X, y, labels) float[source]¶
The smallest penalty that zeroes every group.
- Parameters:
X – design matrix with one column per guide.
y – response vector with one value per design row.
labels – gene label for each design column.
Useful as the top of a path: any larger value gives the same all-zero fit, and starting a path above it wastes every iteration spent there.
- spacr.group_lasso.penalty_path(X, y, labels, *, points: int = PATH_POINTS, depth: float = PATH_DEPTH)[source]¶
Penalties from
max_lambda()down, log-spaced.Starting at the ceiling rather than at an absolute number is what makes the path scale-free: a design of fractions and a design of counts have ceilings three orders of magnitude apart, and a fixed grid that suits one is entirely above or entirely below the other.
- spacr.group_lasso.stability_selection(X, y, labels, *, lam: float = 0.05, n_boot: int = 100, fraction: float = 0.5, seed: int = 0, **kwargs)[source]¶
How often each gene survives a re-fit on half the wells.
- Parameters:
X – two-dimensional design matrix with wells in rows and guides in columns.
y – response values aligned to the rows of
X.labels – gene or other group label for every guide column.
n_boot – subsamples.
fraction – the share of ROWS in each subsample. Half is Meinshausen & Buhlmann’s choice and the one their error bound is stated for.
- Returns:
a DataFrame with a
selection_frequencyper gene, sorted.
THIS IS THE HONEST ANSWER TO “WHAT IS THE P VALUE”. There is not one – the sampling distribution of a penalised coefficient is not the OLS one, and a P value computed as though it were is wrong in the direction that flatters the fit. A selection frequency is a statement about how reproducible the selection is, which is what a screen actually wants to know, and it needs no distributional assumption at all.
Nested helpers¶
- choose_lambda.score(candidates)¶
(errors, selects) for one path, both aligned to
candidates.THE ERROR IS HELD-OUT AND THE SELECTION IS NOT. Whether a penalty predicts well is a question about unseen wells; whether it selects anything is a question about THE FIT THE CALLER WILL GET, which is fitted on every well. Reading selection off the folds instead picked a penalty at which four fifths of the wells kept a gene and all of them together kept none – measured on the tsg101 screen, where it chose 0.000183 and the full fit was empty.
spacr/group_lasso.py:250