spacr.figures.distributions¶
Well-level distribution panels for regression reports.
This module plots normalized guide representation and response distributions
from the table used to fit a model. These panels remain separate from
spacr.figures.panels, whose registry consumes coefficient tables.
Returned Panel records include the plotted data
and the statistics needed to interpret each distribution.
Functions¶
|
One distribution panel on its own figure. |
|
The guide-share column. |
|
The Gini coefficient of a non-negative sample. |
|
Is the library evenly represented within a well? |
|
The response once per well, when it is a per-well quantity. |
|
Each guide's share of its well, divided by an equal split of that well. |
|
Is the fitted family's distributional assumption plausible? |
|
The response column, and it is worth saying how it is decided. |
|
Write both distributions into a run's results folder. |
|
Skewness, excess kurtosis and the word that goes with them. |
|
The well identifier, or None when the frame does not carry one. |
Module Contents¶
- spacr.figures.distributions.build_panel(key: str, frame, *, target: str | None = None, figsize=(3.4, 2.6), **kwargs)[source]¶
One distribution panel on its own figure.
(figure, Panel).The same shape and the same margins as
spacr.figures.sheet.build_panel()so a saved distribution sits beside a saved volcano at the same size on the grid. The style is a CONTEXT MANAGER, as everywhere in this package: spaCR draws from a long-lived GUI and a global rcParams write would restyle every later figure in the session.- Parameters:
key – distribution-panel name from
REGISTRY.frame – data table consumed by the selected panel.
- spacr.figures.distributions.fraction_column(frame) str | None[source]¶
The guide-share column.
- Parameters:
frame – table whose columns are searched for a known guide-share name.
- spacr.figures.distributions.gini(values) float[source]¶
The Gini coefficient of a non-negative sample.
0 when every value is equal, approaching 1 when one value holds everything. The standard evenness statistic for a pooled library, and the same one
spacr.plot.plot_lorenz_curves()reports.THE SAME STATISTIC, NOT THE SAME NUMBER, and the difference has to be stated because both are labelled “Gini”. The Lorenz curves take raw gRNA counts pooled over a plate; this panel takes each guide’s share of its own well divided by that well’s equal split. On the tsg101 screen those are 0.20 and 0.32 – the panel’s is larger because normalising by the well removes the between-well spread that flattens the pooled curve. A reader who takes one for the other will read a change in the question as a change in the library.
NaN for an empty sample or one that sums to zero, rather than a ZeroDivisionError or a silent 0.0 – an evenness of zero would be read as “perfectly even”, which is the opposite of “there was nothing to measure”.
- Parameters:
values – guide-representation values; non-finite values are ignored.
- spacr.figures.distributions.guide_fraction(ax, frame, *, well: str | None = None, bins: int | None = None, relative: bool = True) spacr.figures.panels.Panel[source]¶
Is the library evenly represented within a well?
Each guide against its own well’s equal share, on a log2 axis because the quantity is a RATIO: half and twice equal representation are the same distance from 1, which they are not on a linear axis, and a library’s abundances are log-normal to begin with.
- Parameters:
ax – Matplotlib axes on which to draw the histogram.
frame – guide-level table containing a recognized share column and, for the relative view, a well identifier.
- spacr.figures.distributions.one_value_per_well(frame, column: str, well: str | None)[source]¶
The response once per well, when it is a per-well quantity.
(values, deduplicated).THE BUG THIS EXISTS FOR. The pipeline hands the response as one row per guide-in-well, and the response is a property of the WELL – on this screen
log_predhas exactly one distinct value in each of the 610 wells, repeated once per guide the well retained. The old histogram therefore counted a 15-guide well fifteen times and stated n = 1,945 for 610 independent observations, overstating the evidence three-fold and reshaping the distribution towards whatever the crowded wells did.Checked rather than assumed: a response that genuinely varies within a well is left alone, because collapsing it would then be the error.
- Parameters:
frame – observation table containing the response and optional well identifier.
column – response column to extract from
frame.well – well-identifier column, or
Noneto retain every response observation.
- spacr.figures.distributions.relative_representation(frame, fraction: str, well: str)[source]¶
Each guide’s share of its well, divided by an equal split of that well.
(values, dropped_wells, dropped_rows). 1.0 means the guide holds exactly its share; 2.0 means twice what an equal split would give it.THE POINT OF DIVIDING. The raw fraction confounds evenness with how many guides were retained in the well: a two-guide well and a fifteen-guide well produce shares an order of magnitude apart with no unevenness whatsoever. Dividing by the well’s own equal share removes exactly that and nothing else, and it is what makes a single reference line legitimate.
- Parameters:
frame – guide-level table containing the share and well columns.
fraction – name of the guide-share column in
frame.well – name of the well-identifier column in
frame.
- spacr.figures.distributions.response(ax, frame, *, column: str | None = None, well: str | None = None, bins: int | None = None, family: str = 'gaussian') spacr.figures.panels.Panel[source]¶
Is the fitted family’s distributional assumption plausible?
The response with a normal of the same mean and SD over it, and the two numbers that decide what a reader does next: skewness and excess kurtosis. A bare “distribution of the response” leaves them to judge symmetry by eye, which is exactly what nobody can do.
- Parameters:
ax – Matplotlib axes on which to draw the histogram.
frame – model-input table containing the response observations.
- spacr.figures.distributions.response_column(frame, column: str | None = None) str | None[source]¶
The response column, and it is worth saying how it is decided.
An explicit name always wins – the pipeline knows its own dependent variable and should pass it. Otherwise a frame with exactly one numeric column is unambiguous, which is the shape
dmatriceshands back. Only past that does this guess fromRESPONSE_COLUMNS, and a panel that got here by guessing names the column it drew on the axis, so the guess is never invisible.- Parameters:
frame – model-input table whose numeric and named response columns are inspected.
- spacr.figures.distributions.save_distributions(frame, dst, *, response_variable: str | None = None, target: str | None = None, order: Sequence[str] = ORDER) Dict[str, str][source]¶
Write both distributions into a run’s results folder.
Returns
{panel_name: path}for panels that were drawn; unavailable panels are omitted. The function never opens an interactive window.The default
target='print'produces page-readable ink for saved files. Pass another target explicitly when the output will be embedded on a GUI surface.- Parameters:
frame – well- or guide-level table consumed by the distribution panels.
dst – results directory in which to write the panel files.
- spacr.figures.distributions.shape_of(values) dict[source]¶
Skewness, excess kurtosis and the word that goes with them.
Returned rather than printed so a test can assert the number a reader is shown, and so the console summary can quote the same one the panel does instead of computing its own.
- Parameters:
values – response values whose finite observations define the shape.