spacr.qt.widgets.control_chart¶
Control charts across a campaign — the statistics, with the decisions argued.
A screening campaign is dozens of plates run over weeks. The positive and negative controls are supposed to be the same thing every time: that is what makes hit calling, plate normalisation and Z’ mean anything at all. When the controls stop being the same — the cell line drifts, a reagent lot changes, the incubator door is left open on a Friday — everything computed downstream is already wrong, and nothing in the analysis says so, because the analysis normalises to the controls it is given.
This module tracks one control value per plate over run order and puts limits
round it, so drift is visible before it ruins a screen rather than after.
It is pure numpy/pandas — no Qt anywhere in it, the same rule
spacr.qt.widgets.graph_spec and spacr.qt.widgets.pca_model
follow — so it is usable from a notebook, testable without a display, and
free for a report or a QC gate to call.
The decision the whole module turns on¶
Sigma comes from the average moving range divided by d2 = 1.128
(D2_MOVING_RANGE), never from the standard deviation of the series.
This is not a stylistic preference and it is not a small effect. The standard deviation of the plate-to-plate series is inflated by exactly the drift the chart exists to detect. Take a control that slides linearly by half a unit per plate over thirty plates with no other noise at all:
the SD of those thirty values is 4.40, so
mean ± 3 SDis[4.0, 30.5]and every single point is comfortably inside. The chart says “in control” while the campaign slides from 10 to 24.5.the moving range only ever sees plate-to-plate change, which for a slow drift is 0.5 and stays 0.5.
MR-bar / 1.128 = 0.443, so the limits are[13.42, 16.08]around a baseline centre of 14.75 — and 24 of the 30 plates fall outside them.
That asymmetry is the reason short-term (within-subgroup) variation is the only
admissible sigma for a control chart, and it is pinned by a test. What the SD
route would have said is computed anyway and reported — see
sd_reference_limits() and ControlChartResult.sd_would_flag —
because the argument is more convincing as two numbers side by side than as a
paragraph.
Three estimators, and which one ran is always named¶
ESTIMATOR_MOVING_RANGEIndividuals / moving range (I-MR). One control value per plate is the normal case in a screen, so this is what
ESTIMATOR_AUTOpicks when every plate contributes a single control well. Sigma ismean(|x_i - x_{i-1}|) / 1.128; 1.128 is d2 for a subgroup of two, which is what a moving range of consecutive points is.ESTIMATOR_SUBGROUP_SX-bar / S, for a plate with several control wells (subgroup
n > 1). The plotted point is the plate’s mean and sigma isS-bar / c4(n), wherec4()is computed from the gamma function rather than looked up in a table. The table in the back of a textbook stops at n = 25, does not interpolate, and is the sort of thing that gets transcribed with a typo into a module nobody re-derives.c4(n) = sqrt(2/(n-1)) · Γ(n/2) / Γ((n-1)/2)is three lines, is exact for every n, and can be checked against the published constant —c4(5) = 0.9400— which is what the test does.Sigma of the plotted statistic is
sigma_within / sqrt(n), and that is what the zones and the rules are stated in. When subgroup sizes differ between plates the limits differ per plate; they are computed per point rather than averaged, and the result says so.ESTIMATOR_ROBUST/ESTIMATOR_MADExplicit, never implicit, and named in the result. spaCR measurements are skewed and one catastrophically bad plate inside the baseline drags the classical centre and MR-bar out far enough to swallow everything after it. The robust variant takes the median as the centre and either the median moving range (
ESTIMATOR_ROBUST) or the MAD (ESTIMATOR_MAD) as the spread.The two robust constants are not interchangeable and the module refuses to pretend they are. 1.4826 (
MAD_SCALE) is1/Φ⁻¹(0.75), the factor that makes the MAD of a normal sample an unbiased estimate of its SD. A moving range is|x_i - x_{i-1}|, the absolute value of a difference of two observations, so its scale issqrt(2)·sigmaand its median issqrt(2)·Φ⁻¹(0.75)·sigma = 0.9539·sigma(MEDIAN_MR_CONSTANT). Multiplying a median moving range by 1.4826 overestimates sigma by exactlysqrt(2)— 41% wider limits, which is the difference between catching a drift and not. Both constants are computed fromΦ⁻¹(0.75)here for the same reason c4 is computed from the gamma function.
Phase I and Phase II — a parameter, not an afterthought¶
Limits are estimated from a stated baseline set of plates and then applied forward. Computing them from all the data including the drift is the standard way a control chart is made useless, and it is easy to do by accident because “just chart everything” is the obvious call.
The baseline is the first DEFAULT_BASELINE plates in run order by
default (twenty, the textbook Phase I minimum for an individuals chart), or an
explicit list of plate ids (ControlChartSpec.baseline_plates), or
everything before a cut-off in the order column
(ControlChartSpec.baseline_before). Whichever it was, the plates that
formed it are in ControlChartResult.baseline_plates and named in
ControlChartResult.report().
MIN_BASELINE is eight, because MR-bar over seven moving ranges is
already an opinion about one or two plate-to-plate jumps and below that it is
an opinion about one. Fewer than that is refused rather than answered.
Limits estimated from an out-of-control baseline are not limits. When the
baseline itself contains a violation the result says so, loudly, in
ControlChartResult.caveats(). Re-estimating without the flagged plates is
offered (ControlChartSpec.reestimate) and is off by default, because
silently deleting the inconvenient plates is how a baseline gets talked into
agreeing with itself. It runs one pass, not iterations to convergence:
iterating is a ratchet that shrinks the limits onto the tightest subset of the
baseline and eventually flags everything.
The rules, and saying which one fired¶
The Nelson / Western Electric set, all eight, each a named constant with what
it physically detects and its approximate false-alarm rate:
RULE_BEYOND_3_SIGMA, RULE_NINE_ONE_SIDE,
RULE_SIX_TRENDING, RULE_FOURTEEN_ALTERNATING,
RULE_TWO_OF_THREE_BEYOND_2, RULE_FOUR_OF_FIVE_BEYOND_1,
RULE_FIFTEEN_WITHIN_1, RULE_EIGHT_BEYOND_1.
Which points a rule flags is a convention, and an unstated one makes two
implementations disagree about the same chart, so it is stated: every point
that is part of the evidence is flagged. For rule 1 that is the point itself.
For the run rules (2, 3, 4, 7, 8) it is every point of the run, reported as one
maximal span rather than as a sliding window’s worth of overlapping alarms —
which is what lets ControlChartResult.report() say “plates P22–P30 are
nine in a row above the centre line” instead of listing three violations. For
the “k of m” rules (5, 6) it is the qualifying points only: a window can
trigger on two points beyond 2 sigma with a third sitting on the centre line,
and flagging that third point would be an accusation against a plate that did
nothing.
The rule set is selectable, and it has to be, because running all eight is
not eight times the sensitivity. Each rule has its own in-control false-alarm
rate; RULE_ALARM_RATE carries them, ControlChartResult.caveats()
does the arithmetic for the set that actually ran, and over a 200-plate
campaign the answer is a number of expected false alarms rather than a
reassurance. RULES_DEFAULT is the Western Electric four (1, 2, 5, 6),
which is the set the classical ARL figure was worked out for; RULES_ALL
is all eight, and RULES_LIMITS_ONLY is rule 1 alone for a user who
wants limits and nothing else.
Ordering¶
“Plate by plate over time” means the x axis is run order, and rules 2, 3,
4, 6, 7 and 8 are statements about a sequence — get the order wrong and they
are statements about nothing. An explicit order or date column
(ControlChartSpec.order) is used when given. Without one the plate id
is sorted with a natural, digit-aware key so plate 2 comes before plate 10, and
ControlChartResult.order_inferred is set, the note is in
ControlChartResult.notes, and ControlChartResult.report() says it
in capitals. An inferred order is a guess about the experiment, not a detail.
The key is a deliberate extension of
spacr.qt.widgets.graph_spec._sort_key(), not a copy of it: that one tries
float(text) and falls back to a plain string compare, which puts P10
before P2 — fine for facet labels, wrong for a run order. Here the string
is split into digit and non-digit runs and the digit runs compare as integers.
Z-prime¶
zprime_frame() computes the per-plate Z-factor when both a positive and a
negative control are named, and zprime_chart() charts it with everything
above — same estimator, same rules, same baseline. It is the number a screener
actually watches, and a Z’ that slides from 0.7 to 0.3 over a campaign is the
same failure this module exists for, one level up.
Exceptions¶
A chart that cannot mean anything, with the way out in the message. |
Classes¶
A chart, its limits, what fired, and everything needed to read it. |
|
What to chart, in what order, from which baseline, under which rules. |
|
One rule firing, over one span of plates. |
Functions¶
|
The unbiasing constant for the SD of a subgroup of |
|
The columns worth offering as the plate id or the control label. |
|
The columns worth offering as the charted measurement, sorted. |
|
Chart |
|
|
|
|
|
Per-plate Z-factor, one row per plate, in run order. |
Module Contents¶
- exception spacr.qt.widgets.control_chart.ControlChartError[source]¶
Bases:
ValueErrorA chart that cannot mean anything, with the way out in the message.
Raised rather than returned as an empty result. Every one of these is a sentence a user can act on — “18 plates carry a control value and the baseline needs 20” — and a caller that swallowed it would draw an empty axis with no explanation. The screen catches it and shows the text.
Initialize self. See help(type(self)) for accurate signature.
- class spacr.qt.widgets.control_chart.ControlChartResult[source]¶
A chart, its limits, what fired, and everything needed to read it.
- Parameters:
plates – one entry per point, in run order. The x axis.
values – the plotted statistic per plate — the control value, or the plate’s mean control value when there is more than one control well.
order_labels – what the order column said for each plate, as text. Equal to
plateswhen the order was inferred.subgroup_sizes – control wells behind each point.
subgroup_sd – within-plate SD per point; NaN where
n < 2.centre – the centre line, estimated from Phase I only.
sigma – sigma of the plotted statistic, which is what the zones and every rule are stated in. For X-bar/S that is
sigma_within / sqrt(n); when subgroup sizes differ it is the median ofsigma_atand the per-point values are the truth.sigma_within – the within-plate sigma, before the
sqrt(n). Equal tosigmafor an individuals chart.sigma_at – sigma per point.
lower – per-point lower limit (centre - 3 sigma).
upper – per-point upper limit.
z –
(value - centre) / sigma_atper point — the currency every rule is written in, so that varying subgroup sizes cost the rules nothing.estimator – which estimator actually ran. Never
"auto".baseline – positional indices of the Phase I plates.
baseline_excluded – plates dropped from Phase I by a re-estimation.
violations – every rule firing, in run order then rule order.
rules – the rule set that was run.
value_column – the column the plotted value was read from.
plate_column – the column that identifies each plate.
order_inferred – the order was guessed from the plate id.
degenerate – sigma came out zero — see
caveats().sd_reference –
(centre, sigma, lower, upper)fromsd_reference_limits()over every point.sd_would_flag – how many points fall outside those SD limits.
- false_alarm_rate() float[source]¶
Approximate probability that some selected rule fires on an in-control point.
1 - Π(1 - p_i)over the rules that ran. The independence it assumes is not true and the error has a known direction — seeRULE_ALARM_RATE— so this overstates the alarm rate, which is the safe direction for a warning.
- points_frame() pandas.DataFrame[source]¶
One row per plate — what the chart draws, and its CSV.
Everything a reader would want to re-plot elsewhere: the order label, the value, the subgroup, the limits that applied to that point, the z, whether it was Phase I and which rules fired on it.
- rules_at(index: int) Tuple[int, ...][source]¶
Which rules fired on point
index, ascending.A point can trip several — a plate beyond 3 sigma is usually also the end of a two-of-three-beyond-2-sigma window — and reporting only the first would hide that the same plate is evidence for two different failures.
- Parameters:
index – the point’s position in run order; outside
0 .. len(self) - 1raisesControlChartError.
- violations_frame() pandas.DataFrame[source]¶
One row per violation — the table the screen shows.
- zone(index: int) int[source]¶
Which sigma band point
indexsits in: 0, 1, 2 or 3.0 is inside 1 sigma, 3 is beyond 3 sigma. Signed nowhere — the side is the sign of
z. What a renderer colours a marker by.- Parameters:
index – the point’s position in run order; outside
0 .. len(self) - 1raisesControlChartError.
- property baseline_plates: Tuple[str, ...][source]¶
The plate ids Phase I was estimated from, in run order.
- property baseline_violations: Tuple[Violation, ...][source]¶
The violations that fell inside Phase I.
- property flagged: numpy.ndarray[source]¶
did any rule fire on it.
- Type:
Boolean per point
- class spacr.qt.widgets.control_chart.ControlChartSpec[source]¶
What to chart, in what order, from which baseline, under which rules.
Frozen and JSON round-tripping, like
GraphSpecandPCASpec, so a QC configuration is something a settings file or a report can carry and re-run against the next campaign — which is the whole point of a chart whose limits were estimated once.- Parameters:
value – the measured column to chart. Required at chart time.
plate – the column identifying the plate. Required at chart time. One point per distinct value of it.
order – an explicit run-order or date column. Strongly preferred: without it the order is inferred from
plateand every run-based rule rests on that inference.control_column – the column saying what each row is (
"condition","gene","well_type").Nonemeans the table is already only the control.control_levels – which level(s) of
control_columnare the control being charted. Required whenevercontrol_columnis set — a control column with no level named is a filter with no predicate.positive_levels – the positive control, for
zprime_frame().negative_levels – the negative control, likewise.
estimator – one of
ESTIMATORS.rules – which of
RULES_ALLto run.baseline_n – Phase I length when neither
baseline_platesnorbaseline_beforeis given.baseline_plates – an explicit Phase I set, by plate id.
baseline_before – everything strictly before this value of the order column is Phase I. A date, for a campaign whose baseline is “the first week”.
reestimate – when the baseline itself trips a rule, drop the flagged plates and estimate once more. Off by default; see the module docstring for why silently deleting the inconvenient plates is not the default behaviour of anything here.
- Raises:
ControlChartError – on an unknown estimator or rule id, a baseline shorter than
MIN_BASELINE, or a control column with no level named — at the point the spec is built, not at render time.
- __post_init__() None[source]¶
Normalise the columns and levels, and validate the chart settings.
- Raises:
ControlChartError – if the estimator is unknown; if a rule is not a number, or not one of rules 1-8; if the baseline is shorter than the module’s minimum – MR-bar over fewer moving ranges is already an opinion about one or two plates rather than an estimate; or if
control_columnandcontrol_levelsare given without each other, which says where to look without saying what for, or the reverse.
- classmethod from_dict(payload: Mapping[str, Any]) ControlChartSpec[source]¶
Rebuild from
to_dict().Unknown keys ignored, missing keys defaulted, so a configuration written by another build still opens: a QC chart nobody can reload is a QC chart nobody re-runs.
- Parameters:
payload – a mapping as written by
to_dict(); unknown keys are ignored and missing ones take their defaults.
- classmethod from_json(text: str) ControlChartSpec[source]¶
Rebuild from
to_json().- Parameters:
text – a JSON string as written by
to_json().
- to_dict() Dict[str, Any][source]¶
A plain JSON-able dict. Every field, always — a stable schema beats a compact one for something a QC configuration is stored as.
- with_baseline(*, n: int | None = None, plates: Sequence[str] | None = None, before: str | None = None) ControlChartSpec[source]¶
A copy with a different Phase I.
- with_columns(*, value: str | None = None, plate: str | None = None, order: str | None = None) ControlChartSpec[source]¶
A copy pointed at different columns.
- with_control(column: str | None, levels: Sequence[str]) ControlChartSpec[source]¶
A copy charting a different control.
- Parameters:
column – the column holding the control labels, or None.
levels – the labels in that column that mark control wells.
- with_estimator(estimator: str) ControlChartSpec[source]¶
A copy using a different sigma estimator.
- Parameters:
estimator – the sigma estimator name, e.g.
'auto','moving_range','subgroup_s','robust'or'mad'; stored without checking.
- with_rules(rules: Sequence[int]) ControlChartSpec[source]¶
A copy running a different rule set.
- Parameters:
rules – the rule numbers (1 to 8) to run.
- class spacr.qt.widgets.control_chart.Violation[source]¶
One rule firing, over one span of plates.
A maximal span, not a sliding window’s worth of overlapping alarms: twelve points in a row above the centre line is one violation of rule 2 covering twelve plates, which is what a reader needs, rather than four violations covering nine plates each, which is an artefact of the implementation.
- Parameters:
rule – the rule number, 1-8.
start – first index of the span, into
ControlChartResult.plates.end – last index of the span, inclusive.
points – the indices actually flagged. Equal to the whole span for the run rules; only the qualifying points for rules 5 and 6, where a window can trigger with an innocent point inside it.
plates – the plate ids of
points.side –
ABOVE,BELOWorEITHER. For rule 3 it is the direction of travel (ABOVEfor a rising trend) rather than a side of the centre line, because a trend can run entirely on one side of the centre or cross it and is a violation either way.in_baseline – whether any flagged point is inside Phase I. Limits estimated from an out-of-control baseline are not limits, so this is carried per violation rather than recomputed by every reader.
- describe() str[source]¶
The sentence
ControlChartResult.report()prints.Names the rule by number and in words, because a report that says “rule 6” to a reader who has not memorised Nelson’s numbering has said nothing, and one that says only “four of five beyond 1 sigma” cannot be cross-referenced against anyone else’s chart.
- spacr.qt.widgets.control_chart.c4(n: int) float[source]¶
The unbiasing constant for the SD of a subgroup of
n.c4(n) = sqrt(2/(n-1)) · Γ(n/2) / Γ((n-1)/2), which isE[s] / sigmafor a normal sample of sizen: the sample SD is a biased estimate of sigma, badly so for the small subgroups a plate’s control wells make, andS-bar / c4(n)is the correction.Computed from the gamma function rather than looked up, because the published table stops at 25, does not interpolate, and is a transcription error waiting to happen in a module nobody re-derives. The values agree with the published ones to every digit they print:
c4(2) = 0.7979,c4(5) = 0.9400,c4(10) = 0.9727.- Parameters:
n – the subgroup size, converted to int; must be at least 2.
- Raises:
ControlChartError – for
n < 2. A subgroup of one has no spread to unbias, which is the whole reason the I-MR chart exists.
- spacr.qt.widgets.control_chart.candidate_key_columns(frame: pandas.DataFrame) Tuple[str, ...][source]¶
The columns worth offering as the plate id or the control label.
Categorical by the same classifier, plus anything named like a plate or a date, because
plateIDis high-cardinality enough in a big campaign that the classifier calls it a key and skips it — and a key is exactly what is wanted here. This is the one place in spaCR where “identifies rather than describes” is a recommendation rather than a disqualification.- Parameters:
frame – the per-well table whose columns are classified by
spacr.qt.widgets.graph_spec.column_kinds(). Columns named like a plate, date, time, batch, run or order are added.
- spacr.qt.widgets.control_chart.candidate_value_columns(frame: pandas.DataFrame) Tuple[str, ...][source]¶
The columns worth offering as the charted measurement, sorted.
Continuous by
spacr.qt.widgets.graph_spec.column_kinds(), which is the one column classifier in this codebase — the Local Data Filter’s rule, re-read. Reused rather than re-derived: the Graph Builder, the PCA screen and this one must agree about whatcell_countis, or the same table presents three different mental models depending on which screen is open.- Parameters:
frame – the per-well table whose columns are classified by
spacr.qt.widgets.graph_spec.column_kinds().
- spacr.qt.widgets.control_chart.control_chart(frame: pandas.DataFrame, spec: ControlChartSpec | None = None) ControlChartResult[source]¶
Chart
spec’s control value plate by plate over run order.The whole policy is in the module docstring; the short version is that one point is drawn per plate in run order, sigma comes from short-term variation and never from the SD of the series, the limits are estimated from a stated baseline and applied forward, and every rule that fires is reported by number, in words, and with the plates it fired on.
- Parameters:
frame – the per-well table holding the spec’s value, plate, order and control columns.
- Raises:
ControlChartError – whenever the chart would be meaningless — with the reason and the way out in the message.
- spacr.qt.widgets.control_chart.sd_reference_limits(values: Sequence[float]) Tuple[float, float, float, float][source]¶
(centre, sigma, lower, upper)the wrong way — mean ± 3 SD.This is not an estimator option and never will be. It exists so a result can print what the SD route would have said next to what the moving-range route did say, because the central argument of this module is far more convincing as two intervals side by side than as a paragraph: on a drifting series the SD limits are wide enough to contain the drift, and a chart drawn with them reports “in control” all the way down.
Uses the sample SD (
ddof=1) over every value given, which is exactly the mistake being illustrated — limits computed from all the data, including the excursion they are supposed to detect.- Parameters:
values – the plotted values; non-finite ones are ignored.
- spacr.qt.widgets.control_chart.zprime_chart(frame: pandas.DataFrame, spec: ControlChartSpec) ControlChartResult[source]¶
zprime_frame(), charted with the same estimator, rules and baseline.A thin composition rather than a second engine: the Z’ series is one value per plate, which is an individuals chart, which is what
control_chart()already is. The order is carried as an explicit integer column, so the Z’ chart never reports an inferred order — the ordering decision was taken once, upstream.- Parameters:
frame – the per-well table holding the spec’s value, plate, order and control columns.
spec – the chart settings; must name
control_column,positive_levelsandnegative_levels. Its estimator, rules and baseline are reused for the Z’ chart.
- spacr.qt.widgets.control_chart.zprime_frame(frame: pandas.DataFrame, spec: ControlChartSpec) pandas.DataFrame[source]¶
Per-plate Z-factor, one row per plate, in run order.
Z' = 1 - 3(sd_pos + sd_neg) / |mean_pos - mean_neg|— the number a screener actually watches, and the one that says whether a plate could have detected anything at all. Charting it with the same limits and the same rules is the same question one level up: a Z’ that slides from 0.7 to 0.3 over a campaign is a screen that stopped working, plate by plate, with no single plate looking wrong.Both controls need at least two wells on a plate for the SD to exist, so a plate with a single positive well produces no Z’ and is left out rather than given a zero.
- Parameters:
frame – the per-well table holding the spec’s value, plate, order and control columns.
spec – the chart settings; must name
control_column,positive_levelsandnegative_levels.
- Returns:
columns
plate,order,order_index,zprime,separation, and the per-control means, SDs and n.- Raises:
ControlChartError – when the spec does not name both controls.
Nested helpers¶
- control_chart._finish(centre: float, sigma_within: float, baseline: np.ndarray, excluded: Tuple[str, ...]) ControlChartResult¶
Assemble the result, flagging a degenerate spread.
A sigma at or below the tolerance means every point is a violation, so it is reported as DEGENERATE rather than charted – limits drawn from no spread say nothing about the process.
spacr/qt/widgets/control_chart.py:1719