spacr.qt.widgets.control_chart

Control charts across a campaign — the statistics, with the decisions argued.

A screening campaign is dozens of plates run over weeks. The positive and negative controls are supposed to be the same thing every time: that is what makes hit calling, plate normalisation and Z’ mean anything at all. When the controls stop being the same — the cell line drifts, a reagent lot changes, the incubator door is left open on a Friday — everything computed downstream is already wrong, and nothing in the analysis says so, because the analysis normalises to the controls it is given.

This module tracks one control value per plate over run order and puts limits round it, so drift is visible before it ruins a screen rather than after. It is pure numpy/pandas — no Qt anywhere in it, the same rule spacr.qt.widgets.graph_spec and spacr.qt.widgets.pca_model follow — so it is usable from a notebook, testable without a display, and free for a report or a QC gate to call.

The decision the whole module turns on

Sigma comes from the average moving range divided by d2 = 1.128 (D2_MOVING_RANGE), never from the standard deviation of the series.

This is not a stylistic preference and it is not a small effect. The standard deviation of the plate-to-plate series is inflated by exactly the drift the chart exists to detect. Take a control that slides linearly by half a unit per plate over thirty plates with no other noise at all:

  • the SD of those thirty values is 4.40, so mean ± 3 SD is [4.0, 30.5] and every single point is comfortably inside. The chart says “in control” while the campaign slides from 10 to 24.5.

  • the moving range only ever sees plate-to-plate change, which for a slow drift is 0.5 and stays 0.5. MR-bar / 1.128 = 0.443, so the limits are [13.42, 16.08] around a baseline centre of 14.75 — and 24 of the 30 plates fall outside them.

That asymmetry is the reason short-term (within-subgroup) variation is the only admissible sigma for a control chart, and it is pinned by a test. What the SD route would have said is computed anyway and reported — see sd_reference_limits() and ControlChartResult.sd_would_flag — because the argument is more convincing as two numbers side by side than as a paragraph.

Three estimators, and which one ran is always named

ESTIMATOR_MOVING_RANGE

Individuals / moving range (I-MR). One control value per plate is the normal case in a screen, so this is what ESTIMATOR_AUTO picks when every plate contributes a single control well. Sigma is mean(|x_i - x_{i-1}|) / 1.128; 1.128 is d2 for a subgroup of two, which is what a moving range of consecutive points is.

ESTIMATOR_SUBGROUP_S

X-bar / S, for a plate with several control wells (subgroup n > 1). The plotted point is the plate’s mean and sigma is S-bar / c4(n), where c4() is computed from the gamma function rather than looked up in a table. The table in the back of a textbook stops at n = 25, does not interpolate, and is the sort of thing that gets transcribed with a typo into a module nobody re-derives. c4(n) = sqrt(2/(n-1)) · Γ(n/2) / Γ((n-1)/2) is three lines, is exact for every n, and can be checked against the published constant — c4(5) = 0.9400 — which is what the test does.

Sigma of the plotted statistic is sigma_within / sqrt(n), and that is what the zones and the rules are stated in. When subgroup sizes differ between plates the limits differ per plate; they are computed per point rather than averaged, and the result says so.

ESTIMATOR_ROBUST / ESTIMATOR_MAD

Explicit, never implicit, and named in the result. spaCR measurements are skewed and one catastrophically bad plate inside the baseline drags the classical centre and MR-bar out far enough to swallow everything after it. The robust variant takes the median as the centre and either the median moving range (ESTIMATOR_ROBUST) or the MAD (ESTIMATOR_MAD) as the spread.

The two robust constants are not interchangeable and the module refuses to pretend they are. 1.4826 (MAD_SCALE) is 1/Φ⁻¹(0.75), the factor that makes the MAD of a normal sample an unbiased estimate of its SD. A moving range is |x_i - x_{i-1}|, the absolute value of a difference of two observations, so its scale is sqrt(2)·sigma and its median is sqrt(2)·Φ⁻¹(0.75)·sigma = 0.9539·sigma (MEDIAN_MR_CONSTANT). Multiplying a median moving range by 1.4826 overestimates sigma by exactly sqrt(2) — 41% wider limits, which is the difference between catching a drift and not. Both constants are computed from Φ⁻¹(0.75) here for the same reason c4 is computed from the gamma function.

Phase I and Phase II — a parameter, not an afterthought

Limits are estimated from a stated baseline set of plates and then applied forward. Computing them from all the data including the drift is the standard way a control chart is made useless, and it is easy to do by accident because “just chart everything” is the obvious call.

The baseline is the first DEFAULT_BASELINE plates in run order by default (twenty, the textbook Phase I minimum for an individuals chart), or an explicit list of plate ids (ControlChartSpec.baseline_plates), or everything before a cut-off in the order column (ControlChartSpec.baseline_before). Whichever it was, the plates that formed it are in ControlChartResult.baseline_plates and named in ControlChartResult.report().

MIN_BASELINE is eight, because MR-bar over seven moving ranges is already an opinion about one or two plate-to-plate jumps and below that it is an opinion about one. Fewer than that is refused rather than answered.

Limits estimated from an out-of-control baseline are not limits. When the baseline itself contains a violation the result says so, loudly, in ControlChartResult.caveats(). Re-estimating without the flagged plates is offered (ControlChartSpec.reestimate) and is off by default, because silently deleting the inconvenient plates is how a baseline gets talked into agreeing with itself. It runs one pass, not iterations to convergence: iterating is a ratchet that shrinks the limits onto the tightest subset of the baseline and eventually flags everything.

The rules, and saying which one fired

The Nelson / Western Electric set, all eight, each a named constant with what it physically detects and its approximate false-alarm rate: RULE_BEYOND_3_SIGMA, RULE_NINE_ONE_SIDE, RULE_SIX_TRENDING, RULE_FOURTEEN_ALTERNATING, RULE_TWO_OF_THREE_BEYOND_2, RULE_FOUR_OF_FIVE_BEYOND_1, RULE_FIFTEEN_WITHIN_1, RULE_EIGHT_BEYOND_1.

Which points a rule flags is a convention, and an unstated one makes two implementations disagree about the same chart, so it is stated: every point that is part of the evidence is flagged. For rule 1 that is the point itself. For the run rules (2, 3, 4, 7, 8) it is every point of the run, reported as one maximal span rather than as a sliding window’s worth of overlapping alarms — which is what lets ControlChartResult.report() say “plates P22–P30 are nine in a row above the centre line” instead of listing three violations. For the “k of m” rules (5, 6) it is the qualifying points only: a window can trigger on two points beyond 2 sigma with a third sitting on the centre line, and flagging that third point would be an accusation against a plate that did nothing.

The rule set is selectable, and it has to be, because running all eight is not eight times the sensitivity. Each rule has its own in-control false-alarm rate; RULE_ALARM_RATE carries them, ControlChartResult.caveats() does the arithmetic for the set that actually ran, and over a 200-plate campaign the answer is a number of expected false alarms rather than a reassurance. RULES_DEFAULT is the Western Electric four (1, 2, 5, 6), which is the set the classical ARL figure was worked out for; RULES_ALL is all eight, and RULES_LIMITS_ONLY is rule 1 alone for a user who wants limits and nothing else.

Ordering

“Plate by plate over time” means the x axis is run order, and rules 2, 3, 4, 6, 7 and 8 are statements about a sequence — get the order wrong and they are statements about nothing. An explicit order or date column (ControlChartSpec.order) is used when given. Without one the plate id is sorted with a natural, digit-aware key so plate 2 comes before plate 10, and ControlChartResult.order_inferred is set, the note is in ControlChartResult.notes, and ControlChartResult.report() says it in capitals. An inferred order is a guess about the experiment, not a detail.

The key is a deliberate extension of spacr.qt.widgets.graph_spec._sort_key(), not a copy of it: that one tries float(text) and falls back to a plain string compare, which puts P10 before P2 — fine for facet labels, wrong for a run order. Here the string is split into digit and non-digit runs and the digit runs compare as integers.

Z-prime

zprime_frame() computes the per-plate Z-factor when both a positive and a negative control are named, and zprime_chart() charts it with everything above — same estimator, same rules, same baseline. It is the number a screener actually watches, and a Z’ that slides from 0.7 to 0.3 over a campaign is the same failure this module exists for, one level up.

Exceptions

ControlChartError

A chart that cannot mean anything, with the way out in the message.

Classes

ControlChartResult

A chart, its limits, what fired, and everything needed to read it.

ControlChartSpec

What to chart, in what order, from which baseline, under which rules.

Violation

One rule firing, over one span of plates.

Functions

c4(→ float)

The unbiasing constant for the SD of a subgroup of n.

candidate_key_columns(→ Tuple[str, ...])

The columns worth offering as the plate id or the control label.

candidate_value_columns(→ Tuple[str, ...])

The columns worth offering as the charted measurement, sorted.

control_chart(→ ControlChartResult)

Chart spec's control value plate by plate over run order.

sd_reference_limits(→ Tuple[float, float, float, float])

(centre, sigma, lower, upper) the wrong way — mean ± 3 SD.

zprime_chart(→ ControlChartResult)

zprime_frame(), charted with the same estimator, rules and baseline.

zprime_frame(→ pandas.DataFrame)

Per-plate Z-factor, one row per plate, in run order.

Module Contents

exception spacr.qt.widgets.control_chart.ControlChartError[source]

Bases: ValueError

A chart that cannot mean anything, with the way out in the message.

Raised rather than returned as an empty result. Every one of these is a sentence a user can act on — “18 plates carry a control value and the baseline needs 20” — and a caller that swallowed it would draw an empty axis with no explanation. The screen catches it and shows the text.

Initialize self. See help(type(self)) for accurate signature.

class spacr.qt.widgets.control_chart.ControlChartResult[source]

A chart, its limits, what fired, and everything needed to read it.

Parameters:
  • plates – one entry per point, in run order. The x axis.

  • values – the plotted statistic per plate — the control value, or the plate’s mean control value when there is more than one control well.

  • order_labels – what the order column said for each plate, as text. Equal to plates when the order was inferred.

  • subgroup_sizes – control wells behind each point.

  • subgroup_sd – within-plate SD per point; NaN where n < 2.

  • centre – the centre line, estimated from Phase I only.

  • sigma – sigma of the plotted statistic, which is what the zones and every rule are stated in. For X-bar/S that is sigma_within / sqrt(n); when subgroup sizes differ it is the median of sigma_at and the per-point values are the truth.

  • sigma_within – the within-plate sigma, before the sqrt(n). Equal to sigma for an individuals chart.

  • sigma_at – sigma per point.

  • lower – per-point lower limit (centre - 3 sigma).

  • upper – per-point upper limit.

  • z – (value - centre) / sigma_at per point — the currency every rule is written in, so that varying subgroup sizes cost the rules nothing.

  • estimator – which estimator actually ran. Never "auto".

  • baseline – positional indices of the Phase I plates.

  • baseline_excluded – plates dropped from Phase I by a re-estimation.

  • violations – every rule firing, in run order then rule order.

  • rules – the rule set that was run.

  • value_column – the column the plotted value was read from.

  • plate_column – the column that identifies each plate.

  • order_inferred – the order was guessed from the plate id.

  • degenerate – sigma came out zero — see caveats().

  • sd_reference – (centre, sigma, lower, upper) from sd_reference_limits() over every point.

  • sd_would_flag – how many points fall outside those SD limits.

__len__() → int[source]

Plates on the chart.

caveats() → Tuple[str, ...][source]

Everything a reader needs before believing the chart.

false_alarm_rate() → float[source]

Approximate probability that some selected rule fires on an in-control point.

1 - Π(1 - p_i) over the rules that ran. The independence it assumes is not true and the error has a known direction — see RULE_ALARM_RATE — so this overstates the alarm rate, which is the safe direction for a warning.

headline() → str[source]

One sentence about the campaign, including the bad news.

points_frame() → pandas.DataFrame[source]

One row per plate — what the chart draws, and its CSV.

Everything a reader would want to re-plot elsewhere: the order label, the value, the subgroup, the limits that applied to that point, the z, whether it was Phase I and which rules fired on it.

report() → str[source]

The whole story, as the screen shows it and a report file writes it.

rules_at(index: int) → Tuple[int, ...][source]

Which rules fired on point index, ascending.

A point can trip several — a plate beyond 3 sigma is usually also the end of a two-of-three-beyond-2-sigma window — and reporting only the first would hide that the same plate is evidence for two different failures.

Parameters:

index – the point’s position in run order; outside 0 .. len(self) - 1 raises ControlChartError.

violations_frame() → pandas.DataFrame[source]

One row per violation — the table the screen shows.

zone(index: int) → int[source]

Which sigma band point index sits in: 0, 1, 2 or 3.

0 is inside 1 sigma, 3 is beyond 3 sigma. Signed nowhere — the side is the sign of z. What a renderer colours a marker by.

Parameters:

index – the point’s position in run order; outside 0 .. len(self) - 1 raises ControlChartError.

property baseline_plates: Tuple[str, ...][source]

The plate ids Phase I was estimated from, in run order.

property baseline_violations: Tuple[Violation, ...][source]

The violations that fell inside Phase I.

property flagged: numpy.ndarray[source]

did any rule fire on it.

Type:

Boolean per point

class spacr.qt.widgets.control_chart.ControlChartSpec[source]

What to chart, in what order, from which baseline, under which rules.

Frozen and JSON round-tripping, like GraphSpec and PCASpec, so a QC configuration is something a settings file or a report can carry and re-run against the next campaign — which is the whole point of a chart whose limits were estimated once.

Parameters:
  • value – the measured column to chart. Required at chart time.

  • plate – the column identifying the plate. Required at chart time. One point per distinct value of it.

  • order – an explicit run-order or date column. Strongly preferred: without it the order is inferred from plate and every run-based rule rests on that inference.

  • control_column – the column saying what each row is ("condition", "gene", "well_type"). None means the table is already only the control.

  • control_levels – which level(s) of control_column are the control being charted. Required whenever control_column is set — a control column with no level named is a filter with no predicate.

  • positive_levels – the positive control, for zprime_frame().

  • negative_levels – the negative control, likewise.

  • estimator – one of ESTIMATORS.

  • rules – which of RULES_ALL to run.

  • baseline_n – Phase I length when neither baseline_plates nor baseline_before is given.

  • baseline_plates – an explicit Phase I set, by plate id.

  • baseline_before – everything strictly before this value of the order column is Phase I. A date, for a campaign whose baseline is “the first week”.

  • reestimate – when the baseline itself trips a rule, drop the flagged plates and estimate once more. Off by default; see the module docstring for why silently deleting the inconvenient plates is not the default behaviour of anything here.

Raises:

ControlChartError – on an unknown estimator or rule id, a baseline shorter than MIN_BASELINE, or a control column with no level named — at the point the spec is built, not at render time.

__post_init__() → None[source]

Normalise the columns and levels, and validate the chart settings.

Raises:

ControlChartError – if the estimator is unknown; if a rule is not a number, or not one of rules 1-8; if the baseline is shorter than the module’s minimum – MR-bar over fewer moving ranges is already an opinion about one or two plates rather than an estimate; or if control_column and control_levels are given without each other, which says where to look without saying what for, or the reverse.

describe() → str[source]

One line, for a caption.

classmethod from_dict(payload: Mapping[str, Any]) → ControlChartSpec[source]

Rebuild from to_dict().

Unknown keys ignored, missing keys defaulted, so a configuration written by another build still opens: a QC chart nobody can reload is a QC chart nobody re-runs.

Parameters:

payload – a mapping as written by to_dict(); unknown keys are ignored and missing ones take their defaults.

classmethod from_json(text: str) → ControlChartSpec[source]

Rebuild from to_json().

Parameters:

text – a JSON string as written by to_json().

to_dict() → Dict[str, Any][source]

A plain JSON-able dict. Every field, always — a stable schema beats a compact one for something a QC configuration is stored as.

to_json() → str[source]

to_dict() as sorted JSON.

with_baseline(*, n: int | None = None, plates: Sequence[str] | None = None, before: str | None = None) → ControlChartSpec[source]

A copy with a different Phase I.

with_columns(*, value: str | None = None, plate: str | None = None, order: str | None = None) → ControlChartSpec[source]

A copy pointed at different columns.

with_control(column: str | None, levels: Sequence[str]) → ControlChartSpec[source]

A copy charting a different control.

Parameters:
  • column – the column holding the control labels, or None.

  • levels – the labels in that column that mark control wells.

with_estimator(estimator: str) → ControlChartSpec[source]

A copy using a different sigma estimator.

Parameters:

estimator – the sigma estimator name, e.g. 'auto', 'moving_range', 'subgroup_s', 'robust' or 'mad'; stored without checking.

with_rules(rules: Sequence[int]) → ControlChartSpec[source]

A copy running a different rule set.

Parameters:

rules – the rule numbers (1 to 8) to run.

class spacr.qt.widgets.control_chart.Violation[source]

One rule firing, over one span of plates.

A maximal span, not a sliding window’s worth of overlapping alarms: twelve points in a row above the centre line is one violation of rule 2 covering twelve plates, which is what a reader needs, rather than four violations covering nine plates each, which is an artefact of the implementation.

Parameters:
  • rule – the rule number, 1-8.

  • start – first index of the span, into ControlChartResult.plates.

  • end – last index of the span, inclusive.

  • points – the indices actually flagged. Equal to the whole span for the run rules; only the qualifying points for rules 5 and 6, where a window can trigger with an innocent point inside it.

  • plates – the plate ids of points.

  • side – ABOVE, BELOW or EITHER. For rule 3 it is the direction of travel (ABOVE for a rising trend) rather than a side of the centre line, because a trend can run entirely on one side of the centre or cross it and is a violation either way.

  • in_baseline – whether any flagged point is inside Phase I. Limits estimated from an out-of-control baseline are not limits, so this is carried per violation rather than recomputed by every reader.

describe() → str[source]

The sentence ControlChartResult.report() prints.

Names the rule by number and in words, because a report that says “rule 6” to a reader who has not memorised Nelson’s numbering has said nothing, and one that says only “four of five beyond 1 sigma” cannot be cross-referenced against anyone else’s chart.

where() → str[source]

The plates in words — "plate P17" or "plates P22-P30".

property span: int[source]

Plates from start to end, inclusive.

spacr.qt.widgets.control_chart.c4(n: int) → float[source]

The unbiasing constant for the SD of a subgroup of n.

c4(n) = sqrt(2/(n-1)) · Γ(n/2) / Γ((n-1)/2), which is E[s] / sigma for a normal sample of size n: the sample SD is a biased estimate of sigma, badly so for the small subgroups a plate’s control wells make, and S-bar / c4(n) is the correction.

Computed from the gamma function rather than looked up, because the published table stops at 25, does not interpolate, and is a transcription error waiting to happen in a module nobody re-derives. The values agree with the published ones to every digit they print: c4(2) = 0.7979, c4(5) = 0.9400, c4(10) = 0.9727.

Parameters:

n – the subgroup size, converted to int; must be at least 2.

Raises:

ControlChartError – for n < 2. A subgroup of one has no spread to unbias, which is the whole reason the I-MR chart exists.

spacr.qt.widgets.control_chart.candidate_key_columns(frame: pandas.DataFrame) → Tuple[str, ...][source]

The columns worth offering as the plate id or the control label.

Categorical by the same classifier, plus anything named like a plate or a date, because plateID is high-cardinality enough in a big campaign that the classifier calls it a key and skips it — and a key is exactly what is wanted here. This is the one place in spaCR where “identifies rather than describes” is a recommendation rather than a disqualification.

Parameters:

frame – the per-well table whose columns are classified by spacr.qt.widgets.graph_spec.column_kinds(). Columns named like a plate, date, time, batch, run or order are added.

spacr.qt.widgets.control_chart.candidate_value_columns(frame: pandas.DataFrame) → Tuple[str, ...][source]

The columns worth offering as the charted measurement, sorted.

Continuous by spacr.qt.widgets.graph_spec.column_kinds(), which is the one column classifier in this codebase — the Local Data Filter’s rule, re-read. Reused rather than re-derived: the Graph Builder, the PCA screen and this one must agree about what cell_count is, or the same table presents three different mental models depending on which screen is open.

Parameters:

frame – the per-well table whose columns are classified by spacr.qt.widgets.graph_spec.column_kinds().

spacr.qt.widgets.control_chart.control_chart(frame: pandas.DataFrame, spec: ControlChartSpec | None = None) → ControlChartResult[source]

Chart spec’s control value plate by plate over run order.

The whole policy is in the module docstring; the short version is that one point is drawn per plate in run order, sigma comes from short-term variation and never from the SD of the series, the limits are estimated from a stated baseline and applied forward, and every rule that fires is reported by number, in words, and with the plates it fired on.

Parameters:

frame – the per-well table holding the spec’s value, plate, order and control columns.

Raises:

ControlChartError – whenever the chart would be meaningless — with the reason and the way out in the message.

spacr.qt.widgets.control_chart.sd_reference_limits(values: Sequence[float]) → Tuple[float, float, float, float][source]

(centre, sigma, lower, upper) the wrong way — mean ± 3 SD.

This is not an estimator option and never will be. It exists so a result can print what the SD route would have said next to what the moving-range route did say, because the central argument of this module is far more convincing as two intervals side by side than as a paragraph: on a drifting series the SD limits are wide enough to contain the drift, and a chart drawn with them reports “in control” all the way down.

Uses the sample SD (ddof=1) over every value given, which is exactly the mistake being illustrated — limits computed from all the data, including the excursion they are supposed to detect.

Parameters:

values – the plotted values; non-finite ones are ignored.

spacr.qt.widgets.control_chart.zprime_chart(frame: pandas.DataFrame, spec: ControlChartSpec) → ControlChartResult[source]

zprime_frame(), charted with the same estimator, rules and baseline.

A thin composition rather than a second engine: the Z’ series is one value per plate, which is an individuals chart, which is what control_chart() already is. The order is carried as an explicit integer column, so the Z’ chart never reports an inferred order — the ordering decision was taken once, upstream.

Parameters:
  • frame – the per-well table holding the spec’s value, plate, order and control columns.

  • spec – the chart settings; must name control_column, positive_levels and negative_levels. Its estimator, rules and baseline are reused for the Z’ chart.

spacr.qt.widgets.control_chart.zprime_frame(frame: pandas.DataFrame, spec: ControlChartSpec) → pandas.DataFrame[source]

Per-plate Z-factor, one row per plate, in run order.

Z' = 1 - 3(sd_pos + sd_neg) / |mean_pos - mean_neg| — the number a screener actually watches, and the one that says whether a plate could have detected anything at all. Charting it with the same limits and the same rules is the same question one level up: a Z’ that slides from 0.7 to 0.3 over a campaign is a screen that stopped working, plate by plate, with no single plate looking wrong.

Both controls need at least two wells on a plate for the SD to exist, so a plate with a single positive well produces no Z’ and is left out rather than given a zero.

Parameters:
  • frame – the per-well table holding the spec’s value, plate, order and control columns.

  • spec – the chart settings; must name control_column, positive_levels and negative_levels.

Returns:

columns plate, order, order_index, zprime, separation, and the per-control means, SDs and n.

Raises:

ControlChartError – when the spec does not name both controls.

Nested helpers

control_chart._finish(centre: float, sigma_within: float, baseline: np.ndarray, excluded: Tuple[str, ...]) → ControlChartResult

Assemble the result, flagging a degenerate spread.

A sigma at or below the tolerance means every point is a violation, so it is reported as DEGENERATE rather than charted – limits drawn from no spread say nothing about the process.

spacr/qt/widgets/control_chart.py:1719