spacr.point_patterns¶
Is this well clustered, dispersed, or neither – against the right null.
WHAT THE PACKAGE COULD NOT ASK. spacr.object_distances says how far
one object is from another, and spacr.bystanders says whether a cell
has an infected neighbour. Neither answers the question a focus of
infection actually poses: are the infected cells in this well arranged more
tightly than chance would put them? That is Ripley’s K, and it is the
statistic that turns “these two cells are close” into “this well is
clustered”.
THE NULL IS THE WHOLE STATISTIC. K is meaningless without something to compare it against, and the comparison has to respect the shape the cells were allowed to occupy. A well is not a rectangle, a field with a torn monolayer is not a rectangle, and a null that scatters points over the bounding box will call every mask with a concave edge “clustered” – the points are crowded into the tissue because there is nowhere else to be. Everything here takes a mask and both the estimator and the simulated null stay inside it.
EDGE CORRECTION IS NOT OPTIONAL AND IT IS THE PART THAT GETS SKIPPED. A
point near the boundary has part of its neighbourhood outside the window,
so it finds fewer neighbours than it has; uncorrected, that deficit reads
as dispersion, and it is worst at exactly the radii anyone cares about.
ripley_k() applies the reduced-sample (border) correction by
default: at radius r, only points at least r inside the boundary
contribute. It costs precision at large r – honestly, by reporting how
many points were left – and it is exact for any window shape, which the
analytic isotropic corrections are not.
THE FRAME EDGE IS A BOUNDARY TOO. The distance-to-boundary field is computed on a mask padded with one row of background all round, so a mask that runs to the edge of the image is treated as cut off there, because it is. Without the pad, cells at the top of the image would be scored as though the monolayer continued past it.
Functions¶
|
The spread of L - r for |
|
Ripley's K and L for one point pattern in one window. |
|
Observed K against a null that respects the mask. |
|
Radii worth computing K at, for this window. |
Module Contents¶
- spacr.point_patterns.csr_envelope(mask, n_points: int, radii, *, spacing=None, simulations: int = DEFAULT_SIMULATIONS, seed: int = 0, correction: str = 'border') pandas.DataFrame[source]¶
The spread of L - r for
n_pointsscattered at random inmask.- Parameters:
mask – the same window the observed pattern was measured in.
n_points – how many points to scatter – the observed count.
radii – the same radii.
spacing – the same pixel size.
simulations – how many random patterns.
seed – fixed, so an envelope is reproducible.
correction – the same correction as the observation used.
- Returns:
r,lo,hiandmeanofl_minus_r.
THE POINTS ARE SCATTERED IN THE MASK, NOT IN ITS BOUNDING BOX. That is the entire reason this function exists rather than a lookup of the analytic CSR variance: the analytic result assumes a rectangle, and a well is a disc with holes in it.
SAME CORRECTION ON BOTH SIDES. An observation corrected against an envelope that was not – or the reverse – compares two different estimators and the difference between them will look like biology.
- spacr.point_patterns.ripley_k(points, mask, radii, *, spacing=None, correction: str = 'border') pandas.DataFrame[source]¶
Ripley’s K and L for one point pattern in one window.
- Parameters:
points –
(n, 2)centroids in array (row, column) order, asspacr.object_distances._centroids()returns them.mask – boolean array, true where a point could have been.
radii – the radii to evaluate, in the same units as
spacing.spacing – pixel size;
Nonemeans the radii are in pixels.correction –
"border"for the reduced-sample correction, or"none"to see what the correction is worth.
- Returns:
r,k,l,l_minus_randn_contributing.
l_minus_rIS THE COLUMN TO READ. K grows as r squared, so two curves are impossible to compare by eye; L is its variance-stabilised square root, and under complete spatial randomness L(r) = r exactly, so L(r) - r is a flat line at zero whatever the density. Above zero is clustering, below is dispersion, and the size of the departure is in the units of the image.n_contributingIS NOT DECORATION. It is the number of points the border correction kept at that radius, and when it falls into single figures the row is arithmetic, not evidence.
- spacr.point_patterns.ripley_test(points, mask, radii=None, *, spacing=None, simulations: int = DEFAULT_SIMULATIONS, seed: int = 0, correction: str = 'border') pandas.DataFrame[source]¶
Observed K against a null that respects the mask.
- Parameters:
points – the observed centroids.
mask – the window.
radii –
Noneaskssuggest_radii().spacing – pixel size.
simulations – how many random patterns behind the envelope.
seed – fixed, so the answer is reproducible.
correction – applied to both the observation and the null.
- Returns:
the columns of
ripley_k(), pluslo,hi,clusteredanddispersed.
clusteredANDdispersedARE PER RADIUS AND THAT IS NOT ONE TEST. Twenty radii give twenty chances to leave a 2% envelope, so a single excursion is close to meaningless; a run of them at adjacent radii is the thing to believe. The columns are deliberately named for what they are – an excursion at this r – rather than “significant”.
- spacr.point_patterns.suggest_radii(mask, *, spacing=None, count: int = 20) numpy.ndarray[source]¶
Radii worth computing K at, for this window.
- Parameters:
mask – boolean array of where a point was allowed to be.
spacing – pixel size, so the radii come back as real lengths.
count – how many radii.
- Returns:
countradii, evenly spaced, ending atMAX_RADIUS_FRACTIONof the window’s shorter side.
NOT STARTING AT ZERO. K(0) is 0 by construction and carries no information; including it only puts a point where every curve agrees.