spacr.plaque_papers¶
Plaque measurements from published papers: a DOI, PMID or PDF in, rows out.
Given a paper, this module finds its figures, finds the plaque images inside them with the zoo’s well detector, reads the text printed around each image, works out which condition each image shows, segments and measures the plaques, and writes everything to a SQLite database together with where every value came from.
THE CONDITION IS THE HARD PART, and it is read two independent ways, always both, because two readings that agree are evidence and one is a guess:
the label – the text nearest the image inside the figure: the column header above it, the row label beside it, anything printed beneath it. A figure that prints
TATi | iGAPDH2over two columns and-ATc | +ATcbeside two rows gives each of its four crops two words, and those two words are the condition.the legend – the panel letter nearest the image keys into the figure legend, which is fetched from Europe PMC, read from the PDF, or pasted by the user when neither has it.
When both are present they are both kept; a label whose words the legend also
uses is strong, one the legend does not echo is medium – a legend that
says “under indicated conditions” is not a disagreement. They DISAGREE when the
panel’s legend passage names conditions this figure prints on its other
images and none of this image’s own: that sets conflict, with the words
that caused it, and the row is kept for a person to settle rather than one
reading being preferred. When neither is present the image is named by its
figure and its row and column in the panel, and marked weak so a dataset
can leave those rows out.
SIZES. A plaque area in pixels depends on the paper’s printing, so every plaque also carries its area relative to the median plaque in the same panel, and an area in mm^2 only when there is a ruler: a scale bar read in or under the image (its length from the label beside it, or from the legend’s “scale bar, 1 mm”), or a whole well of a plate format the settings or the legend state. A stated magnification is recorded but is never a ruler, because a printed figure has been rescaled since the picture was taken. Rows with no ruler say their sizes are in pixels.
DUPLICATES. A figure is identified by its paper (the DOI when there is one)
and its image hash. The same image seen again – the same bytes, or the same
pixels in another file format, under another name or another paper, such as a
preprint and its published version – is measured once and recorded in the
duplicates table against the copy that was measured. A paper already
measured from one source (Europe PMC) is not measured again from another (its
PDF). An image that merely LOOKS the same (a near-identical perceptual hash,
as a re-encoded or resized copy would have) is recorded as a possible
duplicate and still measured, because on figures that is a resemblance, not
a proof.
TEXT. A PDF’s own text layer is used for legends and labels before any OCR; OCR is asked only when the text layer has nothing near the plaque images, which is what a figure pasted into the PDF as one picture looks like.
The optional pieces – ultralytics for the detector, rapidocr for the
text in figure images, pdfplumber for a PDF’s text layer – are the
spacr[papers] extra. Everything here that touches them takes the callable
as a parameter, so the logic is testable without them.
Exceptions¶
The figure reader must be installed, or installed again, first. |
Classes¶
What one plaque image shows, and how that was decided. |
|
One figure image and its legend. |
|
One paper, however it was named. |
|
One plaque image the detector found inside a figure. |
|
How the text around a plaque image is turned into its condition. |
|
One piece of text in a figure, in image pixels. |
Functions¶
|
Propose a condition for every plaque image in one figure. |
|
An annotation as plain values, for JSON and the review dialog. |
|
Put a person's edits and approvals onto one figure's proposals. |
|
Parse an optional finite calibration value without treating invalid input as missing. |
|
Return per-well size, scale and timing metadata for preview and saved tables. |
|
Ask at the terminal for a legend nobody could fetch. |
|
Show each proposed condition at the terminal and take OK, an edit, or skip. |
|
The detector call Figure mode uses: in process, or in the reader. |
|
Download a paper's figures from Europe PMC with their legends. |
|
Put a paper's figures in a folder Plaque Assay's Figure mode can read. |
|
The folders Figure mode reads for |
|
Every page of a PDF as a figure image, with its text layer attached. |
|
Every figure image directly in |
|
Plaque images in one figure, asked at every inference size given. |
|
Figure legends found in a paper's running text. |
|
|
|
Figure mode of Plaque Assay: find, read, annotate and measure a folder. |
|
Measure the plaques in every figure of every paper given. |
|
One row per plaque in a segmented plaque image. |
|
Open (and create, or bring up to date) the results database. |
|
Say what kind of reference a user typed. |
|
Conditions a person edited or approved, keyed by figure and image. |
|
Figure legends supplied beside a folder of figures. |
|
The text printed in a figure image, with where it is. |
|
|
|
The figure reader's own environment when it is installed, else None. |
|
Whether Figure mode can read now, and what to do when it cannot. |
|
Read the text around each grid of plaque images again, enlarged. |
|
Look a reference up in Europe PMC, or describe a local PDF. |
|
Hex sha256 of some bytes. |
|
A figure legend cut into one passage per panel letter. |
|
The text a reader would take as this image's label. |
|
|
|
Save reviewed conditions for |
Module Contents¶
- exception spacr.plaque_papers.ReaderNeedsInstall(message: str, *, reinstall: bool = False)[source]¶
Bases:
ImportErrorThe figure reader must be installed, or installed again, first.
An
ImportError, so every caller that already reported a missing reader still does; Plaque Assay catches this one to offer the install in place.- Parameters:
message – what is missing and why.
reinstall – True when the reader is installed but was built before a package it now needs was pinned.
Keep whether this is a first install or a reinstall.
- class spacr.plaque_papers.Annotation[source]¶
What one plaque image shows, and how that was decided.
- Parameters:
region – the plaque image’s box in the figure.
panel – the panel letter read nearest the image, or
None.label_text – strategy 1 – the text around the image.
legend_text – strategy 2 – the legend’s sentence for the panel.
condition – the proposed condition.
source –
'label+legend','label','legend','manual'or'position'.strength –
'strong'when both readings agree,'medium'for one reading,'weak'for position only,'manual'when a person wrote it.conflict – True when the two readings disagree – the panel’s legend names conditions the figure prints elsewhere and none of this image’s (see
annotate_regions()).conflict_terms – the legend’s words that caused
conflict.approved –
True/Falseonce a person has reviewed it,Nonewhen nobody has.
- class spacr.plaque_papers.Figure[source]¶
One figure image and its legend.
- Parameters:
path – the image file on disk.
label – the figure’s name in the paper, e.g.
'Fig 7'.caption – the legend, or
''when none was found.legend_source – where
captioncame from:'europepmc','pdf','pasted'or'none'.words – text already known with coordinates, e.g. a PDF’s text layer; OCR fills this when it is empty.
- class spacr.plaque_papers.Paper[source]¶
One paper, however it was named.
- Parameters:
key – the identifier rows are filed under: the DOI when there is one, else the PMC id, else the PDF’s sha256.
source –
'europepmc','pdf'or'folder'.doi_from – where the DOI was read:
'europepmc','pdf metadata'or'pdf text'; None when there is no DOI.
- class spacr.plaque_papers.Region[source]¶
One plaque image the detector found inside a figure.
- Parameters:
x0 – left edge of the box, in figure pixels.
y0 – top edge of the box.
x1 – right edge of the box.
y1 – bottom edge of the box.
confidence – the detector’s confidence, 0 to 1.
sizes – every inference size that found it, so a run records which size a region depended on.
- class spacr.plaque_papers.TextOptions[source]¶
How the text around a plaque image is turned into its condition.
Text near a panel is usually read correctly but not always assigned to the right image, so every choice the assignment makes is exposed here. These are the knobs the reading has; the defaults are the values measured on PMC9744290 Fig 6 and Fig 7 (8 of 8 read).
- Variables:
reach_above – how far above the grid a column header may sit, in image heights.
reach_left – how far left of the grid a row label may sit, in image widths.
reach_below – how far below the grid text may sit, in image heights.
use_above – take the column header.
use_left – take the row label.
use_below – take text printed under the images.
panel_reach – how far up and left of the grid’s corner a panel letter may sit, in image sizes.
min_confidence – OCR words scoring below this are ignored.
ignore – regular expressions; a word matching any is ignored (scale bars such as
5 μm, axis numbers).order – which labels come first in the condition, e.g.
("above", "left", "below").separator – what joins the labels into one condition.
reread – read the text around each grid again, enlarged.
reread_scale – how much to enlarge for that second reading.
- class spacr.plaque_papers.Word[source]¶
One piece of text in a figure, in image pixels.
- Parameters:
text – the text as read.
x0 – left edge of its box.
y0 – top edge of its box.
x1 – right edge of its box.
y1 – bottom edge of its box.
confidence – the reader’s confidence in the text, 0 to 1.
- spacr.plaque_papers.annotate_regions(regions: Sequence[Region], words: Sequence[Word], *, caption: str = '', figure_label: str = '', options: TextOptions | None = None) List[Annotation][source]¶
Propose a condition for every plaque image in one figure.
Both strategies run for every image. See the module docstring for how they combine and what
source,strengthandconflictmean.- Parameters:
regions – the figure’s plaque images.
words – the figure’s text.
caption – the figure legend,
''when unknown.figure_label – the figure’s name, used when nothing else is known.
options – how the text is read (
TextOptions); defaults toDEFAULT_TEXT_OPTIONS.
- Returns:
one annotation per region, in the order given.
- spacr.plaque_papers.annotation_as_dict(annotation: Annotation) Dict[str, Any][source]¶
An annotation as plain values, for JSON and the review dialog.
- Parameters:
annotation – the annotation.
- Returns:
its fields, the region flattened.
- spacr.plaque_papers.apply_overrides(stem: str, annotations: List[Annotation], overrides: Mapping[Tuple[str, int], Mapping[str, Any]], *, confirm_each: bool = False) List[Annotation][source]¶
Put a person’s edits and approvals onto one figure’s proposals.
A conflict flag stays on an edited image: the record keeps that the two readings disagreed, and the edit and OK record how a person settled it.
- Parameters:
stem – the figure’s file stem.
annotations – the proposals, in reading order.
overrides – from
read_annotation_overrides().confirm_each – when True, an image nobody approved is marked
approved=Falseand is not measured.
- Returns:
the same annotations.
- spacr.plaque_papers.calibration_number(value, *, name, allow_zero=False)[source]¶
Parse an optional finite calibration value without treating invalid input as missing.
- spacr.plaque_papers.calibration_values(annotation, scale=None, formation_hours=None)[source]¶
Return per-well size, scale and timing metadata for preview and saved tables.
The well diameter is the mean detected bounding-box extent, not a fitted circle or a segmentation-model diameter. It remains available without a physical ruler. Time is entered metadata, never inferred from plaque size.
- spacr.plaque_papers.console_legend_prompt(figure: Figure, annotations: Sequence[Annotation], *, ask: Callable[[str], str] = input) str | None[source]¶
Ask at the terminal for a legend nobody could fetch.
- Parameters:
figure – the figure whose legend is missing.
annotations – what was read so far, to show which panels matter.
ask – the prompt function.
- Returns:
the pasted legend, or None to annotate by hand instead.
- spacr.plaque_papers.console_review(figure: Figure, annotations: List[Annotation], *, ask: Callable[[str], str] = input) List[Annotation][source]¶
Show each proposed condition at the terminal and take OK, an edit, or skip.
- Parameters:
figure – the figure the annotations belong to.
annotations – the proposals; changed in place.
ask – the prompt function.
- Returns:
the same annotations, reviewed.
- spacr.plaque_papers.default_detect() Callable[source]¶
The detector call Figure mode uses: in process, or in the reader.
- Returns:
spacr.plaque.detect_wells()when ultralytics imports in spaCR itself, elsereader_detect().
- spacr.plaque_papers.fetch_figures(paper: Paper, dest: Any, *, get: Callable | None = None) List[Figure][source]¶
Download a paper’s figures from Europe PMC with their legends.
The images come from the
supplementaryFilesbundle, which carries the main figures as well as the supplements; one file per figure survives, the largest, because journals also ship thumbnails and a detector shown a thumbnail has not been asked the question. Legends come from the full-text XML.- Parameters:
paper – a paper with a PMC id.
dest – directory for the images; created if missing.
get –
fn(url, **kw) -> response; defaults to requests.
- Returns:
the figures, sorted by file name; empty when Europe PMC has no bundle for the paper.
- spacr.plaque_papers.fetch_paper_to_folder(reference: Any, dest: Any, *, get: Callable | None = None, pdf_opener: Callable | None = None) Dict[str, Any][source]¶
Put a paper’s figures in a folder Plaque Assay’s Figure mode can read.
When a figure’s panel letters are detected, its legend is gathered automatically. A DOI, PMID or PMC id is fetched from Europe PMC with its JATS legends; a PDF is rendered page by page with the legends found in its text. Either way the images land in
destand each figure’s legend indest/legends.csv, which Figure mode andmeasure_figure_folder()read – so a paper becomes an ordinary folder of figures, annotated with its own legends. A PDF’s text layer is kept indest/text_layer.jsonso its labels are read from the PDF itself before any OCR, and the paper’s identity indest/paper.json.- Parameters:
reference – a DOI, PMID, PMC id or PDF path.
dest – the folder to fill; created if missing.
get – HTTP getter for Europe PMC.
pdf_opener – passed to
figures_from_pdf().
- Returns:
{'folder', 'paper', 'figures', 'with_legend', 'licence', 'text_layer'}–text_layercounts the pages that have one.
- spacr.plaque_papers.figure_folders(src: Any) List[pathlib.Path][source]¶
The folders Figure mode reads for
src, in reading order.A paper folder, or a folder of figure images, is read as it is. A folder that holds paper folders – several PDFs read at once, each into a folder of its own – is read paper by paper: its own figure images first when it has any, then each paper folder by name.
- Parameters:
src – the folder Figure mode was given.
- Returns:
the folders;
[src]when it holds no paper folders.
- spacr.plaque_papers.figures_from_pdf(pdf: Any, dest: Any, *, dpi: int = 200, opener: Callable | None = None) List[Figure][source]¶
Every page of a PDF as a figure image, with its text layer attached.
A page is rendered whole rather than its embedded images extracted, because a published figure is often assembled from many embedded images and vector labels, and the labels are what the condition is read from. The text layer comes back as
Wordobjects in the rendered image’s pixels, so no OCR is needed for a PDF that has one.Rendering needs
pdfplumber: in this process when it is installed here, else in the figure reader’s own environment, which installs it.- Parameters:
pdf – the PDF path.
dest – directory for the page images.
dpi – render resolution.
opener –
fn(path) -> pdfplumber-like document; defaults topdfplumber.open().
- Returns:
one figure per page, legend attached from the text when the page names a figure.
- Raises:
ImportError – when pdfplumber is available neither way.
- spacr.plaque_papers.figures_in_folder(src: Any, legends: Mapping[str, str] | None = None, text_layer: Mapping[str, List[Word]] | None = None) List[Figure][source]¶
Every figure image directly in
src, with its legend when known.- Parameters:
src – the folder.
legends –
{stem: legend}.text_layer –
{stem: [Word]}, a PDF’s words for its pages; read fromsrc/text_layer.jsonwhen None.
- Returns:
the figures, by file name.
- spacr.plaque_papers.find_plaque_regions(image: numpy.ndarray, weights: Any, *, imgsz: Sequence[int] = DEFAULT_IMGSZ, confidence: float = 0.25, iou: float = 0.5, detect: Callable | None = None) List[Region][source]¶
Plaque images in one figure, asked at every inference size given.
One size is not right for everything a literature corpus prints: at 640 the detector finds whole faint panels and misses dilution-spot strips, at 1280 the reverse (
features/data/424_imgsz_sweep_2026-09-20.json). So each size is asked and the boxes are merged, a box found at two sizes kept once with the higher score and both sizes recorded.- Parameters:
image – the figure,
H x W x 3RGB.weights – detector checkpoint path.
imgsz – the inference sizes to ask at.
confidence – minimum detector score.
iou – boxes overlapping more than this are the same region.
detect –
fn(image, weights, confidence=, imgsz=, min_axis_ratio=) -> boxes with x0, y0, x1, y1, confidence; defaults tospacr.plaque.detect_wells(). Nothing is dropped for not being square: a figure crop need not be a round well.
- Returns:
the regions, top-to-bottom then left-to-right.
- spacr.plaque_papers.legends_from_text(text: str) Dict[str, str][source]¶
Figure legends found in a paper’s running text.
- Parameters:
text – the text of a PDF, pages joined.
- Returns:
{figure number: legend}, the LONGEST paragraph startingFig Nfor each number, because the legend is the paragraph and a short hit is a mention in the body.
- spacr.plaque_papers.main(argv: Sequence[str] | None = None) int[source]¶
python -m spacr.plaque_papers DOI|PMID|PMC|PDF ... --dst DIR.- Parameters:
argv – arguments; defaults to
sys.argv[1:].- Returns:
the exit status.
- spacr.plaque_papers.measure_figure_folder(src: Any, dst: Any = None, *, detector: str = DEFAULT_DETECTOR, segmenter: str = DEFAULT_SEGMENTER, imgsz: Sequence[int] = DEFAULT_IMGSZ, confidence: float = 0.25, confirm_each: bool = False, plate_format: str | None = None, pixels_per_um=None, formation_hours=None, growth_settings=None, legends: Any = None, annotations: Any = None, read_text: Callable | None = None, detect: Callable | None = None, segment: Callable | None = None, text_options: TextOptions | None = None) Dict[str, Any][source]¶
Figure mode of Plaque Assay: find, read, annotate and measure a folder.
Nothing here stops to ask. Legends come from
legends(default<src>/legends.csv), a PDF’s text layer from<src>/text_layer.json, the paper from<src>/paper.json, and a person’s edits and approvals fromannotations(default<src>/figure_annotations.csv), which the Figure preview writes. Withconfirm_eachon, only approved images are measured and the rest are counted as waiting. A figure measured again replaces its earlier rows; the same image under a second name is measured once and recorded as a duplicate – the copy with a legend is the one measured, since figures with legends are taken first.- Parameters:
src – the folder of figure images.
dst – output folder; default
<src>/plaque_figures.detector – zoo key or checkpoint for the plaque-image detector.
segmenter – zoo key or checkpoint for the plaque segmenter.
imgsz – detector inference sizes.
confidence – minimum detector score.
confirm_each – measure only images a person approved.
plate_format – plate format for whole-well images.
pixels_per_um – optional positive manual pixel scale; per-well annotations override this value, then automatic rulers are considered.
formation_hours – optional nonnegative elapsed time in hours; per-well manual times take precedence. Stored as metadata, not inferred.
growth_settings – optional experimental growth settings consumed by
spacr.plaque_growth.estimates_from_settings(); disabled by default.legends – CSV of legends, see
read_legends().annotations – CSV of reviews, see
read_annotation_overrides().read_text –
fn(path) -> [Word].detect – passed to
find_plaque_regions().segment –
fn(crop) -> labels.text_options – how the text is read (
TextOptions).
- Returns:
the summary, with
awaiting_approvaladded.
- spacr.plaque_papers.measure_plaques_from_papers(references: Iterable[Any], dst: Any, *, detector: str = DEFAULT_DETECTOR, segmenter: str = DEFAULT_SEGMENTER, imgsz: Sequence[int] = DEFAULT_IMGSZ, confidence: float = 0.25, confirm_each: bool = False, plate_format: str | None = None, ask_legend: Callable | None = None, review: Callable | None = None, read_text: Callable | None = None, detect: Callable | None = None, segment: Callable | None = None, get: Callable | None = None, pdf_opener: Callable | None = None) Dict[str, Any][source]¶
Measure the plaques in every figure of every paper given.
A figure already measured – the same image in another paper, or the same paper from another source – is recorded in the
duplicatestable and not measured again (see the module docstring).- Parameters:
references – DOIs, PMC ids, PMIDs or PDF paths.
dst – output folder:
plaque_papers.db,figures/andcrops/are written under it.detector – zoo key or checkpoint for the plaque-image detector.
segmenter – zoo key or checkpoint for the plaque segmenter.
imgsz – detector inference sizes, all asked, results merged.
confidence – minimum detector score.
confirm_each – when True, every proposed condition is shown to a person (
review) before it is stored as approved.plate_format – plate format for whole-well images, which gives them a ruler; None lets the legend name one, and otherwise keeps every area in pixels and ratios.
ask_legend –
fn(figure, annotations) -> legend or None, called when a panel letter was read but no legend could be fetched. Defaults toconsole_legend_prompt().review –
fn(figure, annotations) -> annotations, called per figure whenconfirm_each. Defaults toconsole_review().read_text –
fn(image path) -> [Word]; defaults toread_words(). Not called for a figure whose PDF text layer already has the words around its plaque images.detect – passed to
find_plaque_regions().segment –
fn(crop) -> labels; defaults to a Cellpose model loaded fromsegmenter.get – HTTP getter for Europe PMC.
pdf_opener – passed to
figures_from_pdf().
- Returns:
a summary: papers, figures, figures skipped as already measured, duplicates (not measured) and possible duplicates (measured), regions, plaques, conflicts, regions with a ruler, figures read from a text layer, the run id and the database path.
- spacr.plaque_papers.measure_region(labels: numpy.ndarray, *, px_per_mm: float | None = None) List[Dict[str, Any]][source]¶
One row per plaque in a segmented plaque image.
- Parameters:
labels – the segmenter’s label image, 0 = background, or
(labels, per_label_metrics)from a diagnostic segmenter.px_per_mm – the image’s scale, when it has a ruler.
- Returns:
[{'label', 'area_px', 'area_mm2'}];area_mm2is None without a ruler.
- spacr.plaque_papers.open_database(path: Any) sqlite3.Connection[source]¶
Open (and create, or bring up to date) the results database.
Every table in
TABLESis created if missing, and a table an older spaCR wrote gets the columns it lacks, so rows already in it stay readable next to new ones.Waits up to 30 seconds for a lock another process holds, as
spacr.database_concurrency.connect()does, rather than failing at SQLite’s five-second default. Opened withsqlite3.connectand not with that helper because the callers commit once per figure: the helper’s autocommit mode would make a figure’s deletes and inserts separate transactions, and a crash between them would leave half a figure in the table.- Parameters:
path – the
.dbfile.- Returns:
an open connection.
- spacr.plaque_papers.parse_reference(ref: Any) Dict[str, str][source]¶
Say what kind of reference a user typed.
- Parameters:
ref – a DOI (bare,
doi:or a doi.org URL), a PMC id, a PMID, or a path to a PDF.- Returns:
{'kind': 'doi'|'pmcid'|'pmid'|'pdf', 'value': ...}.- Raises:
ValueError – when it is none of these.
- spacr.plaque_papers.read_annotation_overrides(path: Any) Dict[Tuple[str, int], Dict[str, Any]][source]¶
Conditions a person edited or approved, keyed by figure and image.
- Parameters:
path – a CSV with
file,region(1-based, in reading order),conditionandapprovedcolumns – what the Figure preview saves. Its other columns are the record of what was proposed and are not read back.- Returns:
{(stem, region): {'condition': str, 'approved': bool}}.
- spacr.plaque_papers.read_legends(path: Any) Dict[str, str][source]¶
Figure legends supplied beside a folder of figures.
- Parameters:
path – a CSV with
fileandlegendcolumns;fileis the figure’s file name or stem.- Returns:
{stem: legend}; empty when the file does not exist.
- spacr.plaque_papers.read_words(image: Any, *, engine: Callable | None = None) List[Word][source]¶
The text printed in a figure image, with where it is.
- Parameters:
image – an image path or an
H x W x 3array.engine –
fn(image) -> (results, elapsed)in RapidOCR’s shape, each result[box points, text, confidence]; defaults to RapidOCR.
- Returns:
the words found, top-to-bottom then left-to-right.
- Raises:
ImportError – when no engine is given and RapidOCR is missing.
- spacr.plaque_papers.reader_detect(image: numpy.ndarray, weights: str, *, confidence: float = 0.25, imgsz: int = 640, min_axis_ratio: float = 0.0) List[Any][source]¶
spacr.plaque.detect_wells(), answered by the reader’s environment.- Parameters:
image – the figure, RGB.
weights – the detector checkpoint.
confidence – minimum score.
imgsz – the inference size.
min_axis_ratio – accepted for the same signature; nothing is dropped for its shape here, as in Figure mode’s in-process call.
- Returns:
boxes with
x0, y0, x1, y1, confidence.
- spacr.plaque_papers.reader_environment() str | None[source]¶
The figure reader’s own environment when it is installed, else None.
- Returns:
the environment folder.
- spacr.plaque_papers.reader_problem(*, pdf: bool = False) Tuple[str, str][source]¶
Whether Figure mode can read now, and what to do when it cannot.
Read from disk only – spaCR’s own packages and the reader’s install record (
spacr._segmentation_backends._stale_requirements()) – so it is cheap enough for the GUI thread.- Parameters:
pdf – the question is about reading a PDF, which needs pdfplumber; otherwise the detector and OCR.
- Returns:
('', '')when ready;('install', why)when the reader is not installed;('reinstall', why)when it is installed without a package spaCR now pins for it.
- spacr.plaque_papers.reread_around(image: numpy.ndarray, regions: Sequence[Region], words: Sequence[Word], *, engine: Callable | None = None, scale: int = 3) List[Word][source]¶
Read the text around each grid of plaque images again, enlarged.
A figure’s labels are small: a lone panel letter or a rotated
+ATcbeside a crop is a few pixels tall, and OCR run on the whole figure misses some (Fig 7 of PMC9744290 loses itsFand one row label). Each grid’s surroundings – one image size up and left, half down and right – are cut out, enlargedscaletimes and read again, and any word not already found is added in figure coordinates.- Parameters:
image – the figure,
H x W x 3.regions – its plaque images.
words – the words already read.
engine – a RapidOCR-shaped engine; defaults to RapidOCR.
scale – the enlargement.
- Returns:
wordsplus whatever the second reading found.
- spacr.plaque_papers.resolve_paper(ref: Any, *, get: Callable | None = None) Paper[source]¶
Look a reference up in Europe PMC, or describe a local PDF.
A PDF is filed under its DOI when it states one and Europe PMC knows that DOI (the three likeliest candidates are asked) – so the same paper given once as a DOI and once as a PDF is one paper, not two – and under its sha256 otherwise, with the DOI it names kept as unconfirmed.
- Parameters:
ref – anything
parse_reference()accepts.get –
fn(url, params=...) -> response; defaults to requests.
- Returns:
the paper. A reference Europe PMC does not know comes back with only the identifier filled in, which still lets a PDF be measured.
- spacr.plaque_papers.sha256_bytes(data: bytes) str[source]¶
Hex sha256 of some bytes.
- Parameters:
data – the bytes.
- Returns:
the digest.
- spacr.plaque_papers.split_legend(caption: str) Dict[str, str][source]¶
A figure legend cut into one passage per panel letter.
Legends mark panels as
(A),A,,A.orA), sometimes glued to the title (...growth.A, strategy), sometimes as a range (B-E, diagnostic PCRs). A letter is accepted as a marker only when it is the NEXT letter after the last marker, which is what stops(B, D)inside panel B-E’s own text, orP. falciparum, from being read as markers.- Parameters:
caption – the legend.
- Returns:
{letter: passage}in upper case, plus''for the title sentence before the first marker. Empty when no marker was found.
- spacr.plaque_papers.text_near(region: Region, words: Sequence[Word], *, regions: Sequence[Region] = (), reach: float = 1.0, options: TextOptions | None = None) Dict[str, Any][source]¶
The text a reader would take as this image’s label.
Distances are measured from the edge of the GRID the image sits in, not from the image itself, because a header printed once above a column labels every row in it and a row label printed once at the left labels every column:
above – the nearest line of text over the image’s column.
left – the nearest text level with the image’s row, left of the grid, including a label printed rotated.
below – the nearest line under the image’s column.
panel – a panel letter above and to the left of the grid’s top-left corner, within one image size of it. Further than that is another panel’s letter, and no letter is better than the wrong one.
Text inside any plaque image (a scale bar, an inset label) is not a label and is left out.
- Parameters:
region – the image.
words – every word in the figure.
regions – every plaque image in the figure, including
region.reach – how far past the grid to look, in image sizes; used when
optionsis not given.options – the reading’s settings (
TextOptions).
- Returns:
{'panel': letter or None, 'above': [...], 'left': [...], 'below': [...]}.
- spacr.plaque_papers.text_options_from_settings(settings: Mapping[str, Any]) TextOptions[source]¶
TextOptionsfrom a settings dict’stext_*keys.- Parameters:
settings – settings; missing keys keep the defaults.
- Returns:
the options.
- spacr.plaque_papers.write_annotation_overrides(path: Any, rows: Iterable[Mapping[str, Any]]) pathlib.Path[source]¶
Save reviewed conditions for
read_annotation_overrides().Rows for figures not in
rowsare kept, so reviewing one figure does not erase another’s approvals. Every column ofANNOTATION_COLUMNSis written; a row that does not give one leaves it blank.- Parameters:
path – the CSV to write.
rows – mappings with
file,region,condition,approvedand optionally the other annotation columns.
- Returns:
the path written.
Nested helpers¶
- _blocks.root(i: int) int¶
The block
ibelongs to, compressing the path on the way.spacr/plaque_papers.py:1131
- _cellpose_segmenter.segment(crop: np.ndarray) Any¶
The Cellpose label mask of one plaque image crop.
spacr/plaque_papers.py:2377
- _doi_candidates.add(doi: str) None¶
Count one sighting of a cleaned candidate.
spacr/plaque_papers.py:382
- _find_duplicate.found(row, match, distance=0, skip=True) Dict[str, Any]¶
The duplicate record for
row.spacr/plaque_papers.py:2238
- _find_duplicate.same_figure(row) bool¶
The row is this very file of this very paper.
spacr/plaque_papers.py:2234
- _grid_positions.ranks(centres: List[float], sizes: List[float]) List[int]¶
A 1-based rank per centre, sharing a rank within half a size.
spacr/plaque_papers.py:1431
- _label_words.size(line: List[Word]) int¶
Words in a line, counting inside each piece of text.
spacr/plaque_papers.py:2558
- measure_figure_folder.review(figure: Figure, found: List[Annotation]) List[Annotation]¶
Apply the saved overrides to a figure and count what still awaits approval.
spacr/plaque_papers.py:3192
- text_options_from_settings.patterns(value)¶
A tuple of patterns from a comma-separated string or a sequence.
spacr/plaque_papers.py:1225
- text_options_from_settings.pick(key, default, cast)¶
settings[key]cast bycast, ordefaultif blank or invalid.spacr/plaque_papers.py:1216
- text_options_from_settings.sides(value)¶
The valid sides named in
value, or the default order if none are.spacr/plaque_papers.py:1230