spacr.annotation

Join bundled Toxoplasma gene annotations to exported results.

The joined fields include gene names, predicted signal peptides and transmembrane domains, published phenotype scores, stage-specific expression and hyperLOPIT localisation.

All supported identifier forms are normalized to a bare gene number: TGGT1_224750, TGME49_224750, gene_fraction:gene[224750] and the guide identifier 224750_2 resolve to gene 224750 through gene_number(). Each source is reduced to one row per gene and merged with many_to_one validation, preventing annotation duplicates from multiplying result rows. A field unavailable in a source is omitted rather than emitted as an all-missing column.

Bundled annotation sources:

toxoplasma_metadata.csv 2.97 MB gene name, product, expression deeptmhmm.csv 1.11 MB signal peptide and transmembrane phenotype.csv 0.48 MB the published CRISPR fitness screens lopit.csv 0.20 MB hyperLOPIT/TAGM compartment uniprot.csv 0.13 MB accession

deeptmhmm.csv contains a DeepTMHMM analysis of 8,140 proteins. Summary fields (dtm_type, n_tm and sp_length) are joined by annotate(); coordinates for as many as 24 transmembrane segments are written separately by supplementary(). Per-segment sequence columns are excluded because they are recoverable from the reference proteome.

Functions

annotate(frame, *[, key_column, quiet])

frame with the bundled Toxoplasma annotation joined on.

annotate_from_uniprot(frame, source, *[, cache_dir, ...])

frame with UniProt's annotation for source joined onto it.

annotate_with(frame, source, *[, cache_dir, ...])

Annotate frame from whatever source names.

clear_cache(→ None)

Forget the bundled tables. For tests, and for a reinstall mid-session.

columns(→ List[str])

Every column annotate() can add, in order.

gene_number(→ Optional[str])

The bare gene number named by value, or None.

supplementary([genes, path])

The full DeepTMHMM table, as its own supplementary data table.

Module Contents

spacr.annotation.annotate(frame, *, key_column: str | None = None, quiet: bool = False)[source]

frame with the bundled Toxoplasma annotation joined on.

Parameters:
  • frame – any exported table – coefficients, significant hits, the montage sidecar.

  • key_column – the column naming a gene. Found automatically when not given, and a table naming no gene comes back UNCHANGED rather than gaining a block of empty columns.

  • quiet – suppress the console line saying what was joined.

Returns:

a COPY. The caller is usually holding the run’s own results, and annotating them in place would move what every other panel and export sees.

Columns already present in frame are not overwritten – a run that computed its own gene_name keeps it, and the annotation’s version is left out rather than silently replacing it.

spacr.annotation.annotate_from_uniprot(frame, source, *, cache_dir=None, key_column: str | None = None, quiet: bool = False)[source]

frame with UniProt’s annotation for source joined onto it.

The non-bundled half of annotate_with(). Joins on the gene NAME – UniProt’s own, including its synonyms – and on the accession, because a screen library names its targets one way or the other and neither is wrong.

Never raises and never multiplies rows: the annotation is collapsed to one row per key before the merge, and the merge is many_to_one, for the reason this module’s header gives.

Returns:

(frame, note). The frame is unchanged when nothing could be joined, and the note says why.

spacr.annotation.annotate_with(frame, source, *, cache_dir=None, key_column: str | None = None, quiet: bool = False)[source]

Annotate frame from whatever source names.

The one call the pipeline makes. source is the annotation_source setting:

  • empty or any spelling of Toxoplasma – the BUNDLED tables, offline, exactly as before. This is the default and it does not touch spacr.uniprot at all;

  • an organism name or taxon id – that organism from UniProt;

  • an accession – that entry.

Returns:

(frame, note); the note is empty on a clean join.

spacr.annotation.clear_cache() → None[source]

Forget the bundled tables. For tests, and for a reinstall mid-session.

spacr.annotation.columns() → List[str][source]

Every column annotate() can add, in order.

Empty when nothing is bundled. Used to state up front what an export will carry, and by the tests, so a source added here cannot be forgotten.

spacr.annotation.gene_number(value) → str | None[source]

The bare gene number named by value, or None.

Parameters:

value – feature, accession, guide, or design-term value to parse.

Accepts every spelling this project has: a design term (gene_fraction:gene[224750], fraction:grna[224750_2]), a gene id in either strain (TGGT1_224750, TGME49_224750), a split gene model (TGME49_201180A), a guide id (224750_2) or the bare number.

Returns:

the number as a string, or None for anything that names no gene – Intercept, rowID[T.r03], an empty cell, NaN.

A SPLIT GENE MODEL COLLAPSES TO ITS PARENT. TGME49_201180A and TGME49_201180B are both 201180, because the screen’s library targets the locus and has no way to tell the two models apart. It is why every source here is deduplicated before it is joined.

spacr.annotation.supplementary(genes=None, path=None)[source]

The full DeepTMHMM table, as its own supplementary data table.

This table contains the complete bundled DeepTMHMM results for all Toxoplasma proteins.

Parameters:
  • genes – restrict to these genes – any spelling gene_number() accepts. None writes all 8,140 proteins.

  • path – write here as CSV as well as returning the frame.

Returns:

the table, or None when DeepTMHMM is not bundled.

SEPARATE FROM annotate() ON PURPOSE. The per-segment coordinates are 72 columns, and a coefficient export carrying them is a table nobody opens twice. What belongs beside a coefficient is “does this protein have a signal peptide, and how many transmembrane helices”; where each helix starts is a different question and gets a different file.

ONLY THE SEGMENTS THAT EXIST. A screen of soluble proteins gets a table ending at n_tm, not 72 columns of nothing – the same rule the rest of this module follows.