spacr.annotation¶
Join bundled Toxoplasma gene annotations to exported results.
The joined fields include gene names, predicted signal peptides and transmembrane domains, published phenotype scores, stage-specific expression and hyperLOPIT localisation.
All supported identifier forms are normalized to a bare gene number:
TGGT1_224750, TGME49_224750, gene_fraction:gene[224750] and the
guide identifier 224750_2 resolve to gene 224750 through
gene_number(). Each source is reduced to one row per gene and merged
with many_to_one validation, preventing annotation duplicates from
multiplying result rows. A field unavailable in a source is omitted rather
than emitted as an all-missing column.
Bundled annotation sources:
toxoplasma_metadata.csv 2.97 MB gene name, product, expression deeptmhmm.csv 1.11 MB signal peptide and transmembrane phenotype.csv 0.48 MB the published CRISPR fitness screens lopit.csv 0.20 MB hyperLOPIT/TAGM compartment uniprot.csv 0.13 MB accession
deeptmhmm.csv contains a DeepTMHMM analysis of 8,140 proteins. Summary
fields (dtm_type, n_tm and sp_length) are joined by
annotate(); coordinates for as many as 24 transmembrane segments are
written separately by supplementary(). Per-segment sequence columns are
excluded because they are recoverable from the reference proteome.
Functions¶
|
|
|
|
|
Annotate |
|
Forget the bundled tables. For tests, and for a reinstall mid-session. |
|
Every column |
|
The bare gene number named by |
|
The full DeepTMHMM table, as its own supplementary data table. |
Module Contents¶
- spacr.annotation.annotate(frame, *, key_column: str | None = None, quiet: bool = False)[source]¶
framewith the bundled Toxoplasma annotation joined on.- Parameters:
frame – any exported table – coefficients, significant hits, the montage sidecar.
key_column – the column naming a gene. Found automatically when not given, and a table naming no gene comes back UNCHANGED rather than gaining a block of empty columns.
quiet – suppress the console line saying what was joined.
- Returns:
a COPY. The caller is usually holding the run’s own results, and annotating them in place would move what every other panel and export sees.
Columns already present in
frameare not overwritten – a run that computed its owngene_namekeeps it, and the annotation’s version is left out rather than silently replacing it.
- spacr.annotation.annotate_from_uniprot(frame, source, *, cache_dir=None, key_column: str | None = None, quiet: bool = False)[source]¶
framewith UniProt’s annotation forsourcejoined onto it.The non-bundled half of
annotate_with(). Joins on the gene NAME – UniProt’s own, including its synonyms – and on the accession, because a screen library names its targets one way or the other and neither is wrong.Never raises and never multiplies rows: the annotation is collapsed to one row per key before the merge, and the merge is
many_to_one, for the reason this module’s header gives.- Returns:
(frame, note). The frame is unchanged when nothing could be joined, and the note says why.
- spacr.annotation.annotate_with(frame, source, *, cache_dir=None, key_column: str | None = None, quiet: bool = False)[source]¶
Annotate
framefrom whateversourcenames.The one call the pipeline makes.
sourceis theannotation_sourcesetting:empty or any spelling of Toxoplasma – the BUNDLED tables, offline, exactly as before. This is the default and it does not touch
spacr.uniprotat all;an organism name or taxon id – that organism from UniProt;
an accession – that entry.
- Returns:
(frame, note); the note is empty on a clean join.
- spacr.annotation.clear_cache() None[source]¶
Forget the bundled tables. For tests, and for a reinstall mid-session.
- spacr.annotation.columns() List[str][source]¶
Every column
annotate()can add, in order.Empty when nothing is bundled. Used to state up front what an export will carry, and by the tests, so a source added here cannot be forgotten.
- spacr.annotation.gene_number(value) str | None[source]¶
The bare gene number named by
value, orNone.- Parameters:
value – feature, accession, guide, or design-term value to parse.
Accepts every spelling this project has: a design term (
gene_fraction:gene[224750],fraction:grna[224750_2]), a gene id in either strain (TGGT1_224750,TGME49_224750), a split gene model (TGME49_201180A), a guide id (224750_2) or the bare number.- Returns:
the number as a string, or
Nonefor anything that names no gene –Intercept,rowID[T.r03], an empty cell, NaN.
A SPLIT GENE MODEL COLLAPSES TO ITS PARENT.
TGME49_201180AandTGME49_201180Bare both201180, because the screen’s library targets the locus and has no way to tell the two models apart. It is why every source here is deduplicated before it is joined.
- spacr.annotation.supplementary(genes=None, path=None)[source]¶
The full DeepTMHMM table, as its own supplementary data table.
This table contains the complete bundled DeepTMHMM results for all Toxoplasma proteins.
- Parameters:
genes – restrict to these genes – any spelling
gene_number()accepts.Nonewrites all 8,140 proteins.path – write here as CSV as well as returning the frame.
- Returns:
the table, or
Nonewhen DeepTMHMM is not bundled.
SEPARATE FROM
annotate()ON PURPOSE. The per-segment coordinates are 72 columns, and a coefficient export carrying them is a table nobody opens twice. What belongs beside a coefficient is “does this protein have a signal peptide, and how many transmembrane helices”; where each helix starts is a different question and gets a different file.ONLY THE SEGMENTS THAT EXIST. A screen of soluble proteins gets a table ending at
n_tm, not 72 columns of nothing – the same rule the rest of this module follows.