spacr.hits¶
Build ranked, annotated hit lists from regression results.
A regression run writes full, level-specific, and selected coefficient tables
to a uniquely named results/<kind>[_n] directory. This module combines the
available tables into one sortable row per gene, adds effect estimates and
uncertainty when the backend provides them, calculates guide-direction
agreement, and joins optional gene metadata.
Metadata joins are validated as many-to-one. Repeated transcript records are collapsed before the join so they cannot duplicate a gene in the hit list. Backends without frequentist p-values are ranked by bootstrap selection frequency; other supported backends are ranked by q-value.
The module is headless and reads files already written by a regression run.
Public API:
from spacr.hits import build_hit_list, load_results
hits = build_hit_list("/data/plate7/results/ols_2")
strong = hits.filter(max_q=0.05, min_agreement=0.66, min_guides=2)
strong.write_csv("/tmp/hits.csv")
print(strong.to_markdown(limit=20))
Classes¶
Functions¶
|
Benjamini-Hochberg FDR q-values for a vector of p-values. |
|
Build the ranked, annotated hit list of one regression run. |
|
Identify the model level associated with each coefficient row. |
|
Label the multiple-testing family of each coefficient term. |
|
Extract the gene identifier from a model term. |
|
How many of each gene's guides push the same way as the gene. |
|
Extract the guide identifier from a model term. |
|
Attach every metadata file to |
|
Read an annotation CSV with at most one row per gene. |
|
Read the coefficient tables a regression results folder holds. |
|
Identify coefficient terms included in multiple testing. |
Module Contents¶
- class spacr.hits.Hit[source]¶
One gene, everything known about it, and why to believe it.
- Parameters:
gene – the gene id parsed from the model term.
feature – the model term itself, so a row can be traced back.
effect – the fitted coefficient — the effect size.
std_err – its standard error, when the backend reported one.
ci_low – lower bound of the 95% interval;
nanwithout an error.ci_high – upper bound of the same.
p_value – the reported p-value;
nanfor a backend that has none.q_value – Benjamini-Hochberg FDR over the genes tested here.
selection_frequency – bootstrap selection frequency, for the penalised backends that rank by it instead of by a p-value.
n_guides – how many of this gene’s guides the per-gRNA table holds.
n_agree – how many of them push the same way as the gene effect.
agreement –
n_agree / n_guides;nanwith no guide rows.agreeing_guides – which guides those are.
n_obs – observations behind the gene term (well x guide rows).
condition –
nc/pc/control/otheras the regression labelled it.direction –
upordown, from the sign of the effect.rank – 1-based position in the list as ranked.
flags – reasons to look twice; see
FLAG_MEANING.annotation – the metadata columns that joined onto this gene.
- class spacr.hits.HitList[source]¶
A ranked list of hits plus everything needed to interpret it.
- Parameters:
hits – the rows, already ranked.
source – the results folder or a description of where it came from.
regression_type – the backend, when it could be determined.
ranking – how the list is ordered —
"q-value"or"selection-frequency".alpha – the FDR the
significantcount is taken at.n_terms – how many model terms the source table held.
n_genes – how many distinct genes were tested.
filters – the filters applied to reach this list, as data.
notes – metadata collapses, missing files, backend caveats.
- filter(*, max_q: float | None = None, max_p: float | None = None, min_effect: float | None = None, min_agreement: float | None = None, min_guides: int | None = None, min_selection: float | None = None, direction: str | None = None, conditions: Iterable[str] | None = None, exclude_controls: bool = False, genes: Iterable[str] | None = None, query: str = '', predicate: Callable[[Hit], bool] | None = None) HitList[source]¶
Return a narrowed list. Every argument is optional and ANDed.
A row whose value for a criterion is missing FAILS that criterion rather than passing it: a gene with no q-value has not been shown to clear an FDR, and letting missing data through a filter is how an untested term ends up in a hit list.
- Parameters:
max_q – keep rows with
q_value <= max_q.max_p – keep rows with
p_value <= max_p.min_effect – keep rows with
abs(effect) >= min_effect.min_agreement – keep rows whose guide agreement is at least this.
min_guides – keep rows with at least this many guides.
min_selection – keep rows with at least this bootstrap selection frequency.
direction –
"up"or"down".conditions – keep only these
conditionvalues.exclude_controls – drop
nc/pc/controlrows.genes – keep only these gene ids.
query – case-insensitive substring, matched against the gene id, the name and every annotation value.
predicate – an arbitrary extra test.
- Returns:
a new
HitList; the receiver is unchanged.
- gene(gene: str) Hit | None[source]¶
The row for one gene id, or
None.- Parameters:
gene – exact gene identifier to look up.
- significant(alpha: float | None = None) HitList[source]¶
The rows that clear the FDR (or the selection threshold).
- summary() Dict[str, Any][source]¶
Counts and settings, for a header line or a run digest.
This is the block
spacr.methods_exportputs in front of the model: every number a results paragraph would quote, computed here rather than by whatever writes the prose.
- to_frame() pandas.DataFrame[source]¶
The list as a DataFrame, one row per gene, in rank order.
- to_html(limit: int = 500) str[source]¶
A standalone HTML table — the form handed to a collaborator.
Self-contained: no stylesheet, no script, no network. It opens in a browser on a machine that has never heard of spaCR, which is the whole requirement.
- to_markdown(limit: int = 50) str[source]¶
A Markdown table of the top
limitrows, with its legend.The form the list travels in: pasted into an email, a lab notebook or an issue. No trailing newline.
- top(n: int) HitList[source]¶
The first
nrows, still ranked.- Parameters:
n – maximum number of ranked rows to retain.
- write_csv(path: str | os.PathLike) str[source]¶
Write the table as CSV and return the path written.
- Parameters:
path – destination CSV path.
Through
spacr.tabular.write_table(), so a hit list is written with the same column spellings every spaCR reader expects to find.
- spacr.hits.benjamini_hochberg(p_values: Sequence[Any]) numpy.ndarray[source]¶
Benjamini-Hochberg FDR q-values for a vector of p-values.
The step-up procedure, with the monotonicity enforced by the running minimum from the largest p-value down — without it a q-value can come out smaller than one belonging to a smaller p-value, which reads as a hit ranking that disagrees with itself.
NaNp-values (a term the backend could not test) stayNaNand are excluded fromm: correcting for a test that was not run inflates every other q-value.- Parameters:
p_values – p-values; anything non-finite is treated as untested.
- Returns:
q-values aligned with the input.
- spacr.hits.build_hit_list(source: str | os.PathLike | Mapping[str, pandas.DataFrame], *, metadata_files: Sequence[str | os.PathLike] = (), metadata_key: str = 'Gene ID', toxoplasma: bool = False, regression_type: str = '', alpha: float = DEFAULT_ALPHA, include_controls: bool = True) HitList[source]¶
Build the ranked, annotated hit list of one regression run.
- Parameters:
source – a results folder, or a
{role: DataFrame}mapping asload_results()returns — the second form is what a caller with the frames already in hand (or a test) uses.metadata_files – annotation CSVs to join, each collapsed to one row per gene first. See
load_gene_metadata().metadata_key – the gene identifier column in those files.
toxoplasma – also join the bundled Toxoplasma annotation — gene name, signal peptide and transmembrane, hyperLOPIT compartment, the published CRISPR fitness scores, and tachyzoite / tissue-cyst / EES1-5 expression. Applied AFTER
metadata_filesso a column the user’s own file supplies is never replaced by the bundle’s.regression_type – the backend, if known. Only affects how the list is ranked: the penalised backends have no p-value, so they rank by bootstrap selection frequency and carry no q-value.
alpha – the FDR (or the selection frequency) a hit is called at.
include_controls – keep the control rows in the list. They belong there by default — a screen whose positive control is not near the top has a problem, and that is visible only if it is listed.
- Returns:
a
HitList, ranked and with exactly one row per gene.- Raises:
FileNotFoundError – when
sourceis a path that is not a folder.ValueError – when no gene-level coefficients could be found.
- spacr.hits.coefficient_levels(frame: pandas.DataFrame | None) pandas.Series[source]¶
Identify the model level associated with each coefficient row.
- Parameters:
frame (pandas.DataFrame or None) – Coefficient table.
Nonereturns an empty series. If neitherlevelnorfeatureis present, every row is assigned''.- Returns:
pandas.Series –
'grna','gene', or''for each row, indexed likeframe. An empty value denotes a nuisance term or a term whose level cannot be inferred.
Notes
The explicit
levelcolumn is preferred. Feature-name inference supports tables written before that column was introduced. This distinction matters forlevel='both'fits because both models contain anInterceptwhose name alone does not identify its model.Examples
>>> import pandas as pd >>> coefficient_levels(pd.DataFrame( ... {"feature": ["Intercept", "fraction:grna[233460_1]"], ... "level": ["grna", "grna"]})).tolist() ['grna', 'grna'] >>> coefficient_levels(pd.DataFrame( ... {"feature": ["Intercept", "gene_fraction:gene[233460]"]})).tolist() ['', 'gene']
- spacr.hits.family_labels(features: Iterable[Any]) numpy.ndarray[source]¶
Label the multiple-testing family of each coefficient term.
- Parameters:
features (iterable of Any) – Design-matrix term names.
- Returns:
numpy.ndarray of object – Labels aligned with
features. Values are'grna'for guide terms,'gene'for gene terms, and''for nuisance terms.
Notes
Runs with
level='both'contain separate guide and gene testing families. Corrections should be applied within each non-empty label rather than across the pooled coefficient table. Explicit:gene[...]and:grna[...]labels take precedence over identifier shape.Examples
>>> family_labels(["Intercept", "fraction:grna[233460_1]", ... "gene_fraction:gene[233460]"]).tolist() ['', 'grna', 'gene']
- spacr.hits.gene_of(feature: Any) str | None[source]¶
Extract the gene identifier from a model term.
- Parameters:
feature (Any) – Design-matrix term containing an identifier in square brackets.
- Returns:
str or None – Gene identifier, or
Nonewhen the term contains no bracketed identifier. Guide suffixes are removed, while the strain prefix and numeric portion of a VEuPathDB accession are retained.
Examples
gene_fraction:gene[233460]andfraction:grna[233460_1]both map to233460.fraction:grna[TGGT1_231640_3]maps toTGGT1_231640.
- spacr.hits.grna_agreement(gene_effects: Mapping[str, float], grna_frame: pandas.DataFrame | None) Dict[str, Tuple[int, int, List[str]]][source]¶
How many of each gene’s guides push the same way as the gene.
- Parameters:
gene_effects –
{gene id: gene-level coefficient}.grna_frame – the per-gRNA coefficient table (
featureandcoefficientcolumns).Noneor empty means “no guide-level evidence”, which is reported as such rather than as agreement.
- Returns:
{gene id: (n_agree, n_guides, [guide ids that agree])}. A guide whose coefficient is exactly zero counts as a guide but agrees with nothing — a penalised backend shrinks non-contributing guides to zero, and counting those as agreement would turn a lasso’s sparsity into corroboration.
- spacr.hits.guide_of(feature: Any) str | None[source]¶
Extract the guide identifier from a model term.
- Parameters:
feature (Any) – Design-matrix term containing an identifier in square brackets.
- Returns:
str or None – Complete guide identifier, or
Nonefor gene and nuisance terms. Explicit:gene[...]and:grna[...]labels take precedence over identifier shape, so an underscore within a VEuPathDB gene accession is not mistaken for a guide suffix.
- spacr.hits.join_metadata(frame: pandas.DataFrame, metadata_files: Sequence[str | os.PathLike] = (), *, key: str = 'Gene ID') Tuple[pandas.DataFrame, List[str]][source]¶
Attach every metadata file to
frameon itsgenecolumn.validate="many_to_one"on every join: many result rows may name one gene (one per guide), but each gene gets one annotation row. pandas raises rather than fanning out if a file ever breaks that, which is the guard that keeps the collapse inload_gene_metadata()honest.- Parameters:
frame – a table with a
genecolumn.metadata_files – annotation CSVs, applied in order. A later file’s columns are suffixed rather than overwriting an earlier one’s.
key – the gene identifier column in the metadata files.
- Returns:
(joined, notes).
- spacr.hits.load_gene_metadata(path: str | os.PathLike, *, key: str = 'Gene ID') Tuple[pandas.DataFrame, List[str]][source]¶
Read an annotation CSV with at most one row per gene.
- Parameters:
path (path-like) – Annotation CSV to read.
key (str, default='Gene ID') – Column containing accessions such as
TGME49_233460. The component after the first underscore is stored in a newgenecolumn.
- Returns:
frame (pandas.DataFrame) – Annotation rows with a
genecolumn and no duplicate gene values. When several transcript rows map to one gene, the first row is kept.notes (list of str) – User-facing descriptions of unparsable rows and duplicate-gene rows removed during normalization.
- Raises:
FileNotFoundError – If
pathis not an existing file.KeyError – If
keyis absent from the CSV.
Notes
Duplicate transcript annotations are collapsed before metadata is joined to results, preventing one gene from becoming several hit-list rows. Annotations from discarded duplicate rows are not merged into the row that is retained.
- spacr.hits.load_results(folder: str | os.PathLike) Dict[str, pandas.DataFrame][source]¶
Read the coefficient tables a regression results folder holds.
- Parameters:
folder – A
results/<kind>[_n]directory written byspacr.ml.perform_regression().- Returns:
{role: DataFrame}for whichever ofRESULT_FILESexist. A folder with none of them yields an empty dict rather than an exception — “that is not a results folder” is something the caller must be able to say to a user.- Raises:
FileNotFoundError – when
folderdoes not exist at all.
- spacr.hits.tested_family(features: Iterable[Any]) numpy.ndarray[source]¶
Identify coefficient terms included in multiple testing.
- Parameters:
features (iterable of Any) – Design-matrix term names.
- Returns:
numpy.ndarray of bool – Boolean mask aligned with
features. Guide and gene terms areTrue; intercept and layout nuisance terms areFalse.
Notes
The mask follows the family corrected by
spacr.ml.perform_regression(). Nuisance terms are fitted as covariates but are excluded from hit plots and multiple-testing correction.Examples
>>> tested_family(["Intercept", "fraction:grna[233460_1]"]).tolist() [False, True]