spacr.hits

Build ranked, annotated hit lists from regression results.

A regression run writes full, level-specific, and selected coefficient tables to a uniquely named results/<kind>[_n] directory. This module combines the available tables into one sortable row per gene, adds effect estimates and uncertainty when the backend provides them, calculates guide-direction agreement, and joins optional gene metadata.

Metadata joins are validated as many-to-one. Repeated transcript records are collapsed before the join so they cannot duplicate a gene in the hit list. Backends without frequentist p-values are ranked by bootstrap selection frequency; other supported backends are ranked by q-value.

The module is headless and reads files already written by a regression run.

Public API:

from spacr.hits import build_hit_list, load_results

hits = build_hit_list("/data/plate7/results/ols_2")
strong = hits.filter(max_q=0.05, min_agreement=0.66, min_guides=2)
strong.write_csv("/tmp/hits.csv")
print(strong.to_markdown(limit=20))

Classes

Hit

One gene, everything known about it, and why to believe it.

HitList

A ranked list of hits plus everything needed to interpret it.

Functions

benjamini_hochberg(→ numpy.ndarray)

Benjamini-Hochberg FDR q-values for a vector of p-values.

build_hit_list(, metadata_key, toxoplasma, ...)

Build the ranked, annotated hit list of one regression run.

coefficient_levels(→ pandas.Series)

Identify the model level associated with each coefficient row.

family_labels(→ numpy.ndarray)

Label the multiple-testing family of each coefficient term.

gene_of(→ Optional[str])

Extract the gene identifier from a model term.

grna_agreement(→ Dict[str, Tuple[int, int, List[str]]])

How many of each gene's guides push the same way as the gene.

guide_of(→ Optional[str])

Extract the guide identifier from a model term.

join_metadata(, *, key, List[str]])

Attach every metadata file to frame on its gene column.

load_gene_metadata(→ Tuple[pandas.DataFrame, List[str]])

Read an annotation CSV with at most one row per gene.

load_results(→ Dict[str, pandas.DataFrame])

Read the coefficient tables a regression results folder holds.

tested_family(→ numpy.ndarray)

Identify coefficient terms included in multiple testing.

Module Contents

class spacr.hits.Hit[source]

One gene, everything known about it, and why to believe it.

Parameters:
  • gene – the gene id parsed from the model term.

  • feature – the model term itself, so a row can be traced back.

  • effect – the fitted coefficient — the effect size.

  • std_err – its standard error, when the backend reported one.

  • ci_low – lower bound of the 95% interval; nan without an error.

  • ci_high – upper bound of the same.

  • p_value – the reported p-value; nan for a backend that has none.

  • q_value – Benjamini-Hochberg FDR over the genes tested here.

  • selection_frequency – bootstrap selection frequency, for the penalised backends that rank by it instead of by a p-value.

  • n_guides – how many of this gene’s guides the per-gRNA table holds.

  • n_agree – how many of them push the same way as the gene effect.

  • agreement – n_agree / n_guides; nan with no guide rows.

  • agreeing_guides – which guides those are.

  • n_obs – observations behind the gene term (well x guide rows).

  • condition – nc / pc / control / other as the regression labelled it.

  • direction – up or down, from the sign of the effect.

  • rank – 1-based position in the list as ranked.

  • flags – reasons to look twice; see FLAG_MEANING.

  • annotation – the metadata columns that joined onto this gene.

to_dict() → Dict[str, Any][source]

A flat, JSON-serializable row: the fields, then the annotation.

property name: str[source]

an annotated one, else the id.

Type:

The most human name available

class spacr.hits.HitList[source]

A ranked list of hits plus everything needed to interpret it.

Parameters:
  • hits – the rows, already ranked.

  • source – the results folder or a description of where it came from.

  • regression_type – the backend, when it could be determined.

  • ranking – how the list is ordered — "q-value" or "selection-frequency".

  • alpha – the FDR the significant count is taken at.

  • n_terms – how many model terms the source table held.

  • n_genes – how many distinct genes were tested.

  • filters – the filters applied to reach this list, as data.

  • notes – metadata collapses, missing files, backend caveats.

__getitem__(index)[source]

Index or slice the rows; a slice returns a HitList.

__iter__()[source]

Iterate the rows in rank order.

__len__() → int[source]

How many rows the list holds.

columns() → List[str][source]

Column order for the table forms, annotation columns last.

filter(*, max_q: float | None = None, max_p: float | None = None, min_effect: float | None = None, min_agreement: float | None = None, min_guides: int | None = None, min_selection: float | None = None, direction: str | None = None, conditions: Iterable[str] | None = None, exclude_controls: bool = False, genes: Iterable[str] | None = None, query: str = '', predicate: Callable[[Hit], bool] | None = None) → HitList[source]

Return a narrowed list. Every argument is optional and ANDed.

A row whose value for a criterion is missing FAILS that criterion rather than passing it: a gene with no q-value has not been shown to clear an FDR, and letting missing data through a filter is how an untested term ends up in a hit list.

Parameters:
  • max_q – keep rows with q_value <= max_q.

  • max_p – keep rows with p_value <= max_p.

  • min_effect – keep rows with abs(effect) >= min_effect.

  • min_agreement – keep rows whose guide agreement is at least this.

  • min_guides – keep rows with at least this many guides.

  • min_selection – keep rows with at least this bootstrap selection frequency.

  • direction – "up" or "down".

  • conditions – keep only these condition values.

  • exclude_controls – drop nc / pc / control rows.

  • genes – keep only these gene ids.

  • query – case-insensitive substring, matched against the gene id, the name and every annotation value.

  • predicate – an arbitrary extra test.

Returns:

a new HitList; the receiver is unchanged.

flag_counts() → Dict[str, int][source]

{flag: how many rows carry it}.

gene(gene: str) → Hit | None[source]

The row for one gene id, or None.

Parameters:

gene – exact gene identifier to look up.

significant(alpha: float | None = None) → HitList[source]

The rows that clear the FDR (or the selection threshold).

summary() → Dict[str, Any][source]

Counts and settings, for a header line or a run digest.

This is the block spacr.methods_export puts in front of the model: every number a results paragraph would quote, computed here rather than by whatever writes the prose.

to_frame() → pandas.DataFrame[source]

The list as a DataFrame, one row per gene, in rank order.

to_html(limit: int = 500) → str[source]

A standalone HTML table — the form handed to a collaborator.

Self-contained: no stylesheet, no script, no network. It opens in a browser on a machine that has never heard of spaCR, which is the whole requirement.

to_markdown(limit: int = 50) → str[source]

A Markdown table of the top limit rows, with its legend.

The form the list travels in: pasted into an email, a lab notebook or an issue. No trailing newline.

top(n: int) → HitList[source]

The first n rows, still ranked.

Parameters:

n – maximum number of ranked rows to retain.

write_csv(path: str | os.PathLike) → str[source]

Write the table as CSV and return the path written.

Parameters:

path – destination CSV path.

Through spacr.tabular.write_table(), so a hit list is written with the same column spellings every spaCR reader expects to find.

write_html(path: str | os.PathLike) → str[source]

Write to_html() to a file and return the path.

Parameters:

path – destination HTML path.

spacr.hits.benjamini_hochberg(p_values: Sequence[Any]) → numpy.ndarray[source]

Benjamini-Hochberg FDR q-values for a vector of p-values.

The step-up procedure, with the monotonicity enforced by the running minimum from the largest p-value down — without it a q-value can come out smaller than one belonging to a smaller p-value, which reads as a hit ranking that disagrees with itself.

NaN p-values (a term the backend could not test) stay NaN and are excluded from m: correcting for a test that was not run inflates every other q-value.

Parameters:

p_values – p-values; anything non-finite is treated as untested.

Returns:

q-values aligned with the input.

spacr.hits.build_hit_list(source: str | os.PathLike | Mapping[str, pandas.DataFrame], *, metadata_files: Sequence[str | os.PathLike] = (), metadata_key: str = 'Gene ID', toxoplasma: bool = False, regression_type: str = '', alpha: float = DEFAULT_ALPHA, include_controls: bool = True) → HitList[source]

Build the ranked, annotated hit list of one regression run.

Parameters:
  • source – a results folder, or a {role: DataFrame} mapping as load_results() returns — the second form is what a caller with the frames already in hand (or a test) uses.

  • metadata_files – annotation CSVs to join, each collapsed to one row per gene first. See load_gene_metadata().

  • metadata_key – the gene identifier column in those files.

  • toxoplasma – also join the bundled Toxoplasma annotation — gene name, signal peptide and transmembrane, hyperLOPIT compartment, the published CRISPR fitness scores, and tachyzoite / tissue-cyst / EES1-5 expression. Applied AFTER metadata_files so a column the user’s own file supplies is never replaced by the bundle’s.

  • regression_type – the backend, if known. Only affects how the list is ranked: the penalised backends have no p-value, so they rank by bootstrap selection frequency and carry no q-value.

  • alpha – the FDR (or the selection frequency) a hit is called at.

  • include_controls – keep the control rows in the list. They belong there by default — a screen whose positive control is not near the top has a problem, and that is visible only if it is listed.

Returns:

a HitList, ranked and with exactly one row per gene.

Raises:
spacr.hits.coefficient_levels(frame: pandas.DataFrame | None) → pandas.Series[source]

Identify the model level associated with each coefficient row.

Parameters:

frame (pandas.DataFrame or None) – Coefficient table. None returns an empty series. If neither level nor feature is present, every row is assigned ''.

Returns:

pandas.Series – 'grna', 'gene', or '' for each row, indexed like frame. An empty value denotes a nuisance term or a term whose level cannot be inferred.

Notes

The explicit level column is preferred. Feature-name inference supports tables written before that column was introduced. This distinction matters for level='both' fits because both models contain an Intercept whose name alone does not identify its model.

Examples

>>> import pandas as pd
>>> coefficient_levels(pd.DataFrame(
...     {"feature": ["Intercept", "fraction:grna[233460_1]"],
...      "level": ["grna", "grna"]})).tolist()
['grna', 'grna']
>>> coefficient_levels(pd.DataFrame(
...     {"feature": ["Intercept", "gene_fraction:gene[233460]"]})).tolist()
['', 'gene']
spacr.hits.family_labels(features: Iterable[Any]) → numpy.ndarray[source]

Label the multiple-testing family of each coefficient term.

Parameters:

features (iterable of Any) – Design-matrix term names.

Returns:

numpy.ndarray of object – Labels aligned with features. Values are 'grna' for guide terms, 'gene' for gene terms, and '' for nuisance terms.

Notes

Runs with level='both' contain separate guide and gene testing families. Corrections should be applied within each non-empty label rather than across the pooled coefficient table. Explicit :gene[...] and :grna[...] labels take precedence over identifier shape.

Examples

>>> family_labels(["Intercept", "fraction:grna[233460_1]",
...                "gene_fraction:gene[233460]"]).tolist()
['', 'grna', 'gene']
spacr.hits.gene_of(feature: Any) → str | None[source]

Extract the gene identifier from a model term.

Parameters:

feature (Any) – Design-matrix term containing an identifier in square brackets.

Returns:

str or None – Gene identifier, or None when the term contains no bracketed identifier. Guide suffixes are removed, while the strain prefix and numeric portion of a VEuPathDB accession are retained.

Examples

gene_fraction:gene[233460] and fraction:grna[233460_1] both map to 233460. fraction:grna[TGGT1_231640_3] maps to TGGT1_231640.

spacr.hits.grna_agreement(gene_effects: Mapping[str, float], grna_frame: pandas.DataFrame | None) → Dict[str, Tuple[int, int, List[str]]][source]

How many of each gene’s guides push the same way as the gene.

Parameters:
  • gene_effects – {gene id: gene-level coefficient}.

  • grna_frame – the per-gRNA coefficient table (feature and coefficient columns). None or empty means “no guide-level evidence”, which is reported as such rather than as agreement.

Returns:

{gene id: (n_agree, n_guides, [guide ids that agree])}. A guide whose coefficient is exactly zero counts as a guide but agrees with nothing — a penalised backend shrinks non-contributing guides to zero, and counting those as agreement would turn a lasso’s sparsity into corroboration.

spacr.hits.guide_of(feature: Any) → str | None[source]

Extract the guide identifier from a model term.

Parameters:

feature (Any) – Design-matrix term containing an identifier in square brackets.

Returns:

str or None – Complete guide identifier, or None for gene and nuisance terms. Explicit :gene[...] and :grna[...] labels take precedence over identifier shape, so an underscore within a VEuPathDB gene accession is not mistaken for a guide suffix.

spacr.hits.join_metadata(frame: pandas.DataFrame, metadata_files: Sequence[str | os.PathLike] = (), *, key: str = 'Gene ID') → Tuple[pandas.DataFrame, List[str]][source]

Attach every metadata file to frame on its gene column.

validate="many_to_one" on every join: many result rows may name one gene (one per guide), but each gene gets one annotation row. pandas raises rather than fanning out if a file ever breaks that, which is the guard that keeps the collapse in load_gene_metadata() honest.

Parameters:
  • frame – a table with a gene column.

  • metadata_files – annotation CSVs, applied in order. A later file’s columns are suffixed rather than overwriting an earlier one’s.

  • key – the gene identifier column in the metadata files.

Returns:

(joined, notes).

spacr.hits.load_gene_metadata(path: str | os.PathLike, *, key: str = 'Gene ID') → Tuple[pandas.DataFrame, List[str]][source]

Read an annotation CSV with at most one row per gene.

Parameters:
  • path (path-like) – Annotation CSV to read.

  • key (str, default='Gene ID') – Column containing accessions such as TGME49_233460. The component after the first underscore is stored in a new gene column.

Returns:

  • frame (pandas.DataFrame) – Annotation rows with a gene column and no duplicate gene values. When several transcript rows map to one gene, the first row is kept.

  • notes (list of str) – User-facing descriptions of unparsable rows and duplicate-gene rows removed during normalization.

Raises:

Notes

Duplicate transcript annotations are collapsed before metadata is joined to results, preventing one gene from becoming several hit-list rows. Annotations from discarded duplicate rows are not merged into the row that is retained.

spacr.hits.load_results(folder: str | os.PathLike) → Dict[str, pandas.DataFrame][source]

Read the coefficient tables a regression results folder holds.

Parameters:

folder – A results/<kind>[_n] directory written by spacr.ml.perform_regression().

Returns:

{role: DataFrame} for whichever of RESULT_FILES exist. A folder with none of them yields an empty dict rather than an exception — “that is not a results folder” is something the caller must be able to say to a user.

Raises:

FileNotFoundError – when folder does not exist at all.

spacr.hits.tested_family(features: Iterable[Any]) → numpy.ndarray[source]

Identify coefficient terms included in multiple testing.

Parameters:

features (iterable of Any) – Design-matrix term names.

Returns:

numpy.ndarray of bool – Boolean mask aligned with features. Guide and gene terms are True; intercept and layout nuisance terms are False.

Notes

The mask follows the family corrected by spacr.ml.perform_regression(). Nuisance terms are fitted as covariates but are excluded from hit plots and multiple-testing correction.

Examples

>>> tested_family(["Intercept", "fraction:grna[233460_1]"]).tolist()
[False, True]

Nested helpers

HitList.filter._ok(hit: Hit) → bool

Return whether hit satisfies every supplied filter.

spacr/hits.py:745

_rank.key(hit: Hit)

Rank finite selection and effect magnitude high, then gene.

Rank finite q low, effect magnitude high, then gene.

spacr/hits.py:1225 spacr/hits.py:1232