spacr.sra¶
Fetch sequencing reads for the published screen straight from ENA/NCBI.
Map Barcodes can fetch the paper’s own sequencing data directly. Each download accepts a read limit per file, so a small representative subset does not require transferring the complete archive.
The reads are NCBI BioProject DEFAULT_BIOPROJECT, runs
SRR33531217-SRR33531220 – the four plates, paired, named hilib_p1 through
hilib_p4 in the submission.
WHY ENA AND NOT NCBI’S OWN DOWNLOAD. The three routes are not equivalent
here. fasterq-dump needs the SRA toolkit installed, which spaCR does not
ship and cannot assume. NCBI’s own FASTQ endpoint does not support fetching a
prefix. ENA mirrors every SRA submission as plain gzipped FASTQ over HTTPS,
which is what makes the read limit meaningful rather than cosmetic: the file
is read as a STREAM and the connection is dropped as soon as enough reads have
arrived.
That distinction is the whole feature. The four runs are 2.2-3.2 GB per mate, about 19 GB in total, and 60-74 million reads each. Measured against the live archive on 2026-09-01: 3,662 reads arrived in 0.13 MB and 1.7 seconds from a 2,833 MB file. Asking for a hundred thousand reads costs a few megabytes, not a few gigabytes, so a laptop on hotel wifi can have a real subset of the real screen in under a minute.
Classes¶
One downloadable FASTQ: a run, and which mate of the pair. |
Functions¶
|
Roughly what |
|
Download |
|
Every FASTQ in |
|
What downloading all of |
Module Contents¶
- class spacr.sra.RunFile[source]¶
One downloadable FASTQ: a run, and which mate of the pair.
- Parameters:
run – archive run accession shared by its mate files.
library – experiment library name used to identify the source plate.
url – HTTPS location of the compressed FASTQ stream.
mate – one-based mate number within a paired-end run.
size_bytes – archive-reported compressed size, or zero when unknown.
read_count – archive-reported reads in the run, or zero when unknown.
- spacr.sra.estimated_bytes(files: Sequence[RunFile], max_reads: int | None) int[source]¶
Roughly what
max_readsfrom each offileswill transfer.Scaled from each run’s own read count rather than from a fixed rate, so a run with longer reads is not under-quoted. Returns the full size when no limit is set, and when a run does not report its read count – guessing low there would understate a multi-gigabyte download.
- Parameters:
files – the run files to download.
max_reads – reads to take from each file, or None for whole files; each file’s share is
max_reads / read_count, capped at 1.
- spacr.sra.fetch_reads(run_file: RunFile, destination, *, max_reads: int | None = None, timeout: float = 60.0, opener=None, progress: Callable[[int, int], None] | None = None, should_stop: Callable[[], bool] | None = None) pathlib.Path[source]¶
Download
run_fileintodestination, stopping aftermax_reads.The stream is decompressed as it arrives and the connection is dropped the moment enough reads are in hand, so a limited request transfers only what it needs.
max_readsofNonefetches the whole file.Written back out as
.fastq.gzbecause that is what thesrcsetting documents for sequencing (“the folder of .fastq.gz reads”), so a subset and a full download are the same shape to everything downstream.- Parameters:
run_file – the archive file to stream; its
urlis read and itsfilenamenames the output.destination – folder the
.fastq.gzfile is written into; created if missing.progress – called with
(reads_so_far, compressed_bytes_so_far).should_stop – polled between chunks; a truthy answer abandons the download and removes the partial file.
- Returns:
the path written.
- Raises:
ValueError – for a non-positive
max_reads.
- spacr.sra.runs_for(accession: str = DEFAULT_BIOPROJECT, *, timeout: float = 30.0, opener=None) tuple[RunFile, ...][source]¶
Every FASTQ in
accession, newest ENA metadata.- Parameters:
accession – a BioProject (
PRJNA...) or a single run (SRR...).opener – replaces the network call; receives the URL and returns a file-like object of TSV bytes.
- Returns:
one
RunFileper mate per run, ordered by run then mate.
Nested helpers¶
- runs_for.cell(name: str, row_cells=cells) str¶
Read one named value from the bound portal row.
- Parameters:
name – ENA header name to look up in the captured index.
row_cells – cells bound when this row’s helper is created, preventing later loop iterations from changing the source row.
- Returns:
the indexed cell, or an empty string when the column is absent or the row is too short.
spacr/sra.py:127