spacr.sra

Fetch sequencing reads for the published screen straight from ENA/NCBI.

Map Barcodes can fetch the paper’s own sequencing data directly. Each download accepts a read limit per file, so a small representative subset does not require transferring the complete archive.

The reads are NCBI BioProject DEFAULT_BIOPROJECT, runs SRR33531217-SRR33531220 – the four plates, paired, named hilib_p1 through hilib_p4 in the submission.

WHY ENA AND NOT NCBI’S OWN DOWNLOAD. The three routes are not equivalent here. fasterq-dump needs the SRA toolkit installed, which spaCR does not ship and cannot assume. NCBI’s own FASTQ endpoint does not support fetching a prefix. ENA mirrors every SRA submission as plain gzipped FASTQ over HTTPS, which is what makes the read limit meaningful rather than cosmetic: the file is read as a STREAM and the connection is dropped as soon as enough reads have arrived.

That distinction is the whole feature. The four runs are 2.2-3.2 GB per mate, about 19 GB in total, and 60-74 million reads each. Measured against the live archive on 2026-09-01: 3,662 reads arrived in 0.13 MB and 1.7 seconds from a 2,833 MB file. Asking for a hundred thousand reads costs a few megabytes, not a few gigabytes, so a laptop on hotel wifi can have a real subset of the real screen in under a minute.

Classes

RunFile

One downloadable FASTQ: a run, and which mate of the pair.

Functions

estimated_bytes(→ int)

Roughly what max_reads from each of files will transfer.

fetch_reads(→ pathlib.Path)

Download run_file into destination, stopping after max_reads.

runs_for(→ tuple[RunFile, ...])

Every FASTQ in accession, newest ENA metadata.

total_bytes(→ int)

What downloading all of files in full would cost.

Module Contents

class spacr.sra.RunFile[source]

One downloadable FASTQ: a run, and which mate of the pair.

Parameters:
  • run – archive run accession shared by its mate files.

  • library – experiment library name used to identify the source plate.

  • url – HTTPS location of the compressed FASTQ stream.

  • mate – one-based mate number within a paired-end run.

  • size_bytes – archive-reported compressed size, or zero when unknown.

  • read_count – archive-reported reads in the run, or zero when unknown.

label() → str[source]

A one-line description for a picker.

Names the LIBRARY as well as the run, because hilib_p3 says which plate this is and SRR33531218 does not.

property filename: str[source]

the archive’s own name.

Type:

What it is saved as

spacr.sra.estimated_bytes(files: Sequence[RunFile], max_reads: int | None) → int[source]

Roughly what max_reads from each of files will transfer.

Scaled from each run’s own read count rather than from a fixed rate, so a run with longer reads is not under-quoted. Returns the full size when no limit is set, and when a run does not report its read count – guessing low there would understate a multi-gigabyte download.

Parameters:
  • files – the run files to download.

  • max_reads – reads to take from each file, or None for whole files; each file’s share is max_reads / read_count, capped at 1.

spacr.sra.fetch_reads(run_file: RunFile, destination, *, max_reads: int | None = None, timeout: float = 60.0, opener=None, progress: Callable[[int, int], None] | None = None, should_stop: Callable[[], bool] | None = None) → pathlib.Path[source]

Download run_file into destination, stopping after max_reads.

The stream is decompressed as it arrives and the connection is dropped the moment enough reads are in hand, so a limited request transfers only what it needs. max_reads of None fetches the whole file.

Written back out as .fastq.gz because that is what the src setting documents for sequencing (“the folder of .fastq.gz reads”), so a subset and a full download are the same shape to everything downstream.

Parameters:
  • run_file – the archive file to stream; its url is read and its filename names the output.

  • destination – folder the .fastq.gz file is written into; created if missing.

  • progress – called with (reads_so_far, compressed_bytes_so_far).

  • should_stop – polled between chunks; a truthy answer abandons the download and removes the partial file.

Returns:

the path written.

Raises:

ValueError – for a non-positive max_reads.

spacr.sra.runs_for(accession: str = DEFAULT_BIOPROJECT, *, timeout: float = 30.0, opener=None) → tuple[RunFile, ...][source]

Every FASTQ in accession, newest ENA metadata.

Parameters:
  • accession – a BioProject (PRJNA...) or a single run (SRR...).

  • opener – replaces the network call; receives the URL and returns a file-like object of TSV bytes.

Returns:

one RunFile per mate per run, ordered by run then mate.

spacr.sra.total_bytes(files: Iterable[RunFile]) → int[source]

What downloading all of files in full would cost.

Parameters:

files – the run files; their size_bytes are summed, so a file of unknown size adds zero.

Nested helpers

runs_for.cell(name: str, row_cells=cells) → str

Read one named value from the bound portal row.

Parameters:
  • name – ENA header name to look up in the captured index.

  • row_cells – cells bound when this row’s helper is created, preventing later loop iterations from changing the source row.

Returns:

the indexed cell, or an empty string when the column is absent or the row is too short.

spacr/sra.py:127