spacr.sim¶
Workflow inputs and outputs¶
Pooled-screen simulation sweep¶
Use the Python sweep API to explore stated screen-design assumptions. It expands combinations, runs simulations in a process pool and writes synthetic summary statistics; it returns None. Set a small explicit max_workers value and a bounded parameter grid before running. The output does not replace measurements or barcode counts for experimental Regression. Use the findings as planning evidence, with their assumptions recorded. Review the synthetic summaries when choosing planning assumptions; transfer those assumptions manually into Power / Design or Experiment Design. Neither tool imports this simulation database, and neither exports the simulation settings grid.
Use from Python: spacr.sim.run_multiple_simulations(). This API-only workflow has no Home tile or menu entry.
Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.
Inputs
Pooled-screen simulation assumptions — Python settings dictionary: iterable sweep values for screen size, occupancy, classifier accuracy and sequencing assumptions, plus replicates, src, name, variable, plot and max_workers. generate_parameters expands their Cartesian product; begin with a small sweep and an explicit worker bound.
Outputs
Synthetic screen simulation summaries — src/<YYMMDD>/<name>/simulations.db; the sweep appends summary rows to simulations and optionally writes plots. These are synthetic performance estimates, not measured experimental hits. Relevant tables, depending on the route:
simulations.
Pooled-screen simulation, evaluation, and visualisation utilities.
Functions¶
|
Append a DataFrame to |
|
Fit a RandomForest and plot permutation-based feature importances for the sweep columns. |
|
Compute cell-level ROC/PR metrics and confusion matrix at the F1-optimal threshold. |
|
Simulate a noisy binary classifier by drawing per-cell scores from two Beta distributions. |
|
Return precision/recall/F1/PR-AUC arrays for a DataFrame of |
|
Return ROC-curve arrays and AUC for a DataFrame of |
|
Ensure a SQLite database file exists at |
|
Draw a length- |
|
Return floats from |
|
Return |
|
Draw |
|
Return |
|
Expand a sweep-settings dict into one settings dict per (Cartesian) simulation. |
|
Return a 384-well plate map DataFrame spanning |
|
Return a normalized power-law probability vector of length |
|
Fit a RandomForest and render a SHAP summary plot over the standard sweep features. |
|
Aggregate cell-level scores into per-well summary rows. |
|
Return the classification threshold that maximises F1 in a PR result dict. |
|
Return the Gini coefficient of |
Return the Gini coefficient of |
|
Return the Gini coefficient of |
|
|
Return |
|
Render a 2x2 confusion matrix as an annotated Seaborn heatmap. |
|
Render a lower-triangular correlation heatmap of the standard sweep + metric columns. |
|
Train a RandomForestRegressor on sweep variables and plot the resulting importances. |
|
Draw a Seaborn density histogram on |
|
Fit a GradientBoostingRegressor and plot partial dependences for every sweep feature. |
|
Plot a ROC or PR curve with a diagonal random-classifier reference line. |
|
Grid-plot PR-AUC vs |
|
Return |
|
Return the |
|
Score regression hits against ground truth and compute ROC/PR metrics. |
Return |
|
Return |
|
|
Worker that runs one simulation, saves outputs, and appends its runtime. |
|
Simulate one cell-level screening experiment and return per-cell + summary tables. |
|
Fan out the sweep from |
|
Run one end-to-end pooled-screen simulation and return every intermediate table. |
|
Persist one simulation's output tables to a SQLite database under |
|
Save a Matplotlib figure to |
|
Save a SHAP figure to |
|
Simulate sequencing of every well and return per-well gene fraction and metadata. |
|
Add a per-threshold predicted-label column and return the confusion matrix. |
|
Clamp per-run Beta variances so the requested mean/variance is feasible. |
|
Save side-by-side histograms of the six per-run distributions in |
|
Render the full 13-panel diagnostic figure for one simulation output. |
Module Contents¶
- spacr.sim.append_database(src, table, table_name)[source]¶
Append a DataFrame to
<src>/simulations.dbundertable_name.- Parameters:
src – Directory containing (or that should contain)
simulations.db.table – DataFrame written with
if_exists='append'.table_name – Target table name in the SQLite database.
- Returns:
None.
- spacr.sim.calculate_permutation_importance(df, target='prauc', exclude=None, n_repeats=10, clean=True, dst=None)[source]¶
Fit a RandomForest and plot permutation-based feature importances for the sweep columns.
- Parameters:
df – DataFrame with sweep columns and
target.target – Column predicted by the regressor. Default
'prauc'.exclude – Column name or list of columns to remove from the feature set.
n_repeats – Number of permutations per feature. Default
10.clean – When True, drop constant columns before fitting.
- Returns:
The generated Matplotlib figure.
- spacr.sim.cell_level_roc_auc(cell_scores)[source]¶
Compute cell-level ROC/PR metrics and confusion matrix at the F1-optimal threshold.
- Parameters:
cell_scores – DataFrame with columns
is_activeandscore.- Returns:
Tuple
(cell_roc_dict_df, cell_pr_dict_df, cell_scores, cell_cm).
- spacr.sim.classifier(positive_mean, positive_variance, negative_mean, negative_variance, classifier_accuracy, df)[source]¶
Simulate a noisy binary classifier by drawing per-cell scores from two Beta distributions.
Used inside the spacr screen simulator to model an imperfect phenotype classifier: rows with
is_active == 1should score high and rows withis_active == 0should score low, but with probability1 - classifier_accuracya row is drawn from the wrong Beta (simulating classifier error). The result is ascorecolumn on the input DataFrame ready to be fed tocompute_roc_auc()/compute_precision_recall().- Parameters:
positive_mean – Mean of the Beta distribution for active cells; must be in
(0, 1).positive_variance – Variance of the Beta for active cells; must satisfy
0 < var < mean * (1 - mean).negative_mean – Mean of the Beta for inactive cells.
negative_variance – Variance of the Beta for inactive cells.
classifier_accuracy – Probability in
[0, 1]that a row is drawn from the Beta matching its trueis_activelabel.df – DataFrame containing an
is_activecolumn (0 / 1).
- Returns:
The input DataFrame with a new
scorecolumn in[0, 1].- Raises:
ValueError – if either mean is outside
(0, 1)or either variance is outside(0, mean * (1 - mean)).
Example
from spacr.sim import classifier df = classifier( positive_mean=0.8, positive_variance=0.02, negative_mean=0.2, negative_variance=0.02, classifier_accuracy=0.9, df=cells, )
See also
- spacr.sim.compute_precision_recall(cell_scores)[source]¶
Return precision/recall/F1/PR-AUC arrays for a DataFrame of
is_active/scorerows.- Parameters:
cell_scores – DataFrame with columns
is_activeandscore.- Returns:
Dict with keys
threshold,precision,recall,f1_score,pr_auc.
- spacr.sim.compute_roc_auc(cell_scores)[source]¶
Return ROC-curve arrays and AUC for a DataFrame of
is_active/scorerows.- Parameters:
cell_scores – DataFrame with columns
is_activeandscore.- Returns:
Dict with keys
threshold,tpr,fpr,roc_auc.
- spacr.sim.create_database(db_path)[source]¶
Ensure a SQLite database file exists at
db_path.- Parameters:
db_path – Filesystem path for the SQLite database.
- Returns:
None.
- spacr.sim.dist_gen(mean, sd, df)[source]¶
Draw a length-
len(df)Poisson sample with gamma-distributed rates.- Parameters:
mean – Mean of the gamma prior on the Poisson rate.
sd – Standard deviation of the gamma prior.
df – DataFrame whose length sets the sample size.
- Returns:
Tuple
(samples, length)wheresamplesis a NumPy array of Poisson draws andlengthislen(df).
- spacr.sim.generate_floats(start, stop, step)[source]¶
Return floats from
startthroughstopwithstepspacing.- Parameters:
start – first numeric value in the generated sequence.
stop – inclusive endpoint when reached by the requested spacing.
step – numeric spacing and source of output decimal precision.
- spacr.sim.generate_gene_list(number_of_genes, number_of_all_genes)[source]¶
Return
number_of_genesrandomly-drawn gene indices without replacement.- Parameters:
number_of_genes – Number of gene indices to draw.
number_of_all_genes – Size of the pool
[0, number_of_all_genes).
- Returns:
List of drawn gene indices.
- spacr.sim.generate_gene_weights(positive_mean, positive_variance, df)[source]¶
Draw
len(df)gene weights from a Beta distribution matched to the given moments.- Parameters:
positive_mean – Target mean of the Beta distribution in
(0, 1).positive_variance – Target variance (must be feasible for the mean).
df – DataFrame whose length sets the sample size.
- Returns:
NumPy array of Beta-distributed weights.
- spacr.sim.generate_integers(start, stop, step)[source]¶
Return
list(range(start, stop + 1, step))(inclusive upper bound).- Parameters:
start – first integer in the generated sequence.
stop – inclusive upper endpoint for the sequence.
step – integer spacing passed to
range.
- spacr.sim.generate_parameters(settings)[source]¶
Expand a sweep-settings dict into one settings dict per (Cartesian) simulation.
- Parameters:
settings – Config dict where each swept key holds an iterable of values.
- Returns:
Shuffled list of per-run settings dicts, already run through
validate_and_adjust_beta_params().
- spacr.sim.generate_plate_map(nr_plates)[source]¶
Return a 384-well plate map DataFrame spanning
nr_platesplates.- Parameters:
nr_plates – Number of plates to enumerate.
- Returns:
DataFrame with
plate_row_column,plate_id,row_id,column_idcolumns (16 rows x 24 columns per plate).
- spacr.sim.generate_power_law_distribution(num_elements, coeff)[source]¶
Return a normalized power-law probability vector of length
num_elements.- Parameters:
num_elements – Length of the returned distribution.
coeff – Positive exponent applied as
i^-coeff.
- Returns:
NumPy array that sums to 1.
- spacr.sim.generate_shap_summary_plot(df, target='prauc', clean=True, dst=None)[source]¶
Fit a RandomForest and render a SHAP summary plot over the standard sweep features.
- Parameters:
df – DataFrame with sweep columns and
target.target – Column predicted by the regressor. Default
'prauc'.clean – When True, drop constant columns before fitting.
- Returns:
The current Matplotlib figure (SHAP creates it as a side effect).
- spacr.sim.generate_well_score(cell_scores)[source]¶
Aggregate cell-level scores into per-well summary rows.
- Parameters:
cell_scores – DataFrame indexed by cells with
plate_row_column,is_activeandgene_idcolumns.- Returns:
DataFrame indexed by
plate_row_columnwithaverage_active_score,gene_list, andscorecolumns.
- spacr.sim.get_optimum_threshold(cell_pr_dict)[source]¶
Return the classification threshold that maximises F1 in a PR result dict.
- Parameters:
cell_pr_dict – Dict as returned by
compute_precision_recall().- Returns:
Threshold value that maximises the F1 score.
- spacr.sim.gini(x)[source]¶
Return the Gini coefficient of
xvia the ranked-sum formulation.Reference: StatsDirect non-parametric methods (https://www.statsdirect.com/help/nonparametric_methods/gini.htm).
- Parameters:
x – 1-D array-like; all values treated equally.
- Returns:
Gini coefficient in
[0, 1].
- spacr.sim.gini_coefficient(x)[source]¶
Return the Gini coefficient of
xvia the pairwise absolute difference formula.- Parameters:
x – 1-D array-like of non-negative values.
- Returns:
Gini coefficient in
[0, 1].
- spacr.sim.gini_gene_well(x)[source]¶
Return the Gini coefficient of
xusing a memory-cheap upper-triangle sum.- Parameters:
x – 1-D array-like of non-negative values.
- Returns:
Gini coefficient in
[0, 1]; 0 is perfect equality, 1 is perfect inequality.
- spacr.sim.normalize_array(arr)[source]¶
Return
arrmin-max scaled into[0, 1].- Parameters:
arr – Input NumPy array.
- Returns:
Normalized array of the same shape.
- spacr.sim.plot_confusion_matrix(data, ax, title)[source]¶
Render a 2x2 confusion matrix as an annotated Seaborn heatmap.
- Parameters:
data – 2x2 NumPy confusion matrix ordered
[[TN, FP], [FN, TP]].ax – Matplotlib axes to draw into.
title – Axes title.
- Returns:
None.
- spacr.sim.plot_correlation_matrix(df, annot=False, cmap=None, clean=True, dst=None)[source]¶
Render a lower-triangular correlation heatmap of the standard sweep + metric columns.
- Parameters:
df – DataFrame containing sweep variables plus
prauc,roc_aucand related outputs.annot – When True, write numeric correlations in each cell.
cmap – Colormap name or object.
Noneuses the house diverging map, which is the one case the style allows a diverging map for – a correlation is genuinely signed. The parameter used to default to'inferno'and then be OVERWRITTEN two lines later, so a caller who passed one got the diverging map anyway and never knew.clean – When True, drop constant columns before computing correlations.
- Returns:
The generated Matplotlib figure.
- spacr.sim.plot_feature_importance(df, target='prauc', exclude=None, clean=True, dst=None)[source]¶
Train a RandomForestRegressor on sweep variables and plot the resulting importances.
- Parameters:
df – DataFrame with sweep columns and
target.target – Column predicted by the regressor. Default
'prauc'.exclude – Column name or list of columns to remove from the feature set.
clean – When True, drop constant columns before fitting.
- Returns:
The generated Matplotlib figure.
- spacr.sim.plot_histogram(data, x_label, ax, color, title, binwidth=0.01, log=False)[source]¶
Draw a Seaborn density histogram on
axfor the given column.- Parameters:
data – Data source passed to
sns.histplot.x_label – Column name plotted on the x-axis.
ax – Matplotlib axes to draw into.
color – Bar/fill color.
title – Axes title.
binwidth – Histogram bin width; falsy uses Seaborn’s default.
log – When True, apply a log scale to the y-axis.
- Returns:
None.
- spacr.sim.plot_partial_dependences(df, target='prauc', clean=True, dst=None)[source]¶
Fit a GradientBoostingRegressor and plot partial dependences for every sweep feature.
- Parameters:
df – DataFrame with sweep columns and
target.target – Column predicted by the regressor. Default
'prauc'.clean – When True, drop constant columns before fitting.
- Returns:
The generated Matplotlib figure.
- spacr.sim.plot_roc_pr(data, ax, title, x_label, y_label)[source]¶
Plot a ROC or PR curve with a diagonal random-classifier reference line.
- Parameters:
data – DataFrame containing the
x_labelandy_labelcolumns.ax – Matplotlib axes to draw into.
title – Axes title.
x_label – Column name plotted on the x-axis.
y_label – Column name plotted on the y-axis.
- Returns:
None.
- spacr.sim.plot_simulations(df, variable, x_rotation=None, legend=False, grid=False, clean=True, verbose=False)[source]¶
Grid-plot PR-AUC vs
variablefor every unique combination of the other sweep dimensions.- Parameters:
df – DataFrame containing
prauc,variableand the standard grouping columns (number_of_active_genes,avg_reads_per_gene, …).variable – Column plotted on the x-axis of each subplot.
x_rotation – Degrees to rotate x-tick labels.
Noneuses 45.legend – When True, show the per-subplot legend.
grid – When True, draw grid lines.
clean – When True, drop grouping columns whose values never vary.
verbose – When True, annotate each subplot with its filter conditions.
- Returns:
The generated Matplotlib figure.
- spacr.sim.power_law_dist_gen(df, avg, well_ineq_coeff)[source]¶
Return
len(df)per-well values sampled from an average-scaled power-law.- Parameters:
df – DataFrame of wells whose length sets the sample size.
avg – Scale factor applied to each drawn probability.
well_ineq_coeff – Power-law exponent (larger = more unequal).
- Returns:
NumPy array of per-well quantities.
- spacr.sim.read_simulations_table(db_path)[source]¶
Return the
simulationstable fromdb_pathas a DataFrame.- Parameters:
db_path – Path to a SQLite database written by
save_data().- Returns:
DataFrame of the
simulationstable, orNoneon failure.
- spacr.sim.regression_roc_auc(results_df, active_gene_list, control_gene_list, alpha=0.05, optimal=False)[source]¶
Score regression hits against ground truth and compute ROC/PR metrics.
Marks each gene as active/inactive/control, derives a hit cutoff from the control coefficients, and returns ROC/PR curves, a confusion matrix and per-run summary statistics.
- Parameters:
results_df – Regression output with
gene,coefandP>|t|.active_gene_list – Gene indices considered truly active.
control_gene_list – Gene indices used to derive the coefficient cutoff.
alpha – Significance threshold applied to p-values. Default
0.05.optimal – When True, use the F1-optimal probability threshold instead of
0.5for the final confusion matrix.
- Returns:
Tuple
(results_df, reg_roc_dict_df, reg_pr_dict_df, reg_cm, sim_stats)wheresim_statsis a single-row DataFrame.
- spacr.sim.remove_columns_with_single_value(df)[source]¶
Return
dfwithout columns whose values are constant across rows.- Parameters:
df – Source DataFrame.
- Returns:
Copy of
dfwith zero-variance columns dropped.
- spacr.sim.remove_constant_columns(df)[source]¶
Return
dflimited to columns that contain more than one unique value.- Parameters:
df – Source DataFrame.
- Returns:
Copy of
dfwith constant columns dropped.
- spacr.sim.run_and_save(i, settings, time_ls, total_sims)[source]¶
Worker that runs one simulation, saves outputs, and appends its runtime.
- Parameters:
i – Simulation index (used for filenames and tagging).
settings – Simulation settings dict.
time_ls – Shared list receiving the elapsed time in seconds.
total_sims – Total simulation count (used for progress display only).
- Returns:
Tuple
(i, sim_time, None).
- spacr.sim.run_experiment(plate_map, number_of_genes, active_gene_list, avg_genes_per_well, sd_genes_per_well, avg_cells_per_well, sd_cells_per_well, well_ineq_coeff, gene_ineq_coeff)[source]¶
Simulate one cell-level screening experiment and return per-cell + summary tables.
Draws per-well gene assignments from a power-law distribution and per-well cell counts from a gamma/Poisson mixture, then labels each cell active/inactive.
- Parameters:
plate_map – DataFrame of wells with plate/row/column identifiers.
number_of_genes – Total number of genes in the pool.
active_gene_list – Gene indices considered active (positive class).
avg_genes_per_well – Mean genes-per-well before power-law scaling.
sd_genes_per_well – Standard deviation of genes-per-well.
avg_cells_per_well – Mean cells-per-well.
sd_cells_per_well – Standard deviation of cells-per-well.
well_ineq_coeff – Power-law exponent for well-level inequality.
gene_ineq_coeff – Power-law exponent for gene-level inequality.
- Returns:
Tuple
(cell_df, genes_per_well_df, wells_per_gene_df, df_ls)wheredf_lscontains per-well gene counts, per-gene well counts, per-well Gini values, per-gene Gini values, gene weights and well weights.
- spacr.sim.run_multiple_simulations(settings)[source]¶
Fan out the sweep from
generate_parameters()across a process pool.Processing workers start ten seconds apart and drive
run_and_save()with captured database operations. One parent-owned writer commits each complete simulation and its retry ticket atomically. Explicit resource overloads receive one serial retry after all primary tasks finish; database overloads use the writer’s separate final queue. Invalid data and cancellation are not retried. An unsuccessful simulation or final database write raises instead of reporting a successful sweep.- Parameters:
settings – Sweep-settings dict. Must include
max_workers.database_write_queue_giboptionally overrides the saved queue RAM allowance, from zero to 64 GiB; zero buffers serialized packets on disk. Pending packets remain below the dated output folder on interruption.- Returns:
None.
- Raises:
ValueError – if the queue RAM allowance is invalid.
RuntimeError – if a final database write fails. Processing errors propagate with their original exception type.
- spacr.sim.run_simulation(settings)[source]¶
Run one end-to-end pooled-screen simulation and return every intermediate table.
Composes
generate_gene_list(),generate_plate_map(),run_experiment(),classifier(), cell/well aggregation,sequence_plates()andregression_roc_auc()into a single call.- Parameters:
settings – Dict of simulation parameters (gene counts, distribution moments, sequencing error, classifier accuracy, …).
- Returns:
Tuple
(cell_scores, cell_roc_dict_df, cell_pr_dict_df, cell_cm, well_score, gene_fraction_map, metadata, results_df, reg_roc_dict_df, reg_pr_dict_df, reg_cm, sim_stats, genes_per_well_df, wells_per_gene_df, dists).
- spacr.sim.save_data(src, output, settings, save_all=False, i=0, variable='all')[source]¶
Persist one simulation’s output tables to a SQLite database under
src.In the default mode only a concatenated summary row (settings + sim_stats + Gini metrics) is appended to a
simulationstable. Whensave_allis True, every intermediate table is written under its canonical name.- Parameters:
src – Output directory containing
simulations.db.output – 14-element list from
run_simulation().settings – Simulation settings dict recorded as the first row.
save_all – When True, write every intermediate table separately.
i – Simulation index used to tag the summary row.
variable – Name of the swept variable used for tagging.
- Returns:
None.
- spacr.sim.save_plot(fig, src, variable, i)[source]¶
Save a Matplotlib figure to
<src>/<variable>/<i>_figure.The output format and resolution follow the user’s figure preferences. A default installation writes PDF.
- Parameters:
fig – Figure to save.
src – Root directory for outputs.
variable – Sub-folder name (the swept variable label).
i – Zero-padded simulation index used in the file name.
- Returns:
None.
- spacr.sim.save_shap_plot(fig, src, variable, i)[source]¶
Save a SHAP figure to
<src>/<variable>/<i>_figure.Through
spacr.plot.save_figure()for the same reasonsave_plot()is: the format follows the preference, not a literal.- Parameters:
fig – Matplotlib figure to save.
src – base output directory.
variable – subdirectory name identifying the plotted variable.
i – identifier prefixed to the output filename.
- spacr.sim.sequence_plates(well_score, number_of_genes, avg_reads_per_gene, sd_reads_per_gene, sequencing_error=0.01)[source]¶
Simulate sequencing of every well and return per-well gene fraction and metadata.
Each gene present in a well accrues a Poisson-distributed read count that may be reassigned to a random well with probability
sequencing_error.- Parameters:
well_score – DataFrame with a
gene_listcolumn per well.number_of_genes – Number of distinct genes in the pool.
avg_reads_per_gene – Mean of the per-gene read count distribution.
sd_reads_per_gene – Standard deviation of that distribution.
sequencing_error – Probability of assigning a read to the wrong well. Default
0.01.
- Returns:
Tuple
(gene_fraction_map, metadata)DataFrames indexed by well.
- spacr.sim.update_scores_and_get_cm(cell_scores, optimum)[source]¶
Add a per-threshold predicted-label column and return the confusion matrix.
- Parameters:
cell_scores – DataFrame with columns
is_activeandscore.optimum – Score threshold used to binarise predictions.
- Returns:
Tuple
(cell_scores, cell_cm)wherecell_cmis a NumPy confusion matrix.
- spacr.sim.validate_and_adjust_beta_params(sim_params)[source]¶
Clamp per-run Beta variances so the requested mean/variance is feasible.
- Parameters:
sim_params – List of per-run parameter dicts with
positive_mean,negative_mean,positive_variance,negative_variance.- Returns:
The same list with any infeasible variances capped to 99% of the theoretical maximum for the requested mean.
- spacr.sim.vis_dists(dists, src, v, i)[source]¶
Save side-by-side histograms of the six per-run distributions in
dists.- Parameters:
dists – Six arrays in order
[genes/well, wells/gene, gini_well, gini_gene, gene_weights, well_weights].src – Output directory used by
save_plot().v – Variable label passed through to
save_plot().i – Simulation index passed through to
save_plot().
- Returns:
None.
- spacr.sim.visualize_all(output)[source]¶
Render the full 13-panel diagnostic figure for one simulation output.
- Parameters:
output – The 14-element list returned by
run_simulation()(all elements beforedists).- Returns:
The generated Matplotlib figure.
Nested helpers¶
- _commit_simulation_packet.dispatch(operation)¶
Write one captured simulation frame through the normal table writer.
spacr/sim.py:1034
- _simulation_unit_bytes._largest(key, default)¶
The largest value swept for
key.spacr/sim.py:1073
- classifier.calc_alpha_beta(mean, variance)¶
Return
(alpha, beta)for a Beta with the given mean and variance.spacr/sim.py:315
- classifier.get_score(is_active)¶
Return a Beta sample from the correct or incorrect class distribution.
spacr/sim.py:327