spacr.qt.widgets.formula

Computed columns — a small expression language, parsed rather than eval-ed.

ratio = cell_area / cell_perimeter ** 2 is the sort of thing a user wants five seconds after seeing a measurement table, and until this module the answer was “export it and open pandas”. The obvious implementation is one line:

frame[name] = eval(expression, {}, frame)          # never

and that line hands anyone who can type into the box — or anyone who can put a saved chart spec in front of the user — the whole interpreter. A spaCR settings file, a shared .spacr project, a macro pasted from a colleague: all of them would become executable. So the expression is tokenised, parsed into an AST of six node types, and walked. There is no eval, no exec, no compile and no attribute access anywhere in the grammar; a hostile string does not fail a blacklist, it fails to parse.

The grammar

expression := or_expr
or_expr    := and_expr ('or' and_expr)*
and_expr   := not_expr ('and' not_expr)*
not_expr   := 'not' not_expr | comparison
comparison := sum (('<' | '<=' | '>' | '>=' | '==' | '!=') sum)?
sum        := product (('+' | '-') product)*
product    := unary (('*' | '/' | '//' | '%') unary)*
unary      := ('-' | '+') unary | power
power      := atom ('**' unary)?
atom       := NUMBER | COLUMN | FUNCTION '(' [expression {',' expression}] ')'
            | '(' expression ')'

COLUMN     := [A-Za-z_][A-Za-z_0-9]*  |  '`' <anything but a backtick> '`'
NUMBER     := digits ['.' digits] [('e'|'E') ['+'|'-'] digits]

Five deliberate absences, each of which is a class of attack or a class of confusion rather than a feature nobody got round to:

  • no string literals — so no payload can be smuggled through as data, and the language has nothing to feed to a function that might want a name;

  • no attribute access, no indexing, no assignment — ., [ and a bare = are rejected by the tokeniser with the reason, so the ().__class__.__bases__ ladder does not get as far as the parser;

  • no lambdas, no comprehensions, no statements — an expression evaluates to one column and nothing else;

  • no arbitrary names — a name is either a column of the frame or one of FUNCTIONS. __import__, open and eval are not blocked as special cases; they are not columns, and the error says so;

  • comparisons do not chain — 0 < area < 5 is refused rather than silently parsed as something numpy would evaluate elementwise into a shape nobody meant. The message says to write the conjunction.

Numbers are read as floats, always. That is not cosmetic: 9 ** 9 ** 9 with Python integers is a multi-second allocation of a number with 300 million digits — a denial of service typed in eleven characters — and with floats it is inf in a nanosecond. The node count and nesting depth are capped as well (MAX_NODES, MAX_DEPTH), so a pathological expression is refused at parse time rather than during a redraw.

What a formula computes over

The whole table, never the filtered view. zscore(area) computed over whatever the Local Data Filter currently shows would change every time a slider moved, which means the column would not be a column: two charts drawn a second apart would disagree, and an exported CSV would record one arbitrary moment. So compute() is given the loaded frame and the aggregates (AGGREGATE_FUNCTIONS) reduce over all of it. Filtering happens afterwards, to the computed column like any other.

Formulas are evaluated in list order, and each one sees the columns added before it — so density = count / area then log_density = log(density) works, and referring to a formula defined further down is an error that says so rather than a NaN column.

What comes out

Arithmetic gives float64; a comparison or an and/or/not gives bool. The bool case is on purpose: a boolean column lands in spacr.qt.widgets.data_filter_panel.classify_columns() as a category, so infected = pathogen_count > 0 immediately becomes a tick box in the filter panel and a colour channel in the Graph Builder, which is what anyone writing that expression wanted next.

Division by zero produces inf rather than an exception — a ratio column with a few infinities is a real answer about a few objects — but the count of non-finite results is carried in ColumnResult.n_nonfinite and said out loud in ColumnResult.notice. A column that is 90% inf is a mistake, and the only way to notice is to be told.

Text columns are not addressable. gene cannot be used in a formula, and the error says to use the Local Data Filter for it. Arithmetic on a gene name has no meaning, and coercing it to NaN silently would produce an all-NaN column with no explanation.

No Qt in here — pure numpy and pandas, like spacr.qt.widgets.graph_spec and spacr.selection, so the grammar can be tested without a display and used from a notebook.

Exceptions

FormulaError

A formula that cannot be computed, with the reason and the place.

Classes

Binary

Every infix operator, arithmetic, comparison and boolean alike.

Call

One of FUNCTIONS, applied to its arguments.

Column

A reference to a column of the frame.

ColumnFormula

One computed column: a name and an expression.

ColumnResult

One computed column's values and what computing them cost.

FormulaSet

An ordered list of ColumnFormula, and the frame they make.

Node

Base of the five node types. Frozen: a parsed formula is a value.

Number

A numeric literal. Always a float — see the module docstring.

Unary

-x, +x or not x.

Functions

compute(→ Tuple[pandas.DataFrame, List[ColumnResult]])

Add one column per formula to a copy of frame.

evaluate(→ Any)

Evaluate node over frame.

parse(→ Node)

Parse expression into an AST.

referenced_columns(→ Tuple[str, ...])

Every column node reads, in first-appearance order, de-duplicated.

tokenize(→ Tuple[_Token, ...])

Split expression into tokens, or say exactly where it stopped making

unparse(→ str)

Print node back as an expression, fully parenthesised.

Module Contents

exception spacr.qt.widgets.formula.FormulaError[source]

Bases: ValueError

A formula that cannot be computed, with the reason and the place.

Every message names the thing that is wrong — the column, the function, the character, its position in the string — because the user is looking at a one-line text box and “invalid syntax” tells them nothing about which of the forty characters to change.

Initialize self. See help(type(self)) for accurate signature.

class spacr.qt.widgets.formula.Binary[source]

Bases: Node

Every infix operator, arithmetic, comparison and boolean alike.

Parameters:
  • op – the operator text: +, -, *, /, //, %, **, a comparison (<, <=, >, >=, ==, !=), and or or.

  • left – the left operand’s node.

  • right – the right operand’s node.

class spacr.qt.widgets.formula.Call[source]

Bases: Node

One of FUNCTIONS, applied to its arguments.

Parameters:
  • func – the function name, a key of FUNCTIONS.

  • args – the argument nodes, in call order.

class spacr.qt.widgets.formula.Column[source]

Bases: Node

A reference to a column of the frame.

Parameters:

name – the column name, read from the frame when the node is evaluated.

class spacr.qt.widgets.formula.ColumnFormula[source]

One computed column: a name and an expression.

Frozen and JSON round-tripping like GraphSpec, for the same reason — a derived column is part of what a saved analysis is, and a chart of ratio that cannot say what ratio was is not reproducible.

Parameters:
  • name – the new column’s name. A plain identifier, so it can be typed into another formula without backticks.

  • expression – the source text.

  • replace – allow overwriting an existing column of the same name. Off by default: silently replacing cell_area with something derived would make every earlier chart of that column unreproducible.

Raises:

FormulaError – on an unusable name, or an expression that does not parse — at construction, so a bad formula never reaches a render.

__post_init__() → None[source]

Normalise the name and expression, and parse the expression now.

Parsing at construction is the point: an unparseable formula cannot then be stored, serialised, or reach a redraw – it fails where the user typed it.

Raises:

FormulaError – if the name is not a usable column name, or the expression will not parse.

describe() → str[source]

The formula, marked when it reads the WHOLE table.

The mark matters: a formula using a table-wide statistic cannot be computed per row, so it behaves differently under filtering and the reader should not have to work that out from the expression.

Returns:

a one-line description.

classmethod from_dict(payload: Mapping[str, Any]) → ColumnFormula[source]

Rebuild a formula from plain data.

UNKNOWN KEYS ARE IGNORED rather than raising, so a saved set from a later version still opens with the parts this one knows.

Parameters:

payload – what to_dict() produced.

Returns:

the rebuilt formula.

classmethod from_json(text: str) → ColumnFormula[source]

Rebuild a formula from JSON text.

Parameters:

text – the JSON text.

Returns:

the rebuilt formula.

inputs() → Tuple[str, ...][source]

The columns this formula reads.

to_dict() → Dict[str, Any][source]

This formula as plain data.

Returns:

a JSON-safe dict.

to_json() → str[source]

This formula as JSON text, keys sorted so the file is diffable.

Returns:

the JSON text.

uses_whole_table() → bool[source]

Whether one object’s value depends on the other objects.

True for every TABLE_DEPENDENT_FUNCTIONS call anywhere in the expression. Worth showing beside the formula, because such a column is not a property of the object: the same formula over two plates gives two different columns, and re-running it after a re-segmentation moves every value.

property ast: Node[source]

The parsed expression.

Held on an attribute that is not a dataclass field, so two formulas compare equal on their name and text — which is what a saved formula is — rather than on two structurally identical trees.

class spacr.qt.widgets.formula.ColumnResult[source]

One computed column’s values and what computing them cost.

Parameters:
  • formula – the formula that produced the column; its name heads notice.

  • values – the column, aligned to the frame it was computed over.

  • n_rows – number of rows in the frame the column was computed over.

  • n_nonfinite – NaN and ±inf in the result. Reported rather than hidden — a ratio column that is a third infinities is a division by a zero the user did not know was there, and it looks identical to a good column on a chart that drops non-finite points.

  • n_input_missing – rows where at least one input was already missing, so a NaN in the output can be attributed rather than guessed at.

property notice: str[source]

what came out, and what did not.

Type:

One line for the panel

class spacr.qt.widgets.formula.FormulaSet[source]

An ordered list of ColumnFormula, and the frame they make.

Ordered, not a set: each formula sees the columns the ones before it added, which is what makes density then log_density work. A formula that refers to one defined below it gets an error saying to move it up, rather than a column of NaN.

Mutable, unlike the specs elsewhere in this package, because it is a list the user edits rather than a value that is diffed — but every entry in it is frozen, and to_dict() round-trips.

__len__() → int[source]

Return how many formulas the set holds.

add(formula: ColumnFormula) → FormulaSet[source]

Append, replacing any formula of the same name.

Parameters:

formula – the formula to append; any existing formula with the same name is removed first.

apply(frame: pandas.DataFrame) → Tuple[pandas.DataFrame, List[ColumnResult]][source]

frame plus one column per formula, and a result per column.

A copy — the loaded table is never mutated, so removing a formula removes its column rather than leaving it behind, and two screens sharing a frame do not grow each other’s columns.

Parameters:

frame – the table to add columns to; it is copied, never modified.

clear() → FormulaSet[source]

Drop every formula. Returns self, so it chains.

Returns:

this set, now empty.

describe() → str[source]

Every formula in evaluation order, or that there are none.

Returns:

a one-line description.

classmethod from_dict(payload: Mapping[str, Any]) → FormulaSet[source]

Rebuild a set from plain data.

Parameters:

payload – what to_dict() produced.

Returns:

the rebuilt set.

classmethod from_json(text: str) → FormulaSet[source]

Rebuild a set from JSON text.

Parameters:

text – the JSON text.

Returns:

the rebuilt set.

remove(name: str) → FormulaSet[source]

Drop the formula called name. Returns self, so it chains.

A name that is not there is not an error: removing something already gone is the state the caller wanted.

Parameters:

name – the formula’s name.

Returns:

this set.

to_dict() → Dict[str, Any][source]

The whole set as plain data.

Returns:

a JSON-safe dict holding every formula.

to_json() → str[source]

The set as JSON text, keys sorted so the file is diffable.

Returns:

the JSON text.

property is_empty: bool[source]

Whether this set computes nothing.

Returns:

True when empty.

property names: Tuple[str, ...][source]

Every formula’s name, in evaluation order.

Returns:

the names.

class spacr.qt.widgets.formula.Node[source]

Base of the five node types. Frozen: a parsed formula is a value.

class spacr.qt.widgets.formula.Number[source]

Bases: Node

A numeric literal. Always a float — see the module docstring.

Parameters:

value – the literal’s value, as a float.

class spacr.qt.widgets.formula.Unary[source]

Bases: Node

-x, +x or not x.

Parameters:
  • op – the operator text: -, + or not.

  • operand – the node the operator applies to.

spacr.qt.widgets.formula.compute(frame: pandas.DataFrame, formulas: Sequence[ColumnFormula]) → Tuple[pandas.DataFrame, List[ColumnResult]][source]

Add one column per formula to a copy of frame.

In list order, each formula seeing what the earlier ones added.

Parameters:
  • frame – the table to add columns to; it is copied, never modified.

  • formulas – formulas to apply in order; each is written to a column named by its name and may read columns added by earlier ones.

Returns:

(frame_with_columns, results).

Raises:

FormulaError – naming the formula that failed. Nothing is added when one fails — a half-applied set would leave the user with some of the columns they asked for and no way to tell which.

spacr.qt.widgets.formula.evaluate(node: Node, frame: pandas.DataFrame) → Any[source]

Evaluate node over frame.

Parameters:
  • node – root of the expression tree to evaluate.

  • frame – the table whose columns the expression reads; its length sets the length of array results.

Returns:

an ndarray the length of the frame, or a python float when the whole expression reduces (mean(area)). The caller broadcasts — keeping scalars scalar is what lets area / mean(area) cost one division rather than two array allocations.

Raises:

FormulaError – naming the column or the function that failed.

spacr.qt.widgets.formula.parse(expression: str) → Node[source]

Parse expression into an AST.

Parameters:

expression – formula text, converted with str() and stripped; an empty text raises FormulaError.

Raises:

FormulaError – for anything that is not a valid expression in the grammar above, with the position and what to write instead.

spacr.qt.widgets.formula.referenced_columns(node: Node) → Tuple[str, ...][source]

Every column node reads, in first-appearance order, de-duplicated.

Parameters:

node – root of the expression tree to scan.

spacr.qt.widgets.formula.tokenize(expression: str) → Tuple[_Token, ...][source]

Split expression into tokens, or say exactly where it stopped making sense.

Parameters:

expression – formula text, converted with str(); at most 2,000 characters, and a backticked name may contain spaces.

Raises:

FormulaError – on an over-long expression, an unterminated backtick or a character the language does not contain — with the character, its position, and what to write instead where there is an alternative.

spacr.qt.widgets.formula.unparse(node: Node) → str[source]

Print node back as an expression, fully parenthesised.

Not for display — for the round trip. parse(unparse(parse(text))) equalling parse(text) is what says the parser and the tree agree about precedence, which is the one property of a hand-written parser that is hard to eyeball and easy to get wrong.

Parameters:

node – root of the expression tree to print; a node of an unknown type raises FormulaError.

Nested helpers

ColumnFormula.uses_whole_table.walk(node: Node) → bool

Whether any node needs the whole column rather than one row.

An aggregate cannot be computed row by row, so this decides whether the formula can stream or has to hold the table.

spacr/qt/widgets/formula.py:1112

_aggregate.call(values)

Reduce the values to one number, ignoring non-finite entries.

An all-non-finite column gives nan rather than raising: an empty aggregate is an answer the caller can carry, and an exception here would take down a whole computed column for one bad group.

spacr/qt/widgets/formula.py:383

_elementwise.call(*args)

Apply the function across the arrays, warnings suppressed.

A column of real data contains zeros and negatives, so a log or a divide will legitimately produce inf and nan; the VALUE is what the formula means, and numpy’s warning about it is not something a user can act on.

spacr/qt/widgets/formula.py:353

evaluate.walk(item: Node) → Any

Evaluate one node of the tree, recursing into its children.

spacr/qt/widgets/formula.py:930

referenced_columns.walk(item: Node) → None

Collect every column the expression names, depth first.

spacr/qt/widgets/formula.py:832