spacr.qt.stall_watch¶
Name the call that freezes the interface.
A STALLED GUI THREAD LEAVES NO TRACEBACK. It is not an exception and not a
fault, so faulthandler cannot see it, the log simply stops, and the
only report anyone can give is “spaCR froze” with nothing under it.
The only thing that can name the call is a sample of the main thread’s
stack taken WHILE it is stuck, from another thread.
SPACR_WATCH_GUI_STALLS=1 spacr
Every time the event loop stops answering for longer than
STALL_SECONDS, the exact Python stack of the GUI thread is written
to stderr and appended to LOG_PATH. The last frame of that stack
is the blocking call.
WHY THIS IS IN THE PACKAGE RATHER THAN IN tools/.
tools/watch_the_gui_thread.py does the same thing and has to build the
QApplication itself to attach before any screen exists. That stopped
working the day spacr.qt.app.launch() began constructing its own, and
Qt refuses a second.
Reaching in from outside was then tried two more ways and failed twice more.
Hosting qt.run() inside another process makes Home itself time out at
thirty seconds.
The other way was a sitecustomize on the startup benchmark path. It is
imported, and what it changes is never the object that runs.
A flag read inside launch is the one place that cannot be missed, and it
composes with every other driver – the benchmark included, which is what
this was written for.
WHAT IT COSTS WHEN OFF: one os.environ.get. When on: a 100 ms timer on
the GUI thread that increments an integer, and a daemon thread that
compares two integers four times a second.
Functions¶
|
Report the GUI thread's stack whenever it stops answering. |
Module Contents¶
- spacr.qt.stall_watch.watch_this_application(app, *, stall_seconds: float | None = None, echo: bool = True)[source]¶
Report the GUI thread’s stack whenever it stops answering.
- Parameters:
app – the live
QApplication. The heartbeat timer is parented to it so the timer lives exactly as long as the application does.stall_seconds – override
STALL_SECONDSfor one call.echo – also write each report to
sys.stderr. PassFalseunder a harness that captures streams. THE FILE IS THE RECORD AND STDERR IS A CONVENIENCE: this writes from a DAEMON THREAD, and a thread writing into a stream the harness is swapping underneath it crashed pytest inside its owncapture.py– not in our write, which is guarded, but in pytest reading a stream that moved while it read. A tool must not write to a stream it does not own when somebody else is holding it.
- Returns:
the watcher thread, or
Noneif Qt could not be reached.
THE HEARTBEAT IS THE MEASUREMENT. A
QTimeron the GUI thread bumps an integer; a daemon thread watches the integer. If it stops moving the GUI thread is not running the event loop, which is exactly the condition being hunted – and the reason a timer cannot report it itself: a wedged loop does not deliver the timer either.
Nested helpers¶
- watch_this_application._flush_samples() None¶
Write what held the thread through the stall that just ended.
spacr/qt/stall_watch.py:186
- watch_this_application._flush_when_the_stall_ends() None¶
Summarise a stall the moment the loop answers again.
THE LAST STALL OF A SESSION WAS NEVER SUMMARISED, and that is the one anybody runs this for.
_flush_sampleswas called only when the NEXT stall began, so a process that wedges once and is then killed – a run against a sleeping autofs mount, say – left the first traceback and no distribution at all. The summary is the part that separates the call HOLDING the thread from the one that merely happened to be running when the sample was taken.The heartbeat resuming is the end of the stall, so that is where the summary belongs.
tickcannot do it: it runs on the GUI thread, and the whole point is that the GUI thread was not running.spacr/qt/stall_watch.py:203
- watch_this_application._where(frame) str¶
The innermost frame, as
file:line function.spacr/qt/stall_watch.py:177
- watch_this_application.tick() None¶
Record that the event loop is still turning.
spacr/qt/stall_watch.py:107
- watch_this_application.watch() None¶
Sample the GUI thread through each stall and summarise it.
ONE SNAPSHOT NAMES WHERE THE THREAD WAS, NOT WHERE THE TIME WENT, and those are different questions whenever the stall is a loop rather than a single blocking call. Measured: four stalls of the same module sweep gave four different last frames – an event filter, a screen constructor, a settings-search install, another event filter – which is a list of suspects rather than an answer.
So the stack is sampled every
POLL_SECONDSFOR AS LONG AS the stall lasts, and the last frames are counted. A call that holds the thread appears in most samples; one that merely happened to be running appears in one.spacr/qt/stall_watch.py:128