runner
tit.jobs.runner ¶
Runner interface + the local subprocess implementation (TODO.md §2.3).
Runner is deliberately narrow (spawn only) so a future SbatchRunner (3.1, HPC) can
drop in: everything else — waiting for exit, killing a process tree — works on a bare pid and
doesn't care how the process was started, which is also what makes cancelling a re-attached
job (one this server process never spawned) possible.
RunRequest
dataclass
¶
RunRequest(job_id: str, argv: list[str], cwd: str, env: dict[str, str] = dict(), stdout_path: str = devnull)
Everything :meth:Runner.spawn needs to start one job's process.
Runner ¶
Bases: ABC
Starts a job's OS process. Waiting/cancelling operate on the resulting pid directly.
spawn
abstractmethod
async
¶
spawn(request: RunRequest) -> Process
LocalPopenRunner ¶
Bases: Runner
asyncio.create_subprocess_exec with the safe-spawn options from TODO.md §2.3.
Its own process group (:func:tit.jobs.processes.spawn_kwargs -- POSIX
start_new_session, Windows CREATE_NEW_PROCESS_GROUP -- so a cancel's SIGTERM/SIGKILL
(or their Windows equivalents) never hits the server itself), stdin DEVNULL, stdout
appended to stdout_path with stderr merged in, close_fds=True. Never preexec_fn
(unsafe with threads) and never a pipe (a dead server would BrokenPipe the runner instead
of leaving it to finish on its own).
ResourceSampler ¶
ResourceSampler(pid: int)
Per-job CPU%/RSS sampler over the job's whole process tree, kept for the job's lifetime.
psutil.Process.cpu_percent(interval=None) is a delta against the previous call on the
same object, so a fresh Process(pid) per poll always reads 0.0. One sampler therefore
owns one Process per pid it has seen (the root and every descendant --
SimNIBS/FastSurfer/PARDISO spawn workers), created on first sight and dropped when it exits,
and every :meth:sample sums CPU% and RSS across the live tree.
Memory is the tree's PSS (proportional set size) where the platform reports it (Linux,
/proc/<pid>/smaps_rollup, ~0.06 ms per process), so the leadfield a dozen joblib workers
share copy-on-write is counted once, not twelve times -- summing plain RSS over an ex-search
tree read 54 GB where PSS reads 9 GB. Elsewhere it falls back to RSS. The field keeps the
name rss on the wire.
Running statistics: cpu_peak/rss_peak are the maxima; cpu_avg/rss_avg are
simple means over samples (samples are taken at a fixed cadence by the manager, so this
is time-weighted to within one interval). The very first CPU reading of every process is
0.0 by psutil convention and is not counted, so a one-sample job does not report 0 %.
Source code in tit/jobs/runner.py
sample ¶
One reading of (CPU %, RSS bytes) summed over the tree; (None, None) if the
root is gone. Updates the running peak/average.
Source code in tit/jobs/runner.py
runner_env ¶
runner_env(job_id: str, events_file: str, *, interface: str = 'api', base_env: dict[str, str] | None = None, cpus: float | None = None, kind: str | None = None, subject_ids: list[str] | tuple[str, ...] | None = None) -> dict[str, str]
The child process environment (TODO.md §2.3): unbuffered/faulthandler Python, job
identity, thread knobs derived from the budget, and TIT_SERVER_TOKEN /
TIT_SERVER_SETTINGS_FILE scrubbed so a job can never read the server's own auth secret
directly, or the 0600 settings file a --reload parent hands its child (which itself
carries the token in plain JSON) -- see tit/server/__main__.py's module docstring,
which promises both are scrubbed (ra_14 finding #10: only the first one actually was).
Also adds -no_signal_handler to PETSC_OPTIONS (see
:data:PETSC_NO_SIGNAL_HANDLER) so cancelling a job does not end its log in a PETSc/MPI
crash block; a value the caller already set is kept and appended to.
Source code in tit/jobs/runner.py
is_alive ¶
True if pid is a live, non-zombie process — and, when create_time is known, it's
still the same process (not pid reuse after the original exited). Never true for pid 1 or
this server's own pid, regardless of what create_time claims (see
:func:_is_untouchable_pid).
Source code in tit/jobs/runner.py
terminate_tree
async
¶
terminate_tree(pid: int, create_time: float | None = None, *, grace_s: float = DEFAULT_GRACE_S) -> None
Snapshot descendants, SIGTERM the tree, wait up to grace_s, SIGKILL survivors.
Uses asyncio.sleep for the grace period so the caller's event loop keeps servicing other
jobs while this one shuts down. Never raises: a pid that's already gone is a no-op. Also a
no-op for pid 1 or this server's own pid (ra_14 finding #5, see :func:_is_untouchable_pid)
-- a re-attached "running" job can only ever point there via a corrupted or crafted
status.json, never a job this server actually spawned itself.
Source code in tit/jobs/runner.py
stop_docker_siblings
async
¶
Best-effort docker stop for any container labelled tit.job_id=<id>.
QSIPrep/QSIRecon's DooD builders add this label (TODO.md §2.3); a job that never spawned a sibling container, or a host without a docker CLI, makes this a silent no-op. Every call is bounded by timeout_s so an unresponsive daemon degrades to a logged warning.
Source code in tit/jobs/runner.py
stop_docker_siblings_via_engine
async
¶
Stop this job's sibling containers through the bounded Engine-API client.
Used by startup reconciliation (JobManager._reconcile_all), which must not shell out:
at startup a docker CLI may not be on PATH at all, and every step before the server
accepts requests has to be hard-bounded. :func:stop_docker_siblings (the CLI form) stays
for the cancel path, whose behaviour and tests predate this. Never raises — an unreachable
or wedged daemon degrades to a logged warning and 0.