Skip to content

eta

tit.jobs.eta

Wall-clock estimates for one job -- PlanCost.eta_minutes.

Why this exists

The UI used to state a job's duration as a constant typed into a button ("Generate (~40 min)") or as a per-kind row in the renderer's step table. Both are wrong the moment anything about the job changes: a leadfield for a 19-electrode cap and one for a 256-electrode cap differ by more than an order of magnitude (one FEM solve per electrode), and the same job on a native x86 host and under Rosetta/QEMU emulation differs by another factor of three.

The model

Every estimate has the same shape::

minutes = (fixed + per_unit * units) * mesh_scale * system.factor / parallel
  • units is what actually drives the run: electrodes in the cap (leadfield), electrode pairs (simulation), candidate evaluations (ex/mEx), multistart runs x iterations (flex), stages (pre-processing).
  • mesh_scale is the subject's head mesh measured against ernie's, using the .msh file size as a cheap proxy for the element count (a stat() call, not a parse), clamped so a missing or unusual mesh cannot produce an absurd number.
  • system.factor folds in the two machine facts that matter: how many cores the container may actually use (:func:tit.cpu.effective_cpus -- the cgroup limit, not the host's core count), and whether it is running emulated (an amd64 image under Rosetta on Apple Silicon, which is where every constant below was measured).

Calibration

The constants are the native (non-emulated, 12-core) numbers implied by real runs on this project's ernie subject, all measured under emulation and therefore divided by :data:EMULATION_FACTOR:

=============== ==================================================== ===================== Kind Measured run (emulated, ernie mesh ~184 MB) Implied constants =============== ==================================================== ===================== leadfield 76-electrode EEG10-10_UI_Jurak_2007: 16:27:26 -> fixed 2.0 min, 16:50:27 = 23.0 min (76 electrode placements at 0.276 min/electrode ~8.2 s, 75 FEM solves at ~6.2 s + ~2.2 s overhead) sim one TI montage (2 pairs): 00:50:33 -> 00:56:10 = 0.8 + 1.7/pair + 1.4 5.6 min (2 solves at ~5.8 s; meshing the pads and the post-processing dominate) ex docs_ex_large, 16 807 combinations: ~14 min 2.0 min + 7.1e-4/eval flex VAL_rthal_flex_focality, n_multistart=2: ~25 min 1.0 + 12.0/multistart at the default budget pre the per-stage figures the renderer's step table charm 15, FastSurfer 30, carried (measured on the same machine) QSIPrep 40, ... =============== ==================================================== =====================

So on the machine they were measured on (12 cores, emulated -> factor 3.0) the model reproduces those wall clocks, and it extrapolates rather than guesses everywhere else: the same subject's 256-electrode cap comes out near 70 min, not the 40 min the button used to claim.

Every number here is an estimate and the UI must label it as one.

SystemProfile dataclass

SystemProfile(cpus: int, emulated: bool, factor: float)

The machine facts an estimate depends on.

detect_system cached

detect_system() -> SystemProfile

This machine's :class:SystemProfile (cached -- it cannot change under us).

Source code in tit/jobs/eta.py
@lru_cache(maxsize=1)
def detect_system() -> SystemProfile:
    """This machine's :class:`SystemProfile` (cached -- it cannot change under us)."""
    cpus = effective_cpus()
    emulated = _is_emulated(platform.machine(), _cpuinfo())
    factor = _cpu_factor(cpus) * (EMULATION_FACTOR if emulated else 1.0)
    return SystemProfile(cpus=cpus, emulated=emulated, factor=round(factor, 3))

mesh_scale

mesh_scale(subject_id: str | None) -> float

The subject's head mesh against ernie's, as a cheap stat() on <sid>.msh.

1.0 when the mesh is unknown (the job would create it, or the project is not readable). Clamped to [0.4, 3.0]: the proxy is a file size, not an element count, and a stray file must not turn an estimate into a fantasy.

Source code in tit/jobs/eta.py
def mesh_scale(subject_id: str | None) -> float:
    """The subject's head mesh against ernie's, as a cheap ``stat()`` on ``<sid>.msh``.

    1.0 when the mesh is unknown (the job would create it, or the project is not readable).
    Clamped to [0.4, 3.0]: the proxy is a file size, not an element count, and a stray file
    must not turn an estimate into a fantasy.
    """
    path = _mesh_path(subject_id)
    if not path:
        return 1.0
    try:
        size = os.path.getsize(path)
    except OSError:
        return 1.0
    return min(3.0, max(0.4, size / REFERENCE_MESH_BYTES))

electrode_count

electrode_count(subject_id: str | None, eeg_net: str | None) -> int | None

Number of electrodes in eeg_net for subject_id, from the subject's cap CSV.

None when the cap cannot be read -- the caller then has no leadfield estimate to give, which is honest, rather than a number invented from a default cap size.

Source code in tit/jobs/eta.py
def electrode_count(subject_id: str | None, eeg_net: str | None) -> int | None:
    """Number of electrodes in *eeg_net* for *subject_id*, from the subject's cap CSV.

    ``None`` when the cap cannot be read -- the caller then has no leadfield estimate to give,
    which is honest, rather than a number invented from a default cap size.
    """
    if not subject_id or not eeg_net:
        return None
    try:
        from tit import get_path_manager
        from tit.catalog import _read_cap_electrode_labels

        pm = get_path_manager()
        name = eeg_net if eeg_net.lower().endswith(".csv") else f"{eeg_net}.csv"
        from tit.paths import resolve_within, validate_name

        path = resolve_within(
            pm.project_dir,
            os.path.join(pm.eeg_positions(subject_id), validate_name(name, "EEG net")),
        )
        labels = _read_cap_electrode_labels(path)
    except Exception:
        return None
    return len(labels) or None

eta_minutes

eta_minutes(kind: str, config: dict[str, Any] | None = None, *, resolved: dict[str, Any] | None = None, subject_id: str | None = None, n_jobs: int = 1, parallel: int = 1, system: SystemProfile | None = None) -> float | None

Estimated wall-clock minutes for the whole plan, or None when it cannot be modelled.

Parameters

kind : str Job kind ("leadfield", "sim", "ex", ...). config : dict, optional The raw request config -- read for the montages, the cap, the DE budget. resolved : dict, optional The planner's own resolved block, so ex/mEx reuse the n_combinations it already counted and pre reuses its stage list rather than re-deriving either. subject_id : str, optional Whose head mesh and electrode cap to measure. Defaults to config["subject_id"]. n_jobs : int Jobs in the plan (a batch of subjects/montages runs the same estimate that many times). parallel : int How many of them run at once. system : SystemProfile, optional Overrides :func:detect_system (tests, and a future "estimate for another machine").

Returns

float or None Minutes, rounded to one decimal. None for a kind with no model, and for a leadfield whose cap cannot be read.

Source code in tit/jobs/eta.py
def eta_minutes(
    kind: str,
    config: dict[str, Any] | None = None,
    *,
    resolved: dict[str, Any] | None = None,
    subject_id: str | None = None,
    n_jobs: int = 1,
    parallel: int = 1,
    system: SystemProfile | None = None,
) -> float | None:
    """Estimated wall-clock minutes for the whole plan, or ``None`` when it cannot be modelled.

    Parameters
    ----------
    kind : str
        Job kind (``"leadfield"``, ``"sim"``, ``"ex"``, ...).
    config : dict, optional
        The raw request config -- read for the montages, the cap, the DE budget.
    resolved : dict, optional
        The planner's own ``resolved`` block, so ex/mEx reuse the ``n_combinations`` it already
        counted and ``pre`` reuses its stage list rather than re-deriving either.
    subject_id : str, optional
        Whose head mesh and electrode cap to measure. Defaults to ``config["subject_id"]``.
    n_jobs : int
        Jobs in the plan (a batch of subjects/montages runs the same estimate that many times).
    parallel : int
        How many of them run at once.
    system : SystemProfile, optional
        Overrides :func:`detect_system` (tests, and a future "estimate for another machine").

    Returns
    -------
    float or None
        Minutes, rounded to one decimal.  ``None`` for a kind with no model, and for a leadfield
        whose cap cannot be read.
    """
    config = config or {}
    sid = subject_id or (
        config.get("subject_id") if isinstance(config.get("subject_id"), str) else None
    )
    sys_profile = system or detect_system()
    scale = mesh_scale(sid)

    if kind == "pre" and (
        config.get("run_freesurfer")
        or any(
            isinstance(stage, dict) and "G2c" in _as_list(stage.get("tags"))
            for stage in _as_list((resolved or {}).get("stages"))
        )
    ):
        # No measured FreeSurfer calibration yet; do not reuse FastSurfer's estimate.
        return None
    if kind == "pre":
        per_job = _pre_minutes(resolved)
        scale = 1.0  # pre-processing BUILDS the mesh; there is nothing to measure yet.
    elif kind == "leadfield":
        n_electrodes = electrode_count(sid, config.get("eeg_net"))
        if n_electrodes is None:
            return None
        per_job = FIXED_MIN["leadfield"] + PER_UNIT_MIN["leadfield"] * n_electrodes
    elif kind == "sim":
        per_job = FIXED_MIN["sim"] + PER_UNIT_MIN["sim"] * _sim_pairs(config)
        # "vn" / "dir" / "mc" are the anisotropic (DTI) conductivity models; "scalar" is the
        # isotropic default (tit.sim.config.SimulationConfig.conductivity).
        if config.get("conductivity") in ("vn", "dir", "mc"):
            per_job *= ANISOTROPY_FACTOR
    elif kind in ("ex", "mex"):
        combos = _search_combinations(resolved)
        if combos is None:
            return None
        per_job = FIXED_MIN[kind] + PER_UNIT_MIN[kind] * combos
    elif kind in ("flex", "flex_adaptive", "flex_pareto"):
        per_job = FIXED_MIN[kind] + PER_UNIT_MIN[kind] * _flex_units(config)
    else:
        return None

    jobs = max(1, n_jobs)
    lanes = min(max(1, parallel), jobs)
    total = per_job * scale * sys_profile.factor * jobs / lanes
    return round(total, 1)