Skip to content

examples

tit.examples

Example data: a content-addressed catalogue of datasets with independent parts.

Modelled on 3D Slicer's SampleData module (and Tetravox's Sample Data dialog): the catalogue (:file:catalog.json, package data) lists every file by its sha256, and the store is one GitHub release (v1 on idossha/ti-toolbox-example-data <https://github.com/idossha/ti-toolbox-example-data>_, a repository of its own) whose assets carry those hashes as names. Nothing is placed in a project until the downloaded bytes hash to the catalogue entry.

A dataset is one head; a part is one thing you can download. Two datasets ship today -- ernie (the SimNIBS example subject) and mni152 (the template) -- each with a nifti part and a headmodel part::

ernie/nifti      sub-ernie/anat/sub-ernie_T1w.nii.gz, _T2w.nii.gz
ernie/headmodel  derivatives/SimNIBS/sub-ernie/m2m_ernie/   (m2m_ernie.tar.gz, unpacked)
mni152/nifti     sub-MNI152/anat/sub-MNI152_T1w.nii.gz
mni152/headmodel derivatives/SimNIBS/sub-MNI152/m2m_MNI152/

Each part is fetched and detected on its own: deleting sub-ernie/anat leaves ernie/headmodel installed, and vice versa. This is what the old four-overlapping-samples catalogue could not say -- its head-model sample contained the NIfTIs, so removing them flipped the head model to "not installed" too.

Where a part lands is declared in the catalogue, not in this module: dest (a project-root template with {subject}) is the directory its files go into -- a .tar.gz is unpacked there -- verify the paths that must exist for it to count as installed (default: the plain file names), and target what :func:fetch returns (default: dest). Adding a third part later (a FreeSurfer tree, a worked simulation) is therefore an edit to :file:catalog.json alone.

This module is plain functions -- :func:catalogue, :func:status, :func:fetch -- and knows nothing about jobs, stages or :mod:tit.jobs.events. Downloading example data is not a pipeline: it was briefly wired as a project_init job, which made asking an established project for a sample reprint the initializer's "New project detected" banner.

Entry points: python -m tit.examples --project DIR [--list] [ernie/headmodel ...], the server routes GET /api/example-data and POST /api/example-data/{dataset_id}/{part_id} (:mod:tit.server.routes.example_data, a background thread and a poll), and the desktop's Add example data? chooser and Help > Example data tab. Stdlib only, so it runs on the host and in the container, over a verified TLS context from :mod:tit.certs. :func:fetch_ernie is kept as the notebook's one-liner and means both ernie parts.

Progress module-attribute

Progress = Callable[[str, str, int, int], None]

progress(part_id, file_name, received_bytes, total_bytes) across the whole part.

Part dataclass

Part(id: str, title: str, meaning: str, dest: str, files: tuple[PartFile, ...], dataset_id: str = '', subject: str = '', target: str = '', verify: tuple[str, ...] = (), derivative: str = '')

One independently downloadable piece of a dataset.

dest/verify/target are project-root-relative templates carrying {subject}; they are the whole of this part's placement policy, which is why a new part kind is a catalogue edit rather than a code change.

full_id property

full_id: str

ernie/headmodel -- how the CLI, the route and the desktop name this part.

catalogue

catalogue() -> list[Dataset]

Every dataset, in catalogue order, each with its parts.

Source code in tit/examples/__init__.py
def catalogue() -> list[Dataset]:
    """Every dataset, in catalogue order, each with its parts."""
    out = []
    for raw in _load_catalog()["datasets"]:
        subject = raw["subject"]
        parts = []
        for p in raw["parts"]:
            files = tuple(PartFile(**f) for f in p["files"])
            parts.append(
                Part(
                    id=p["id"],
                    title=p["title"],
                    meaning=p["meaning"],
                    dest=p["dest"],
                    files=files,
                    dataset_id=raw["id"],
                    subject=subject,
                    target=p.get("target", p["dest"]),
                    verify=tuple(p.get("verify", [f.name for f in files])),
                    derivative=p.get("derivative", ""),
                )
            )
        out.append(
            Dataset(**{k: v for k, v in raw.items() if k != "parts"}, parts=tuple(parts))
        )
    return out

parse_part_id

parse_part_id(part_id: str) -> tuple[str, str]

"ernie/headmodel" -> ("ernie", "headmodel"); also accepts ernie:headmodel.

Source code in tit/examples/__init__.py
def parse_part_id(part_id: str) -> tuple[str, str]:
    """``"ernie/headmodel"`` -> ``("ernie", "headmodel")``; also accepts ``ernie:headmodel``."""
    text = part_id.replace(":", "/")
    dataset, sep, part = text.partition("/")
    if not sep or not dataset or not part:
        raise KeyError(
            f"{part_id!r} is not a DATASET/PART id; see `python -m tit.examples --list`"
        )
    return dataset, part

part_by_id

part_by_id(dataset_id: str, part_id: str | None = None) -> Part

The part named either as part_by_id("ernie", "headmodel") or ("ernie/headmodel").

Source code in tit/examples/__init__.py
def part_by_id(dataset_id: str, part_id: str | None = None) -> Part:
    """The part named either as ``part_by_id("ernie", "headmodel")`` or ``("ernie/headmodel")``."""
    if part_id is None:
        dataset_id, part_id = parse_part_id(dataset_id)
    for p in dataset_by_id(dataset_id).parts:
        if p.id == part_id:
            return p
    raise KeyError(
        f"unknown example part {dataset_id}/{part_id}; see `python -m tit.examples --list`"
    )

parts

parts() -> list[Part]

Every part of every dataset, in catalogue order.

Source code in tit/examples/__init__.py
def parts() -> list[Part]:
    """Every part of every dataset, in catalogue order."""
    return [p for d in catalogue() for p in d.parts]

status

status(project_dir: str | Path) -> list[dict[str, Any]]

Per-part {id, dataset, part, installed, bytes} from the disk alone -- no network.

Source code in tit/examples/__init__.py
def status(project_dir: str | Path) -> list[dict[str, Any]]:
    """Per-part ``{id, dataset, part, installed, bytes}`` from the disk alone -- no network."""
    root = Path(project_dir).expanduser()
    return [
        {
            "id": p.full_id,
            "dataset": p.dataset_id,
            "part": p.id,
            "installed": _installed(root, p),
            "bytes": p.bytes,
        }
        for p in parts()
    ]

fetch

fetch(dataset_id: str, part_id: str | None = None, project_dir: str | Path | None = None, *, force: bool = False, progress: Progress | None = None, log=print) -> Path

Download one part into project_dir and return the directory it filled.

Called either as fetch("ernie", "headmodel", project) or fetch("ernie/headmodel", project) -- the two-argument form is what the CLI and the notebook use.

Idempotent: returns at once when the part's verify paths are already in place unless force. Each file is streamed to a temporary directory and sha256-verified before it touches the project; a mismatch raises :class:ValueError and leaves the project as it was.

Source code in tit/examples/__init__.py
def fetch(
    dataset_id: str,
    part_id: str | None = None,
    project_dir: str | Path | None = None,
    *,
    force: bool = False,
    progress: Progress | None = None,
    log=print,
) -> Path:
    """Download one part into *project_dir* and return the directory it filled.

    Called either as ``fetch("ernie", "headmodel", project)`` or ``fetch("ernie/headmodel",
    project)`` -- the two-argument form is what the CLI and the notebook use.

    Idempotent: returns at once when the part's ``verify`` paths are already in place unless
    *force*. Each file is streamed to a temporary directory and sha256-verified **before** it
    touches the project; a mismatch raises :class:`ValueError` and leaves the project as it was.
    """
    if project_dir is None:
        # Two-argument form: fetch("ernie/headmodel", project) -- the second positional is the
        # project directory, not a part.
        if part_id is None:
            raise TypeError("fetch() needs a project directory")
        project_dir = part_id
        dataset_id, part_id = parse_part_id(dataset_id)
    part = part_by_id(dataset_id, part_id)
    root = Path(project_dir).expanduser().resolve()
    if _installed(root, part) and not force:
        log(f"{part.full_id} already present: {_target_dir(root, part)}", flush=True)
        return _target_dir(root, part)

    total = max(part.bytes, 1)
    done = 0
    last_pct = -1

    def report(name: str, base: int, received: int) -> None:
        nonlocal last_pct
        if progress:
            progress(part.full_id, name, base + received, total)
        pct = (base + received) * 100 // total
        if pct != last_pct:
            last_pct = pct
            log(f"download {pct}%", flush=True)

    with tempfile.TemporaryDirectory(prefix="tit-examples-") as tmp:
        staged: list[tuple[PartFile, Path]] = []
        for f in part.files:
            tmp_file = Path(tmp) / f.name
            log(f"downloading {f.name} ({f.bytes / 1e6:.0f} MB)", flush=True)
            _download(f, tmp_file, lambda n, name=f.name, base=done: report(name, base, n))
            staged.append((f, tmp_file))
            done += f.bytes
        log("checksums ok; installing", flush=True)
        for f, tmp_file in staged:
            _place(root, part, f, tmp_file)

    _register(root, part)
    target = _target_dir(root, part)
    log(f"{part.full_id} ready: {target}", flush=True)
    return target

fetch_ernie

fetch_ernie(project_dir: str | Path, *, force: bool = False, log=print) -> Path

The example notebook's one-liner: both ernie parts (T1+T2 and m2m_ernie).

Returns the head model directory, which is what the notebook goes on to simulate on.

Source code in tit/examples/__init__.py
def fetch_ernie(project_dir: str | Path, *, force: bool = False, log=print) -> Path:
    """The example notebook's one-liner: **both** ernie parts (T1+T2 and ``m2m_ernie``).

    Returns the head model directory, which is what the notebook goes on to simulate on.
    """
    target = None
    for part in dataset_by_id(ERNIE).parts:
        target = fetch(ERNIE, part.id, project_dir, force=force, log=log)
    assert target is not None  # noqa: S101 - the catalogue always ships ernie with parts
    return target