storage
tit.storage ¶
How much disk a TI-Toolbox project is using, and what is using it.
The System page shows the machine's disk limit; this answers the other half — what of that is us, broken down the way a person thinks about their project ("the flex searches are 180 GB") rather than by directory.
Three things this module is careful about, because each one silently produces a wrong number:
Real disk usage, not apparent size. st_blocks * 512 is what the
filesystem actually charges for a file. st_size over-counts sparse files
and under-counts the tail block of every small one; on a project with a hundred
thousand small meshes those disagree by gigabytes.
Hardlinks counted once. A (st_dev, st_ino) set is kept for the whole
walk. QSIPrep and several of our own steps hardlink outputs rather than copy
them, so summing per-file sizes double-counts them into a total larger than the
volume itself — which is exactly the kind of figure that destroys trust in a
storage page.
Classification comes from :class:~tit.paths.PathManager, never from globs
written here. Every kind's root is asked for by name, so a layout change in
paths.py moves this module with it instead of leaving it quietly attributing
simulations to "Other". Classification is longest-prefix-wins, which is what
makes nested kinds work: .../Simulations/<sim>/Analyses/ is an analysis, and
the rest of Simulations/ is a simulation.
The walk is O(files) and genuinely slow on a large project (a 900 GB volume can
take minutes), so it is never on a request path that something is waiting on:
:func:scan_project is called from a background thread and its result is cached
on disk (:func:load_cache / :func:save_cache).
ItemUsage
dataclass
¶
One named thing inside a kind — a subject, or a simulation.
kind_prefixes ¶
(path, kind) pairs, longest first, from the PathManager.
Every entry is a directory the PathManager names. Nothing here is a glob or
a hand-written path fragment, so a layout change in paths.py cannot leave
this attributing real outputs to "Other" — except for the per-subject entries
below, which are the PathManager's own per-subject accessors applied to the
subjects it lists.
Source code in tit/storage.py
classify ¶
The kind owning path: the longest matching prefix, else "other".
Source code in tit/storage.py
scan_project ¶
scan_project(project_dir: str | None = None, *, pm: Any | None = None, should_stop: Callable[[], bool] | None = None) -> ProjectStorage
Walk the project once and total real disk usage per kind.
should_stop is polled between directories so a scan can be abandoned when
the server is shutting down or the project has changed underneath it; the
result is then marked partial rather than silently short.
Source code in tit/storage.py
252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 | |
cache_path ¶
save_cache ¶
save_cache(result: ProjectStorage, pm: Any | None = None) -> None
Write atomically: a half-written cache would be read as a real scan.