locks
tit.jobs.locks ¶
Advisory, filesystem-based locks (TODO.md §2.4).
Keys are held by the runner process itself (with tit.jobs.locks.hold(...), a few lines
per __main__) and predicted by the scheduler by reading the same lock directory with
:func:holders — so a bare simnibs_python -m tit.sim cfg.json from a second shell and a
server-submitted job share one policy and one lock directory. Direct Python calls
(run_simulation() in a notebook) remain advisory: nothing stops them, they just don't
register a hold.
Storage: one directory per held (resource, mode, holder) triple,
jobs/.locks/<sha1(discriminator)[:16]>/lock.json, created with :func:os.mkdir (atomic on
bind mounts — no : or arbitrary user text in path components; NTFS-hostile characters never
appear because the directory name is a hash). Reader/writer semantics: a "write" request
conflicts with any existing holder of the same resource (read or write) other than itself; a
"read" request conflicts only with an existing "write" holder of the same resource — so
concurrent readers each get their own directory (discriminated by job id) and never collide.
The full lock-key table (which kind holds which resource, in which mode) lives in
:func:keys_for, transcribed from TODO.md §2.4.
LockConflictError ¶
Bases: RuntimeError
Raised by :func:hold on a write-against-write conflict (always), and on any other
conflict in strict mode (TIT_LOCKS=strict).
LockRequest
dataclass
¶
holders ¶
Every currently-held lock, as descriptor dicts ({key, resource, mode, job_id, pid,
create_time, ts}). Stale entries (process gone) are removed as a side effect when
reconcile_stale is true (the default; boot-time reconciliation passes it explicitly, but
every read benefits since a crashed holder never cleans up after itself).
Source code in tit/jobs/locks.py
parse_key ¶
parse_key(key: str) -> LockRequest
Inverse of :attr:LockRequest.key ("<resource>:<mode>"); unknown/missing mode
suffix defaults to "write" (the conservative choice — treat it as exclusive).
Source code in tit/jobs/locks.py
release_job ¶
Forcibly remove every lock directory held by job_id, regardless of liveness.
Used by :meth:tit.jobs.manager.JobManager.force — a job being forced to a terminal state
isn't waited on, so its locks must be dropped immediately rather than on the next lazy
:func:holders reconciliation. Returns the number of directories removed.
Source code in tit/jobs/locks.py
reconcile ¶
Boot-time sweep: drop lock directories whose holder process is gone.
Returns the number of stale directories removed.
Source code in tit/jobs/locks.py
match_conflicts ¶
match_conflicts(requests: list[LockRequest], current_holders: list[dict[str, Any]], *, self_job_id: str | None = None) -> list[dict[str, Any]]
Pure function: which of current_holders block requests (reader/writer rule, §2.4).
Takes an already-fetched holder list (rather than reading the lock directory itself) so the scheduler can reuse one on-disk scan across every queued job in a tick, and so this rule is unit-testable without any filesystem I/O.
Source code in tit/jobs/locks.py
conflicts ¶
conflicts(project_dir: str, requests: list[LockRequest], *, self_job_id: str | None = None) -> list[dict[str, Any]]
Holders that block requests (reader/writer rule, §2.4), excluding self_job_id.
Source code in tit/jobs/locks.py
hold ¶
hold(project_dir: str, job_id: str, requests: list[LockRequest], *, pid: int | None = None, strict: bool | None = None) -> Iterator[None]
Acquire requests for the duration of the context (called by the runner process).
Policy has two tiers. A write against a live write always raises
:class:LockConflictError, whatever the mode: two processes writing one subject's head
model is the corruption this whole module exists to prevent, and "warn and proceed" is not
a policy for it. Every other conflict (a reader against a writer, a writer against readers)
keeps the historical warn-and-continue default, and strict=True / TIT_LOCKS=strict
raises on those too.
A write lock's directory is named after the resource alone, so two jobs wanting the same
exclusive resource want the same directory. It is never taken from a live owner: the
descriptor of a different, still-running job is left exactly as it was, so holders
keeps naming the real owner and that owner's own release does not free someone else's lock.
Source code in tit/jobs/locks.py
217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 | |
keys_for ¶
keys_for(kind: str, subject_ids: list[str], config: dict[str, Any] | None = None) -> list[LockRequest]
Lock requests one job of kind needs, given its (still-serialized) config.
Best-effort against configs whose exact dataclass shape is owned by another lane (marked below): falls back to a coarse subject-scoped lock rather than requesting nothing, so two jobs of an unrecognised shape still serialize instead of silently racing.