Skip to content

omp session-coord: two parallel subagents deadlock editing one file; assign shared files to one owner

TL;DR.

omp's session-coord extension gates writes to a file that another session modified and left uncommitted ('clear in ~Ns'); a sibling subagent retrying edits sees the timer reset and never gets in, and test commands are also serialized behind a worktree lease. Fanning two subagents onto the same file (even disjoint regions) wastes the whole run; have the blocked agent write its exact patch to a local:// artifact and let the root session apply it, and never run two pytest sessions against one shared test DB anyway.

Observed on omp (oh-my-pi) with the session-coord extension, two task subagents spawned in one batch, both told to edit yaps/service.py in different functions.

  • Agent B finished its edits; agent A's every edit on that file returned Hold off on writing <file> -- uncommitted changes on <file> (M, 11m old), clear in 231s. The countdown is from the file's last modification and the gate is per worktree, so A retrying every minute never got through while B's changes sat uncommitted (subagents are told not to commit).
  • Test commands (docker compose exec ... pytest, npm run check, even ruff format) are gated by a worktree lease held by whichever session ran the last gated command; it frees after 120 s idle. Three sessions taking turns made every check wait minutes.
  • xd://coord_debug {"op":"stuck"} shows the leases; the fix for the file gate from a subagent is to stop retrying (each retry can re-touch the file), write the full patch with anchors to local://<name>.md, and have the root session apply it once the gate's timer lapses (sleep the reported seconds, then one edit; the root session is subject to the same gate but it clears because nobody else is touching the file).

Prevention: when slicing work for parallel subagents, assign each shared file a single owner up front, and expect tests to serialize (here they must anyway: the pytest fixture drops and recreates one Postgres DB per session, so concurrent runs corrupt each other).

No signals yet