Skip to content

Run coordination: stale run.lock never reaped; renameSync fails with EPERM on Windows #28

Description

@LeonardoTemporal

Run coordination: stale run.lock is never reaped, and rename fails with EPERM on Windows

Version: 0.6.1 (ac37ed1)
Platform: issue 1 is cross-platform; issue 2 is Windows-only (11 10.0.26200, Node v24.11.1)

Two independent defects in src/runs.ts. Both surface in multi-node runs, which is exactly
what --run / --node / --team is for.


1. Stale run.lock is never reaped — run becomes permanently unusable

acquireRunLock writes the holder's pid into the lock file:

fd = openSync(lockPath, "wx", privateFileMode);
...
writeFileSync(fd, `${process.pid}\n`);

but nothing ever reads it back. On contention the loop only sleeps:

} catch {
  if (Date.now() >= deadline) {
    throw new Error(`run is locked: ${runId}`);
  }
  sleepSync(10);
}

The release path is a closure that runs closeSync + rmSync. If the holder dies before
reaching it — crash, SIGKILL, terminal closed, CI timeout — run.lock stays on disk
forever.

Effect: every later operation on that run blocks for the full 30 s deadline and then
throws run is locked: <runId>. run view, run message, run mark, run wait are all
affected. The run is bricked until someone manually deletes .../run.lock. There is no
command to clear it and the error does not say the lock is stale or where it lives.

The recorded pid shows the liveness check was intended; it just isn't wired up.

Suggested fix

Reap the lock when the recorded pid is gone:

function processAlive(pid: number): boolean {
  try {
    process.kill(pid, 0);
    return true;
  } catch (error) {
    return (error as NodeJS.ErrnoException).code === "EPERM";
  }
}

EPERM from kill(pid, 0) means the process exists but is owned by another user — alive,
not dead. Treating it as dead would let two holders in at once.

Then in the retry loop, attempt reapDeadProcessLock(lockPath) before sleeping: read the
pid, and if it is not a positive integer or processAlive(pid) is false, unlink the lock
and retry immediately. A malformed or empty lock file should also be reaped — it means a
holder died between openSync and writeFileSync.


2. renameSync over run.json fails with EPERM on Windows

writeRun writes a temp file and renames it into place:

const tmpPath = `${path}.tmp-${process.pid}`;
writeFileSync(tmpPath, `${JSON.stringify(run, null, 2)}\n`, { mode: privateFileMode });
chmodSync(tmpPath, privateFileMode);
renameSync(tmpPath, path);

On POSIX, rename(2) over an open file is atomic and always succeeds. On Windows,
MoveFileEx fails with EPERM / EACCES / EBUSY when any other process holds a handle
to the destination — including a reader doing a plain readRun.

The lock does not prevent this: readers do not take the run lock, so a single concurrent
run view is enough to make the writer throw. The temp file is also left behind, since
there is no cleanup on the failure path.

Suggested fix

Bounded retry, Windows-only so POSIX behaviour and latency are unchanged:

function replaceFileWithRetry(source: string, destination: string): void {
  const maxAttempts = process.platform === "win32" ? 20 : 1;
  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    try {
      renameSync(source, destination);
      return;
    } catch (error) {
      const code = (error as NodeJS.ErrnoException).code;
      const retryable =
        process.platform === "win32" &&
        (code === "EPERM" || code === "EACCES" || code === "EBUSY");
      if (!retryable || attempt === maxAttempts - 1) {
        throw error;
      }
      sleepSync(25);
    }
  }
}

20 attempts at 25 ms is ~500 ms worst case, which covers the realistic contention window.
maxAttempts = 1 off Windows keeps non-Windows behaviour byte-for-byte identical.

Wrap the call in try/finally and rmSync(tmpPath, { force: true }) so a failed replace
does not leave run.json.tmp-<pid> behind.

This is a mitigation, not a guarantee — under sustained reader pressure a writer could still
exhaust the attempts. A complete fix would open the destination with
FILE_SHARE_DELETE semantics or use ReplaceFileW.


Diagnostics

A stderr trace gated behind an env var makes the interleaving visible without adding cost
when unset:

function debugRunLock(message: string): void {
  if (process.env.HEADLESS_DEBUG_RUN_LOCK === "1") {
    process.stderr.write(`[${process.pid}] ${message}\n`);
  }
}

Notes

Both fixes are implemented and running locally against 0.6.1. Happy to open a PR with tests
in tests/run-coordination.test.ts if useful.

Separately: agent autodetection does not work on Windows. It probes for extensionless
binaries in PATH, so it never matches claude.cmd, codex.cmd or opencode.cmd, and the
agent has to be named explicitly. Worth its own issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions