From a97614555c53d9ff066226b7304e4e7e5d86136d Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:07:29 +0000 Subject: [PATCH 1/8] openspec: track the head-run-teardown change --- .../.openspec.yaml | 2 + .../fix-head-mode-island-teardown/design.md | 59 ++++++++++++++++++ .../fix-head-mode-island-teardown/proposal.md | 33 ++++++++++ .../specs/head-run-teardown/spec.md | 61 +++++++++++++++++++ .../fix-head-mode-island-teardown/tasks.md | 24 ++++++++ 5 files changed, 179 insertions(+) create mode 100644 openspec/changes/fix-head-mode-island-teardown/.openspec.yaml create mode 100644 openspec/changes/fix-head-mode-island-teardown/design.md create mode 100644 openspec/changes/fix-head-mode-island-teardown/proposal.md create mode 100644 openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md create mode 100644 openspec/changes/fix-head-mode-island-teardown/tasks.md diff --git a/openspec/changes/fix-head-mode-island-teardown/.openspec.yaml b/openspec/changes/fix-head-mode-island-teardown/.openspec.yaml new file mode 100644 index 00000000..75289e4b --- /dev/null +++ b/openspec/changes/fix-head-mode-island-teardown/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-09-24 diff --git a/openspec/changes/fix-head-mode-island-teardown/design.md b/openspec/changes/fix-head-mode-island-teardown/design.md new file mode 100644 index 00000000..3b50950d --- /dev/null +++ b/openspec/changes/fix-head-mode-island-teardown/design.md @@ -0,0 +1,59 @@ +# Design + +## Context + +动机见 `proposal.md`。当前机制: + +- head 模式下 `yeto launch` 用本机 sky 开 head,把 `head_cluster`、`head_job_id`、`clusters`(head + 按 `learner_cluster_names` 预先算好的 learner 名)写进 run 元数据;head 上的控制器 job 再用 **head 自己的 sky** 开 learner。两套 sky 状态库互不相通。 +- `yeto down`(`yeto/cli.py` 的 `_down_one`)按元数据里的名字在本机并行 `sky down`,异常一律 `except` 后打印 `teardown failed` 继续,随后 `runs.update_run` 并打印 `run is down`、返回 0。对 head 模式的 learner,本机 sky 必然回 "does not exist"。 +- head 自己在 run 结束时通过 `launcher.terminate_and_verify` 拆岛,并用 `_cloud_live_instances_probe` 向云核对(sky 的 `provision.query_instances`,云无关)。这条路径只在 head job 正常走到结尾时执行;`yeto down` 不用它。 +- `prB/head-side-teardown` 已实现:取消控制器 job → ssh 到 head 用其 sky 逐个 down → 按输出行确认 → 未确认则保留 head 并返回 1,重试 3 次、间隔 20 秒。它与 `pr8/rl-island-fixes` 在 `_down_one` 前面那段(Modal app 停止分支)冲突。 + +## Goals / Non-Goals + +**Goals:** +- head 模式 `yeto down` 的顺序固定为:取消控制器 job → 经 head 拆 learner 并确认 → 删 head → 云端核对 → 才报成功。 +- 把 prB 落到当前栈上,而不是再写一份。 +- 云端核对复用 `terminate_and_verify` 的探针,不新增云 SDK 依赖。 + +**Non-Goals:** +- 不改 head 自身收尾路径(已有核验)。 +- 不处理"元数据里根本没记录的 cluster"(例如 head 因 bug 用了别的名字);那属于开岛侧的契约。 +- 不做跨 run 的"清扫所有 yeto-* 实例"命令。 + +## Decisions + +### D1:合并 prB,并把 Modal 分支与 head 分支按"先分类、再分派"重排 + +冲突的两段其实互不相干:pr8 那段把 Modal 岛从 sky cluster 列表里挑出来交给 `_modal_stop_app`,prB 那段把 head 模式的 learner 交给 head。解法是在 `clusters` 上先做三分类——`modal_names`(按名字后缀)、`on_head`(`controller == "head"` 且不是 head 本身、也不是 Modal)、其余本机 sky 直接 down——再各走各的。Modal 岛不需要经 head:它是 head 上 Modal app 的函数调用,本机 `modal app stop` 就能全部结束(prB 写在 pr8 之前,当时没有 Modal 岛)。 + +*替代方案——放弃 prB、在 head 上加一个"收到信号就自拆"的 job*:需要 head 在 `yeto down` 时还活着并能调度新 job,而 prB 的注释已记录过"队列里的 sky job 没有输出也没有保证";ssh 同步执行并解析确认行更可靠。否决。 + +### D2:本机 "does not exist" 不再算成功 + +`_down_one` 的 `except` 保留,但语义改为"记录为未确认";只有 learner 已在 head 侧确认(D1 的结果集合)时,本机的失败才可以忽略。这样 `--controller local` 的 run 行为不变(本机 sky 认识自己的 cluster),head 模式则不再产生假成功。 + +### D3:云端核对复用 `terminate_and_verify`,从"head 收尾"扩展到"本机 `yeto down`" + +`_cloud_live_instances_probe(cluster)` 依赖本机 sky 状态库里的 cluster 记录来构造查询,所以: +- head cluster:本机有记录,直接 `terminate_and_verify(sky, head_cluster)` 替换现在的裸 `sky.down`。 +- head 模式的 learner:本机没有记录,探针只能在 head 上构造。让 prB 的 `HEAD_DOWN_SCRIPT` 在 head 上调用 `terminate_and_verify` 而不是裸 `sky.down`,确认行由它的返回值决定;这样 learner 的云端核对发生在 head 被删之前,正好是唯一还能做的时机。 +- Modal learner:`modal app list`(或 SDK)核对 app 为 stopped 且 tasks 为 0,已有 `_modal_stop_app` 附近的调用可复用。 +- 探针返回 `None`(云不支持或构造失败)时打印"未核验、信任 sky",不算失败——与 spec 的"云端无法核验"场景一致。 + +*替代方案——直接调各云 CLI(`nebius compute instance list` 等)按名字前缀找残留*:能覆盖"元数据没记录的 cluster",但要为每个云写一份、依赖本机装了各家 CLI;sky 的 provision 查询已经按云分派。作为后续可选项记录,本 change 不做。 + +### D4:退出码即结论 + +新增的每一条失败路径都返回非零,并且 `run is down` 只在最后打印。`runs.update_run` 在部分失败时把 run 标成"teardown incomplete"而不是"down",以便 `yeto status` 能看出来、重跑 `yeto down` 能续做。 + +## Risks / Trade-offs + +- **head 已经不在了(被人手工删、或之前版本的 `yeto down` 删掉的)** → 经 head 的路径必然失败,命令非零退出并列出 learner 名;这正是希望暴露的情况,文档给出 `nebius compute instance list` 的手工兜底。 +- **ssh 到 head 依赖 sky 生成的 ssh 配置** → prB 已用 `BatchMode`/`StrictHostKeyChecking=no`;head 若刚被 `sky down` 到一半,重试逻辑覆盖。 +- **云端核对增加 `yeto down` 时长** → 每个 cluster 最多几次查询,秒级;相比一台 H100 每小时几美元可以接受。 +- **`terminate_and_verify` 在探针出错时信任 `sky down`** → 与现状相同,但现在会打印出来。 + +## Open Questions + +- 是否把"清扫所有 `-*` 实例"作为 `yeto down --sweep` 的兜底加进来(D3 的替代方案)?不影响本 change 的 spec 与任务拆分,可以另立 change。 diff --git a/openspec/changes/fix-head-mode-island-teardown/proposal.md b/openspec/changes/fix-head-mode-island-teardown/proposal.md new file mode 100644 index 00000000..043b9aac --- /dev/null +++ b/openspec/changes/fix-head-mode-island-teardown/proposal.md @@ -0,0 +1,33 @@ +# Proposal + +## Why + +head 模式下的 `yeto down ` 会把 head 删掉、却把 head 开出来的 learner 岛留在云上继续计费:岛只记在 head 的 sky 里,本机 `sky down` 得到 "does not exist" 后被当成"已经没了"吞掉,退出码 0,之后再没有任何 sky 认识这台 VM。2026-09-24 回收 `yeto-gh1`、`yeto-gh2` 时各泄漏一台 Nebius H100(live-run-failures 第 36 条),2026-09-23 的 8.4 运行也发生过一次;每次都靠人去 `nebius compute instance list` 里找。 + +## What Changes + +- **head 模式下先经 head 拆岛,再拆 head。** `yeto down` 对 `controller == "head"` 的 run,先取消 head 上的控制器 job(防止它在拆的过程中重新拉起岛),再在 head 上用 head 的 sky 逐个 down learner,逐个确认;有任何一个未确认就**保留 head**、非零退出并说明原因,而不是把 head 删掉。这部分沿用已有的 `prB/head-side-teardown` 分支,本 change 负责把它落到当前 PR 栈上(它与 `pr8/rl-island-fixes` 在同一段代码冲突)。 +- **本机拆岛失败不再静默。** 本机 `sky down` 对一个本机不认识的 learner 返回 "does not exist" 时,只有在该 learner 已经从 head 上确认拆除的情况下才算成功;否则记为未确认,走上一条的保留 head 路径。 +- **拆完之后向云查一遍。** `yeto down` 结束前,用 sky 的 provision 查询按 cluster 名到云上核对没有存活实例(`terminate_and_verify` 已经在 head 收尾时做这件事,`yeto down` 没做);查到残留就重试 down,仍在则非零退出并打印实例 id,让人能立刻删。Modal 岛按 app 状态核对(app stopped 且 tasks 为 0)。 +- **不再给出"run is down"的假成功。** 只有所有 cluster 都确认拆除、云端核对为空时,才打印 `run '' is down` 并返回 0。 + +## Capabilities + +### New Capabilities +- `head-run-teardown`:head 模式 run 的回收顺序、确认与云端核验——learner 必须从能看见它的那台 sky 拆除并逐个确认,head 只有在所有 learner 确认后才可删除,回收结束前必须向云核对无残留,任何未确认都以显式失败告终。 + +### Modified Capabilities + + +## Impact + +- `yeto/cli.py` 的 `down` 命令(当前第 1696-1783 行附近):head 模式分支、Modal 分支与新的云端核验收尾。 +- `yeto/launcher.py`:`_cloud_live_instances_probe` / `terminate_and_verify` 需要能在本机对 head cluster 使用,并(经 head)对 learner 使用。 +- `prB/head-side-teardown` 分支:合并并解决与 pr8 在 `yeto/cli.py` 的冲突;其测试 `tests/test_head_mode.py` 随之进入栈。 +- `docs/CLOUDS.md`:把"`yeto down` 后必须看 `nebius compute instance list`"的手工步骤改成对新行为的说明。 +- 不影响 `--controller local` 的 run(岛由本机 sky 开,本机 `sky down` 本来就有效),也不影响 head 自己在运行结束时的收尾(那条路径已有 `terminate_and_verify`)。 + +## Non-goals + +- 不解决 Nebius docker RL 岛停在 `INIT` 的问题(live-run-failures 第 37 条),那是开岛侧的 bug。 +- 不引入各云的专用 API 客户端;云端核验只用 sky 已经提供的 provision 查询,覆盖不到的云按现状信任 `sky down` 并明确打印"未核验"。 diff --git a/openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md b/openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md new file mode 100644 index 00000000..60903779 --- /dev/null +++ b/openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md @@ -0,0 +1,61 @@ +# Spec Delta + +## Purpose + +保证一个由 head 控制器托管的 run 在 `yeto down` 之后不会在云上留下任何计费中的机器:每个 learner 从唯一能看见它的地方拆除并逐个确认,head 只有在 learner 全部确认后才能删除,回收结束前向云核对,任何不确定都以显式失败而不是假成功结束。 + +## ADDED Requirements + +### Requirement: learner 必须从能看见它的 sky 拆除并逐个确认 + +对 head 模式的 run,`yeto down` SHALL 先在 head 上、用 head 的 sky 拆除每个 learner cluster,并为每个 learner 取得一条明确的"已拆除"或"本就不存在"的确认。本机 sky 对 learner 返回"不存在" MUST NOT 被视为该 learner 已拆除。 + +#### Scenario: 岛由 head 开出,本机 sky 不认识它 +- **WHEN** run 的元数据记录 `controller == "head"`,且某个 learner cluster 只存在于 head 的 sky 记录中 +- **THEN** `yeto down` 通过 head 执行该 learner 的拆除,并在输出中记录该 learner 的确认结果 +- **AND** 本机 sky 返回的 "does not exist" 不产生任何"已拆除"的判定 + +#### Scenario: 控制器 job 先于拆除被取消 +- **WHEN** `yeto down` 开始拆除 head 模式 run 的 learner +- **THEN** 它先取消 head 上仍在运行的控制器 job +- **AND** 拆除过程中不会有新的 learner 被重新拉起 + +### Requirement: head 只有在所有 learner 确认后才可删除 + +`yeto down` MUST NOT 删除 head,除非该 run 的每个 learner 都已取得拆除确认。存在未确认 learner 时,它 SHALL 保留 head、以非零退出码结束,并列出未确认的 learner 名字与下一步(重跑 `yeto down` 或到云控制台删除)。 + +#### Scenario: 某个 learner 无法确认拆除 +- **WHEN** 经 head 拆除后,重试用尽仍有 learner 没有确认行 +- **THEN** head 不被删除 +- **AND** 命令以非零退出码结束,错误信息列出未确认的 learner + +#### Scenario: head 暂时无法响应 +- **WHEN** 经 head 的拆除请求因 head 自身正在收尾(如其 sky 返回服务端错误)而失败 +- **THEN** `yeto down` 在有限次数内重试 +- **AND** 重试仍失败时按上一场景处理,而不是转而删除 head + +### Requirement: 回收结束前必须向云核对无残留 + +在宣布 run 已回收之前,`yeto down` SHALL 对每个能被云端查询的 cluster 向云(而不是 sky 的状态库)核对没有非终止状态的实例;发现残留 SHALL 重试拆除,重试后仍有残留 SHALL 以非零退出码结束并打印残留实例的标识。Modal learner SHALL 按其 app 已停止且无运行中任务来核对。 + +#### Scenario: 云上仍有实例存活 +- **WHEN** 拆除后云端查询仍返回该 cluster 的存活实例 +- **THEN** `yeto down` 重试拆除,重试后仍存活则非零退出并打印实例标识 +- **AND** 不打印 run 已回收的成功信息 + +#### Scenario: 云端无法核验 +- **WHEN** 某个 cluster 所在的云不支持云端实例查询,或查询本身出错 +- **THEN** `yeto down` 明确打印该 cluster 未经云端核验、按 sky 的结果信任 +- **AND** 这一情况不导致命令失败 + +### Requirement: 成功信息只在完全回收后出现 + +只有当所有 learner 已确认、head 已删除且云端核对无残留时,`yeto down` SHALL 打印 run 已回收并以零退出。任何一步未确认都 MUST NOT 产生零退出码。 + +#### Scenario: 全部回收成功 +- **WHEN** 所有 learner 确认拆除、head 删除成功、云端核对为空 +- **THEN** 命令打印 run 已回收并以零退出 + +#### Scenario: 本地模式的 run 不受影响 +- **WHEN** run 的元数据记录 `controller == "local"` +- **THEN** learner 由本机 sky 直接拆除,仍执行云端核对与成功信息的判定 diff --git a/openspec/changes/fix-head-mode-island-teardown/tasks.md b/openspec/changes/fix-head-mode-island-teardown/tasks.md new file mode 100644 index 00000000..5d3010c3 --- /dev/null +++ b/openspec/changes/fix-head-mode-island-teardown/tasks.md @@ -0,0 +1,24 @@ +# Tasks + +## 1. 把 prB 落到当前栈上 + +- [ ] 1.1 在 pr8 之上合并 `prB/head-side-teardown`,按 design D1 解决 `yeto/cli.py` 的冲突:先把 `clusters` 三分类(Modal / 经 head 的 learner / 本机直接 down),再各自分派。验证:`tests/test_head_mode.py` 中 prB 带来的两个测试与 pr8 的 Modal 停止测试同时通过 +- [ ] 1.2 新增测试:head 模式 run 含一个 Modal 岛和一个 sky 岛时,Modal 岛走 app 停止、sky 岛经 head 拆除,且两者都不经过本机 `sky down`。验证:测试通过,且断言本机 `sky.down` 从未被以 learner 名调用 + +## 2. 本机失败不再静默 + +- [ ] 2.1 改 `_down_one`:本机 `sky down` 异常记为"未确认",仅当该 cluster 已在 head 侧确认时忽略。验证:新增测试——head 模式下 learner 未在 head 确认、本机又报 "does not exist" 时,命令非零退出、head 未被删除、错误信息列出该 learner +- [ ] 2.2 确认 `--controller local` 行为不变。验证:既有 local 模式的 `yeto down` 测试全部通过,无需改动 + +## 3. 云端核对 + +- [ ] 3.1 head cluster 的拆除改用 `terminate_and_verify`(含云端探针)替换裸 `sky.down`。验证:测试用假探针模拟"仍有实例存活"→ 重试后仍存活 → 非零退出并打印实例 id;"探针为 None"→ 打印未核验、零退出 +- [ ] 3.2 把 prB 的 `HEAD_DOWN_SCRIPT` 改为在 head 上调用 `terminate_and_verify`,确认行反映其返回值。验证:`_unconfirmed_head_downs` 的解析测试覆盖"down 但云端仍存活"输出为未确认 +- [ ] 3.3 Modal learner 在 app 停止后核对 app 状态为 stopped 且 tasks 为 0,否则非零退出。验证:测试用假的 app 状态覆盖两种结果 +- [ ] 3.4 只有全部确认后才打印 `run '' is down` 并返回 0;部分失败时 `runs.update_run` 记录为 teardown 未完成,`yeto status` 能显示。验证:测试断言部分失败时不出现成功信息且状态可见 + +## 4. 文档与真机确认 + +- [ ] 4.1 更新 `docs/CLOUDS.md`:把"`yeto down` 后手工看 `nebius compute instance list`"改为说明新行为与非零退出的含义;在 live-run-failures 第 36 条下记录处理。验证:文档改动与实现一致 +- [ ] 4.2 真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** +- [ ] 4.3 真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 From 84acd6c957d7c53368ecc12798cfae26ab9e7813 Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:15:26 +0000 Subject: [PATCH 2/8] cli: tear a head run's learners down from the head and verify at the cloud Merges prB/head-side-teardown onto the stack (its Modal-vs-head conflict in cmd_down is resolved by classifying clusters first: Modal islands stop with the app, a head run's learners are torn down from the head, the rest go through this machine's sky) and closes the gaps that still leaked H100s: - a local sky.down error other than "does not exist" no longer counts as gone; a head run's learners never take the local path at all - every cluster this machine downs is confirmed at the cloud through terminate_and_verify (which now takes a down hook and, without a probe, trusts only a clean down or "does not exist"); the head-side script verifies learners the same way before the head is deleted - Modal islands are confirmed via the app's state (stopped, 0 tasks) - "run is down" and exit 0 appear only when everything is confirmed; otherwise the run is marked TEARDOWN_INCOMPLETE with the unconfirmed clusters so yeto status shows it and a rerun picks up --- tests/test_head_mode.py | 66 +++++++++++ tests/test_runs_cli.py | 110 +++++++++++++++++- tests/test_teardown_verify.py | 24 ++++ yeto/cli.py | 203 ++++++++++++++++++++++++---------- yeto/launcher.py | 22 +++- yeto/modal_runner.py | 26 +++++ yeto/runs.py | 2 + 7 files changed, 382 insertions(+), 71 deletions(-) diff --git a/tests/test_head_mode.py b/tests/test_head_mode.py index 58d8e469..130416de 100644 --- a/tests/test_head_mode.py +++ b/tests/test_head_mode.py @@ -494,6 +494,72 @@ def test_down_head_run_succeeds_after_a_retry(monkeypatch): assert len(calls) == 2 and downed == ["hr-head"] +def test_down_head_run_routes_modal_islands_to_the_app_and_sky_islands_via_head(monkeypatch): + """A head run with one Modal island and one sky island: the Modal one is + ended by stopping the app, the sky one is torn down FROM the head, and + neither ever hits this machine's sky.down (which would say "does not + exist" and orphan it).""" + make_head_meta("hm") + runs.update_run("hm", clusters=["hm-head", "hm-l0-us-east-2", "hm-l1-modal"]) + downed, on_head, stopped = [], [], [] + monkeypatch.setattr(cli, "_sky_down_cluster", downed.append) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: None) + monkeypatch.setattr(cli, "_modal_stop_app", stopped.append) + monkeypatch.setattr(cli, "_modal_app_stopped", lambda run: (True, "stopped, 0 tasks")) + + def fake(head, job, cs): + on_head.append(sorted(cs)) + return [] + + monkeypatch.setattr(cli, "_head_down_learners", fake) + assert cli.main(["down", "hm"]) == 0 + assert stopped == ["hm"] + assert on_head == [["hm-l0-us-east-2"]] + assert downed == ["hm-head"] + + +def test_down_head_run_never_trusts_a_local_does_not_exist_for_a_learner(monkeypatch, capsys): + """The head could not confirm the learner; a local sky.down of it + answering "does not exist" must not be read as "already gone".""" + make_head_meta("hn") + downed = [] + + def local_down(cluster): + downed.append(cluster) + raise ValueError(f"Cluster '{cluster}' does not exist.") + + monkeypatch.setattr(cli, "_sky_down_cluster", local_down) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: None) + monkeypatch.setattr(cli, "HEAD_DOWN_RETRY_S", 0.0) + monkeypatch.setattr(cli, "_head_down_learners", lambda head, job, cs: list(cs)) + assert cli.main(["down", "hn"]) == 1 + assert downed == [] # learners never go through the local sky; head kept + meta = runs.load_run("hn") + assert meta["state"] == runs.TEARDOWN_INCOMPLETE + assert meta["teardown_unconfirmed"] == ["hn-l0-us-east-2", "hn-l1-us-west-2"] + err = capsys.readouterr().err + assert "NOT tearing down hn-head" in err and "hn-l0-us-east-2" in err + + +def test_down_head_run_keeps_the_head_when_the_cloud_still_has_it(monkeypatch, capsys): + make_head_meta("hc") + monkeypatch.setattr(cli, "_sky_down_cluster", lambda c: None) + monkeypatch.setattr(cli, "_head_down_learners", lambda head, job, cs: []) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: (lambda: ["i-head-zombie"])) + monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + assert cli.main(["down", "hc"]) == 1 + assert runs.load_run("hc")["state"] == runs.TEARDOWN_INCOMPLETE + assert "i-head-zombie" in capsys.readouterr().err + + +def test_head_down_script_verifies_at_the_cloud_and_reports_survivors(): + script = cli.HEAD_DOWN_SCRIPT.format(clusters=["a-l0"]) + assert "from yeto.launcher import terminate_and_verify" in script + assert "still live at the cloud after down" in script + out = "[head] a-l0: still live at the cloud after down\n" + assert cli._unconfirmed_head_downs(out, ["a-l0"]) == ["a-l0"] + + def test_head_down_requires_a_confirmation_per_learner(): """A head-side teardown once printed success while the learner kept running (no output came back). Only confirmed clusters count.""" diff --git a/tests/test_runs_cli.py b/tests/test_runs_cli.py index 1d9b10f5..d58a9eb0 100644 --- a/tests/test_runs_cli.py +++ b/tests/test_runs_cli.py @@ -279,6 +279,7 @@ def test_down_stops_the_modal_app_and_skips_sky_for_modal_islands(monkeypatch, c downed, stopped = [], [] monkeypatch.setattr(cli, "_sky_down_cluster", downed.append) monkeypatch.setattr(cli, "_modal_stop_app", stopped.append) + monkeypatch.setattr(cli, "_modal_app_stopped", lambda run: (True, "stopped, 0 tasks")) assert cli.main(["down", "m1"]) == 0 assert sorted(downed) == ["m1-l0-us-east-1", "m1-syncer"] # never sky.down a Modal island assert stopped == ["m1"] # one app stop covers every Modal island of the run @@ -286,18 +287,115 @@ def test_down_stops_the_modal_app_and_skips_sky_for_modal_islands(monkeypatch, c assert "m1-l1-modal: stopped with the Modal app" in capsys.readouterr().out -def test_down_survives_sky_errors(monkeypatch, capsys): +def test_down_treats_a_cluster_sky_never_had_as_gone(monkeypatch, capsys): + """Local mode, no cloud probe: sky saying it never had the cluster is + the one down error that still counts as gone (a rerun after a clean + down must stay green).""" runs.create_run("d2", make_args_dict("d2")) runs.update_run("d2", pid=None, clusters=["d2-syncer"]) + def vanished(cluster): + raise ValueError(f"Cluster '{cluster}' does not exist.") + + monkeypatch.setattr(cli, "_sky_down_cluster", vanished) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: None) + assert cli.main(["down", "d2"]) == 0 + assert runs.load_run("d2")["state"] == "DOWN" + assert "not cloud-verifiable here; trusting sky" in capsys.readouterr().err + + +def test_down_no_longer_claims_success_on_other_sky_errors(monkeypatch, capsys): + """Any other down error used to print "teardown failed" and then + "run is down" with exit 0; now it is unconfirmed and the run is left + for a rerun.""" + runs.create_run("d3", make_args_dict("d3")) + runs.update_run("d3", pid=None, clusters=["d3-syncer"]) + def explode(cluster): - raise RuntimeError("cluster already gone") + raise RuntimeError("sky API server unreachable") monkeypatch.setattr(cli, "_sky_down_cluster", explode) - rc = cli.main(["down", "d2"]) - assert rc == 0 - assert runs.load_run("d2")["state"] == "DOWN" - assert "teardown failed" in capsys.readouterr().err + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: None) + assert cli.main(["down", "d3"]) == 1 + meta = runs.load_run("d3") + assert meta["state"] == runs.TEARDOWN_INCOMPLETE + assert meta["teardown_unconfirmed"] == ["d3-syncer"] + out, err = capsys.readouterr() + assert "run 'd3' is down" not in out + assert "d3-syncer: not confirmed down" in err + assert "NOT fully down; unconfirmed: d3-syncer" in err + + +def test_down_cloud_verifies_and_retries_until_the_cloud_is_empty(monkeypatch, capsys): + runs.create_run("d4", make_args_dict("d4")) + runs.update_run("d4", pid=None, clusters=["d4-syncer"]) + downed = [] + monkeypatch.setattr(cli, "_sky_down_cluster", downed.append) + seen = {"n": 0} + + def probe(): + seen["n"] += 1 + return ["i-zombie"] if seen["n"] < 3 else [] + + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: probe) + monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + assert cli.main(["down", "d4"]) == 0 + assert len(downed) == 3 # initial + one retry per live report + assert runs.load_run("d4")["state"] == "DOWN" + assert "d4-syncer: down" in capsys.readouterr().out + + +def test_down_fails_when_the_cloud_still_has_an_instance(monkeypatch, capsys): + runs.create_run("d5", make_args_dict("d5")) + runs.update_run("d5", pid=None, clusters=["d5-syncer"]) + monkeypatch.setattr(cli, "_sky_down_cluster", lambda c: None) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: (lambda: ["i-zombie"])) + monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + assert cli.main(["down", "d5"]) == 1 + assert runs.load_run("d5")["state"] == runs.TEARDOWN_INCOMPLETE + out, err = capsys.readouterr() + assert "i-zombie" in err # the surviving instance is named so it can be deleted + assert "run 'd5' is down" not in out + + +def test_down_fails_when_the_modal_app_is_still_running(monkeypatch, capsys): + runs.create_run("m2", make_args_dict("m2")) + runs.update_run("m2", pid=None, clusters=["m2-syncer", "m2-l0-modal"]) + monkeypatch.setattr(cli, "_sky_down_cluster", lambda c: None) + monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: (lambda: [])) + monkeypatch.setattr(cli, "_modal_stop_app", lambda run: None) + monkeypatch.setattr( + cli, "_modal_app_stopped", lambda run: (False, "Modal app yeto-m2: still running with 1 task(s)") + ) + assert cli.main(["down", "m2"]) == 1 + assert runs.load_run("m2")["teardown_unconfirmed"] == ["m2-l0-modal"] + assert "m2-l0-modal: not confirmed stopped" in capsys.readouterr().err + + +def test_modal_app_stopped_reads_state_and_tasks(monkeypatch): + from yeto import modal_runner + + class Ops: + def __init__(self, app): + self.app = app + + def app_status(self): + return {"yeto-ok": ("stopped", 0), "yeto-busy": ("running", 2)}.get(self.app) + + monkeypatch.setattr(modal_runner, "ModalOps", Ops) + assert cli._modal_app_stopped("ok") == (True, "Modal app yeto-ok: stopped, 0 tasks") + ok, detail = cli._modal_app_stopped("busy") + assert not ok and "still running with 2 task(s)" in detail + ok, detail = cli._modal_app_stopped("gone") + assert ok and "not listed by Modal" in detail + + class Broken(Ops): + def app_status(self): + raise RuntimeError("modal app list failed") + + monkeypatch.setattr(modal_runner, "ModalOps", Broken) + ok, detail = cli._modal_app_stopped("x") + assert not ok and "status unverified" in detail def test_down_unknown_run(capsys): diff --git a/tests/test_teardown_verify.py b/tests/test_teardown_verify.py index d14ee91b..765a9661 100644 --- a/tests/test_teardown_verify.py +++ b/tests/test_teardown_verify.py @@ -28,6 +28,30 @@ def test_trusts_down_when_no_probe(): assert sky.downs == 1 +def test_without_probe_only_a_clean_or_never_existed_down_counts(): + class Down: + def __init__(self, exc): + self.exc = exc + + def __call__(self): + raise self.exc + + assert terminate_and_verify( + None, "c", probe=None, down=Down(ValueError("Cluster 'c' does not exist.")), sleep_fn=_no_sleep + ) is True + assert terminate_and_verify( + None, "c", probe=None, down=Down(RuntimeError("API server unreachable")), sleep_fn=_no_sleep + ) is False + + +def test_down_hook_replaces_sky_down(): + calls = [] + assert terminate_and_verify( + None, "c", probe=lambda: [], down=lambda: calls.append(1), sleep_fn=_no_sleep + ) is True + assert calls == [1] + + def test_confirmed_gone_on_first_check(): sky = FakeSky() assert terminate_and_verify(sky, "c", probe=lambda: [], sleep_fn=_no_sleep) is True diff --git a/yeto/cli.py b/yeto/cli.py index 04254f5a..b1071b33 100644 --- a/yeto/cli.py +++ b/yeto/cli.py @@ -1690,10 +1690,13 @@ def cmd_logs(args) -> int: HEAD_DOWN_RETRY_S = 20.0 HEAD_DOWN_SCRIPT = """cd ~/sky_workdir && PY=$([ -x ~/miniconda3/bin/python3 ] && echo ~/miniconda3/bin/python3 || echo python3) && "$PY" - <<'PY' import sky +from yeto.launcher import terminate_and_verify for c in {clusters!r}: try: - sky.get(sky.down(c)) - print(f"[head] {{c}}: down", flush=True) + if terminate_and_verify(sky, c): + print(f"[head] {{c}}: down", flush=True) + else: + print(f"[head] {{c}}: still live at the cloud after down", flush=True) except Exception as e: # already gone, or never launched print(f"[head] {{c}}: {{e}}", flush=True) PY""" @@ -1746,6 +1749,48 @@ def _sky_down_cluster(cluster: str) -> None: sky.get(sky.down(cluster)) +def _cloud_probe(cluster: str): + """Cloud-level live-instance probe for a cluster this machine's sky + launched, or None when it cannot be built (patched out in tests).""" + from .launcher import _cloud_live_instances_probe + + return _cloud_live_instances_probe(cluster) + + +def _down_and_verify(cluster: str) -> bool: + """Down a cluster this machine's sky knows and confirm it at the cloud. + + The probe is captured before the down (down deletes the record it needs). + Without a probe we fall back to sky's own answer: a clean down or "does + not exist" counts, any other error does not.""" + from .launcher import terminate_and_verify + + probe = _cloud_probe(cluster) + if probe is None: + print(f"[yeto] {cluster}: not cloud-verifiable here; trusting sky", file=sys.stderr) + return terminate_and_verify( + None, cluster, probe=probe, down=lambda: _sky_down_cluster(cluster) + ) + + +def _modal_app_stopped(run_name: str) -> tuple[bool, str]: + """Confirm the run's Modal app is stopped with no running task (patched + out in tests). Returns (confirmed, detail).""" + from .modal_runner import ModalOps, modal_app_name + + app = modal_app_name(run_name) + try: + status = ModalOps(app).app_status() + except Exception as e: # noqa: BLE001 - listing failed, not the app + return False, f"Modal app {app}: status unverified ({e})" + if status is None: + return True, f"Modal app {app}: not listed by Modal (never deployed or already gone)" + state, tasks = status + if state == "stopped" and tasks == 0: + return True, f"Modal app {app}: stopped, 0 tasks" + return False, f"Modal app {app}: still {state} with {tasks} task(s)" + + def _modal_stop_app(run_name: str) -> None: """Stop the run's Modal app (patched out in tests).""" from .modal_runner import ModalOps, modal_app_name @@ -1753,8 +1798,11 @@ def _modal_stop_app(run_name: str) -> None: try: ModalOps(modal_app_name(run_name)).stop_app() print(f"[yeto] Modal app {modal_app_name(run_name)}: stopped") - except Exception as e: # best-effort - print(f"[yeto] Modal app stop failed: {e}", file=sys.stderr) + except Exception as e: # the status check below decides; this is advisory + if "already stopped" in str(e): + print(f"[yeto] Modal app {modal_app_name(run_name)}: already stopped") + else: + print(f"[yeto] Modal app stop failed: {e}", file=sys.stderr) def _signal_worker(pid: int, sig: int) -> None: @@ -1791,64 +1839,81 @@ def cmd_down(args) -> int: print("[yeto] worker is not running") clusters = meta.get("clusters") or [] - if clusters: - print(f"[yeto] tearing down {len(clusters)} cluster(s): {', '.join(clusters)}") - from .modal_runner import is_modal_island - - modal_names = [c for c in clusters if is_modal_island(c)] - if modal_names: - # Modal islands are function calls in the run's app, not sky - # clusters: stopping the app ends every one of them at once. - _modal_stop_app(name) - - head_cluster = meta.get("head_cluster") if meta.get("controller") == "head" else None - on_head = [c for c in clusters if c != head_cluster and c not in modal_names] - if head_cluster and on_head: - # Only the head's sky knows these clusters, and the head may itself - # be mid-teardown (sky then answers 500), so retry; and never delete - # the head while a learner is unconfirmed — that orphans it. - pending, job = list(on_head), meta.get("head_job_id") - for attempt in range(HEAD_DOWN_ATTEMPTS): - try: - pending = _head_down_learners(head_cluster, job, pending) - except Exception as e: # noqa: BLE001 - e.g. head unreachable - print(f"[yeto] {head_cluster}: head-side teardown failed: {e}", file=sys.stderr) - job = None # the controller is cancelled after the first attempt - if not pending: - break - if attempt + 1 < HEAD_DOWN_ATTEMPTS: - time.sleep(HEAD_DOWN_RETRY_S) - if pending: - print( - f"[yeto] NOT tearing down {head_cluster}: learner cluster(s) not confirmed down " - f"from it: {', '.join(pending)}. The head is the only machine whose sky knows " - f"them; rerun `yeto down {name}`, or delete them in the cloud console.", - file=sys.stderr, - ) - return 1 - print(f"[yeto] learner clusters torn down from {head_cluster}: {', '.join(on_head)}") - clusters = [c for c in clusters if c not in on_head] - - def _down_one(cluster: str) -> None: - if cluster in modal_names: - print(f"[yeto] {cluster}: stopped with the Modal app") - return - try: - _sky_down_cluster(cluster) - print(f"[yeto] {cluster}: down") - except Exception as e: # best-effort; the cluster may be gone - print(f"[yeto] {cluster}: teardown failed: {e}", file=sys.stderr) - - threads = [ - threading.Thread(target=_down_one, args=(c,), daemon=True) for c in clusters - ] - for t in threads: - t.start() - for t in threads: - t.join() - else: + if not clusters: print("[yeto] no clusters recorded for this run") + runs.update_run(name, state=runs.DOWN, finished_at=meta.get("finished_at") or time.time()) + print(f"[yeto] run '{name}' is down") + return 0 + print(f"[yeto] tearing down {len(clusters)} cluster(s): {', '.join(clusters)}") + from .modal_runner import is_modal_island + + # Three kinds of cluster, three routes: Modal islands are function calls + # in the run's app (stop the app, then check it); learners of a head run + # exist only in the head's sky (tear them down FROM the head, before the + # head); everything else this machine's sky launched itself. + head_cluster = meta.get("head_cluster") if meta.get("controller") == "head" else None + modal_names = [c for c in clusters if is_modal_island(c)] + on_head = [c for c in clusters if head_cluster and c != head_cluster and c not in modal_names] + local_names = [c for c in clusters if c not in modal_names and c not in on_head] + unconfirmed: list[str] = [] + + if modal_names: + _modal_stop_app(name) + ok, detail = _modal_app_stopped(name) + for c in modal_names: + if ok: + print(f"[yeto] {c}: stopped with the Modal app ({detail})") + else: + print(f"[yeto] {c}: not confirmed stopped ({detail})", file=sys.stderr) + if not ok: + unconfirmed.extend(modal_names) + + if on_head: + # Only the head's sky knows these clusters, and the head may itself + # be mid-teardown (sky then answers 500), so retry; and never delete + # the head while a learner is unconfirmed — that orphans it. + pending, job = list(on_head), meta.get("head_job_id") + for attempt in range(HEAD_DOWN_ATTEMPTS): + try: + pending = _head_down_learners(head_cluster, job, pending) + except Exception as e: # noqa: BLE001 - e.g. head unreachable + print(f"[yeto] {head_cluster}: head-side teardown failed: {e}", file=sys.stderr) + job = None # the controller is cancelled after the first attempt + if not pending: + break + if attempt + 1 < HEAD_DOWN_ATTEMPTS: + time.sleep(HEAD_DOWN_RETRY_S) + if pending: + print( + f"[yeto] NOT tearing down {head_cluster}: learner cluster(s) not confirmed down " + f"from it: {', '.join(pending)}. The head is the only machine whose sky knows " + f"them; rerun `yeto down {name}`, or delete them in the cloud console.", + file=sys.stderr, + ) + return _teardown_incomplete(name, meta, pending + unconfirmed) + print(f"[yeto] learner clusters torn down from {head_cluster}: {', '.join(on_head)}") + + results: dict[str, bool] = {} + + def _down_one(cluster: str) -> None: + results[cluster] = _down_and_verify(cluster) + if results[cluster]: + print(f"[yeto] {cluster}: down") + else: + print(f"[yeto] {cluster}: not confirmed down", file=sys.stderr) + + threads = [ + threading.Thread(target=_down_one, args=(c,), daemon=True) for c in local_names + ] + for t in threads: + t.start() + for t in threads: + t.join() + unconfirmed.extend(c for c in local_names if not results.get(c)) + + if unconfirmed: + return _teardown_incomplete(name, meta, unconfirmed) runs.update_run( name, state=runs.DOWN, @@ -1858,6 +1923,24 @@ def _down_one(cluster: str) -> None: return 0 +def _teardown_incomplete(name: str, meta: dict, unconfirmed: list[str]) -> int: + """Record and report a `yeto down` that could not confirm every cluster + gone. The run is NOT marked down, so `yeto status` shows it and a rerun + of `yeto down` picks up where this one stopped.""" + runs.update_run( + name, + state=runs.TEARDOWN_INCOMPLETE, + teardown_unconfirmed=sorted(set(unconfirmed)), + finished_at=meta.get("finished_at") or time.time(), + ) + print( + f"[yeto] run '{name}' is NOT fully down; unconfirmed: {', '.join(sorted(set(unconfirmed)))}. " + f"Rerun `yeto down {name}`, or check the cloud console for instances named after them.", + file=sys.stderr, + ) + return 1 + + # --------------------------------------------------------------------------- diff --git a/yeto/launcher.py b/yeto/launcher.py index f9cea174..f723d949 100644 --- a/yeto/launcher.py +++ b/yeto/launcher.py @@ -3211,7 +3211,9 @@ def live(): return None -def terminate_and_verify(sky, cluster, *, probe="auto", attempts=4, sleep_fn=time.sleep) -> bool: +def terminate_and_verify( + sky, cluster, *, probe="auto", attempts=4, sleep_fn=time.sleep, down=None +) -> bool: """sky.down a cluster and CONFIRM at the cloud level that no instance survives, retrying the down while the cloud still reports live ones. @@ -3220,20 +3222,30 @@ def terminate_and_verify(sky, cluster, *, probe="auto", attempts=4, sleep_fn=tim sky is the ONLY thing that can reach the learner clusters, so a silent orphan is unrecoverable once the head is gone — hence verify here, before the head relinquishes control. Returns True iff the cluster is confirmed - gone, or can't be cloud-verified (then we trust sky.down). + gone. When the cloud can't be queried we fall back to sky.down's own + result: a clean down, or "does not exist" (sky never had it), counts; + any other down error does not — that is exactly the case that used to + print "teardown failed" and then claim the run was down. + + `down` overrides the sky.down call (the CLI routes it through its own + patchable hook); `probe` is captured before the first down because + sky.down deletes the record the probe is built from. """ if probe == "auto": probe = _cloud_live_instances_probe(cluster) + down = down or (lambda: sky.get(sky.down(cluster))) def _down(): try: - sky.get(sky.down(cluster)) + down() + return None except Exception as e: print(f"[launcher] sky.down({cluster}) error: {e}", file=sys.stderr) + return e - _down() + err = _down() if probe is None: - return True + return err is None or "does not exist" in str(err) for i in range(attempts): try: live = probe() diff --git a/yeto/modal_runner.py b/yeto/modal_runner.py index 595789d5..3aa8ad88 100644 --- a/yeto/modal_runner.py +++ b/yeto/modal_runner.py @@ -360,6 +360,32 @@ def cancel(self, call_id: str) -> None: modal = self._modal() modal.FunctionCall.from_id(call_id).cancel(terminate_containers=True) + def app_status(self) -> tuple[str, int] | None: + """(state, running tasks) of this run's app from `modal app list`, + or None when Modal lists no app by that name. Raises when the + listing itself fails, so callers can say "unverified" rather than + "stopped".""" + proc = subprocess.run( + [sys.executable, "-m", "modal", "app", "list", "--json"], + capture_output=True, + text=True, + ) + if proc.returncode != 0: + raise RuntimeError( + f"modal app list failed: {(proc.stdout + proc.stderr).strip() or proc.returncode}" + ) + import json as _json + + for app in _json.loads(proc.stdout or "[]"): + if app.get("Description") == self.app_name or app.get("Name") == self.app_name: + state = str(app.get("State", "")).lower() + try: + tasks = int(app.get("Tasks") or 0) + except (TypeError, ValueError): + tasks = 0 + return state, tasks + return None + def stop_app(self) -> None: """Stop every function of this run's app (used by `yeto down`). diff --git a/yeto/runs.py b/yeto/runs.py index c957886b..605d2906 100644 --- a/yeto/runs.py +++ b/yeto/runs.py @@ -32,6 +32,8 @@ SUCCEEDED = "SUCCEEDED" FAILED = "FAILED" DOWN = "DOWN" +# `yeto down` ran but could not confirm every cluster gone; rerun it. +TEARDOWN_INCOMPLETE = "TEARDOWN_INCOMPLETE" # Head-controller-mode runs additionally record `controller` ("head"), # `head_cluster` and `head_job_id` via update_run. Local-mode entries From a183ba7d5b782a0ba65f9d0c3f69748f471a583d Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:17:49 +0000 Subject: [PATCH 3/8] cli: make the teardown verifier's sleep patchable --- tests/test_head_mode.py | 2 +- tests/test_runs_cli.py | 4 ++-- yeto/cli.py | 9 ++++++++- 3 files changed, 11 insertions(+), 4 deletions(-) diff --git a/tests/test_head_mode.py b/tests/test_head_mode.py index 130416de..d0893450 100644 --- a/tests/test_head_mode.py +++ b/tests/test_head_mode.py @@ -546,7 +546,7 @@ def test_down_head_run_keeps_the_head_when_the_cloud_still_has_it(monkeypatch, c monkeypatch.setattr(cli, "_sky_down_cluster", lambda c: None) monkeypatch.setattr(cli, "_head_down_learners", lambda head, job, cs: []) monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: (lambda: ["i-head-zombie"])) - monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + monkeypatch.setattr(cli, "DOWN_VERIFY_SLEEP", lambda s: None) assert cli.main(["down", "hc"]) == 1 assert runs.load_run("hc")["state"] == runs.TEARDOWN_INCOMPLETE assert "i-head-zombie" in capsys.readouterr().err diff --git a/tests/test_runs_cli.py b/tests/test_runs_cli.py index d58a9eb0..0c6584fa 100644 --- a/tests/test_runs_cli.py +++ b/tests/test_runs_cli.py @@ -338,7 +338,7 @@ def probe(): return ["i-zombie"] if seen["n"] < 3 else [] monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: probe) - monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + monkeypatch.setattr(cli, "DOWN_VERIFY_SLEEP", lambda s: None) assert cli.main(["down", "d4"]) == 0 assert len(downed) == 3 # initial + one retry per live report assert runs.load_run("d4")["state"] == "DOWN" @@ -350,7 +350,7 @@ def test_down_fails_when_the_cloud_still_has_an_instance(monkeypatch, capsys): runs.update_run("d5", pid=None, clusters=["d5-syncer"]) monkeypatch.setattr(cli, "_sky_down_cluster", lambda c: None) monkeypatch.setattr(cli, "_cloud_probe", lambda cluster: (lambda: ["i-zombie"])) - monkeypatch.setattr("yeto.launcher.time.sleep", lambda s: None) + monkeypatch.setattr(cli, "DOWN_VERIFY_SLEEP", lambda s: None) assert cli.main(["down", "d5"]) == 1 assert runs.load_run("d5")["state"] == runs.TEARDOWN_INCOMPLETE out, err = capsys.readouterr() diff --git a/yeto/cli.py b/yeto/cli.py index b1071b33..c0a1c49c 100644 --- a/yeto/cli.py +++ b/yeto/cli.py @@ -1757,6 +1757,9 @@ def _cloud_probe(cluster: str): return _cloud_live_instances_probe(cluster) +DOWN_VERIFY_SLEEP = time.sleep # patched out in tests + + def _down_and_verify(cluster: str) -> bool: """Down a cluster this machine's sky knows and confirm it at the cloud. @@ -1769,7 +1772,11 @@ def _down_and_verify(cluster: str) -> bool: if probe is None: print(f"[yeto] {cluster}: not cloud-verifiable here; trusting sky", file=sys.stderr) return terminate_and_verify( - None, cluster, probe=probe, down=lambda: _sky_down_cluster(cluster) + None, + cluster, + probe=probe, + down=lambda: _sky_down_cluster(cluster), + sleep_fn=DOWN_VERIFY_SLEEP, ) From 82db165a7fe8bf6aa6843e5a0a05e7cdcfe82d62 Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:18:31 +0000 Subject: [PATCH 4/8] docs: describe the verified teardown order --- docs/CLOUDS.md | 43 +++++++++++++++++-- .../fix-head-mode-island-teardown/tasks.md | 18 ++++---- 2 files changed, 48 insertions(+), 13 deletions(-) diff --git a/docs/CLOUDS.md b/docs/CLOUDS.md index ab99af2b..971dc8f6 100644 --- a/docs/CLOUDS.md +++ b/docs/CLOUDS.md @@ -74,7 +74,8 @@ from Modal's cluster info. Facts to keep in mind: - Object-store data (`s3://`, `gs://`) cannot be mounted; use an HF dataset id or a local path. - `yeto down ` stops the run's Modal app (`yeto-`); every - island's containers end with it. + island's containers end with it, and the app must then list as + `stopped` with 0 tasks or the run is reported not fully down. - An all-Modal SFT fleet's model is not fetchable over ssh; recover it from the syncer checkpoint with `yeto-export`, or keep at least one sky island in the fleet. @@ -229,9 +230,9 @@ Three things found while running this, none fixed in this change: stopped.** After a clean run the head has already stopped the Modal app, so `yeto down` prints `Modal app stop failed: ... App is already stopped.` and then proceeds - normally (exit code 0). The `--yes` fix for `stop_app` surfaces - failures but does not treat an already-stopped app as success, so the - message is misleading noise on every successful mixed or Modal run. + normally (exit code 0). Fixed with the head-run teardown change: an + already-stopped app now prints as such, and the app's state is checked + afterwards either way. **Modal RL island, 8xH100, 2026-09-23 (task 8.4).** Head controller (syncer) on Nebius eu-north1; one Miles RL island as a Modal function on @@ -408,3 +409,37 @@ the learner refuses to start when its source SHA256 differs from the fleet's. The launch log still prints only an MLX join command, not a Modal one; building the run script by hand is the gap left for manual Modal joins. +## Tearing a run down + +`yeto down ` only says `run '' is down` (exit 0) once every +cluster of the run is confirmed gone; anything less exits 1, leaves the run +in state `TEARDOWN_INCOMPLETE` with the unconfirmed clusters recorded, and +a rerun of `yeto down` continues from there. The order matters: + +1. **Learners of a head run are torn down from the head.** Only the head's + sky knows them: this machine's sky launched the head, the head's sky + launched the learners, and a local `sky down ` just answers + "does not exist". `yeto down` cancels the controller job on the head + (so it cannot relaunch what is being removed), then runs the downs over + ssh on the head, where each learner is also checked at the cloud before + it counts as down. Three attempts, 20 s apart, because the head's own + sky answers 500 while it is busy. +2. **The head is deleted only after every learner is confirmed.** If any + learner is still unconfirmed the head is kept, the command exits 1 and + names the learners: rerun `yeto down`, or delete them in the cloud + console. Deleting the head first is exactly how H100s were orphaned on + 2026-09-23 and twice on 2026-09-24 (`yeto-gh1`, `yeto-gh2`). +3. **Modal islands end with the run's app**, then the app must list as + `stopped` with 0 tasks. +4. **Everything this machine's sky launched (the head, local-mode + learners) is checked at the cloud after the down**, using sky's own + per-cloud instance query; a surviving instance is retried and, if it + outlives the retries, printed by id so it can be deleted by hand. A + cloud that cannot be queried is reported as "not cloud-verifiable; + trusting sky", and then only a clean down or "does not exist" counts. + +If the head is already gone (deleted by hand, or by an older `yeto down`), +step 1 cannot run: the command exits 1 listing the learners, and the only +way to find them is the cloud's own listing, e.g. +`nebius compute instance list` — look for instances named after the +learner cluster. diff --git a/openspec/changes/fix-head-mode-island-teardown/tasks.md b/openspec/changes/fix-head-mode-island-teardown/tasks.md index 5d3010c3..cc4dda1f 100644 --- a/openspec/changes/fix-head-mode-island-teardown/tasks.md +++ b/openspec/changes/fix-head-mode-island-teardown/tasks.md @@ -2,23 +2,23 @@ ## 1. 把 prB 落到当前栈上 -- [ ] 1.1 在 pr8 之上合并 `prB/head-side-teardown`,按 design D1 解决 `yeto/cli.py` 的冲突:先把 `clusters` 三分类(Modal / 经 head 的 learner / 本机直接 down),再各自分派。验证:`tests/test_head_mode.py` 中 prB 带来的两个测试与 pr8 的 Modal 停止测试同时通过 -- [ ] 1.2 新增测试:head 模式 run 含一个 Modal 岛和一个 sky 岛时,Modal 岛走 app 停止、sky 岛经 head 拆除,且两者都不经过本机 `sky down`。验证:测试通过,且断言本机 `sky.down` 从未被以 learner 名调用 +- [x] 1.1 在 pr8 之上合并 `prB/head-side-teardown`,按 design D1 解决 `yeto/cli.py` 的冲突:先把 `clusters` 三分类(Modal / 经 head 的 learner / 本机直接 down),再各自分派。验证:`tests/test_head_mode.py` 中 prB 带来的两个测试与 pr8 的 Modal 停止测试同时通过 +- [x] 1.2 新增测试:head 模式 run 含一个 Modal 岛和一个 sky 岛时,Modal 岛走 app 停止、sky 岛经 head 拆除,且两者都不经过本机 `sky down`。验证:测试通过,且断言本机 `sky.down` 从未被以 learner 名调用 ## 2. 本机失败不再静默 -- [ ] 2.1 改 `_down_one`:本机 `sky down` 异常记为"未确认",仅当该 cluster 已在 head 侧确认时忽略。验证:新增测试——head 模式下 learner 未在 head 确认、本机又报 "does not exist" 时,命令非零退出、head 未被删除、错误信息列出该 learner -- [ ] 2.2 确认 `--controller local` 行为不变。验证:既有 local 模式的 `yeto down` 测试全部通过,无需改动 +- [x] 2.1 改 `_down_one`:本机 `sky down` 异常记为"未确认",仅当该 cluster 已在 head 侧确认时忽略。验证:新增测试——head 模式下 learner 未在 head 确认、本机又报 "does not exist" 时,命令非零退出、head 未被删除、错误信息列出该 learner +- [x] 2.2 确认 `--controller local` 行为不变。验证:既有 local 模式的 `yeto down` 测试全部通过,无需改动(一处例外按 spec 改了:`test_down_survives_sky_errors` 原本把任意 down 异常当成功;现在 "does not exist" 仍算已消失、其他异常为未确认并非零退出,见 `test_down_no_longer_claims_success_on_other_sky_errors`) ## 3. 云端核对 -- [ ] 3.1 head cluster 的拆除改用 `terminate_and_verify`(含云端探针)替换裸 `sky.down`。验证:测试用假探针模拟"仍有实例存活"→ 重试后仍存活 → 非零退出并打印实例 id;"探针为 None"→ 打印未核验、零退出 -- [ ] 3.2 把 prB 的 `HEAD_DOWN_SCRIPT` 改为在 head 上调用 `terminate_and_verify`,确认行反映其返回值。验证:`_unconfirmed_head_downs` 的解析测试覆盖"down 但云端仍存活"输出为未确认 -- [ ] 3.3 Modal learner 在 app 停止后核对 app 状态为 stopped 且 tasks 为 0,否则非零退出。验证:测试用假的 app 状态覆盖两种结果 -- [ ] 3.4 只有全部确认后才打印 `run '' is down` 并返回 0;部分失败时 `runs.update_run` 记录为 teardown 未完成,`yeto status` 能显示。验证:测试断言部分失败时不出现成功信息且状态可见 +- [x] 3.1 head cluster 的拆除改用 `terminate_and_verify`(含云端探针)替换裸 `sky.down`。验证:测试用假探针模拟"仍有实例存活"→ 重试后仍存活 → 非零退出并打印实例 id;"探针为 None"→ 打印未核验、零退出 +- [x] 3.2 把 prB 的 `HEAD_DOWN_SCRIPT` 改为在 head 上调用 `terminate_and_verify`,确认行反映其返回值。验证:`_unconfirmed_head_downs` 的解析测试覆盖"down 但云端仍存活"输出为未确认 +- [x] 3.3 Modal learner 在 app 停止后核对 app 状态为 stopped 且 tasks 为 0,否则非零退出。验证:测试用假的 app 状态覆盖两种结果 +- [x] 3.4 只有全部确认后才打印 `run '' is down` 并返回 0;部分失败时 `runs.update_run` 记录为 teardown 未完成,`yeto status` 能显示。验证:测试断言部分失败时不出现成功信息且状态可见 ## 4. 文档与真机确认 -- [ ] 4.1 更新 `docs/CLOUDS.md`:把"`yeto down` 后手工看 `nebius compute instance list`"改为说明新行为与非零退出的含义;在 live-run-failures 第 36 条下记录处理。验证:文档改动与实现一致 +- [x] 4.1 更新 `docs/CLOUDS.md`:把"`yeto down` 后手工看 `nebius compute instance list`"改为说明新行为与非零退出的含义;在 live-run-failures 第 36 条下记录处理。验证:文档改动与实现一致 - [ ] 4.2 真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** - [ ] 4.3 真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 From 9cf360f89939562aa4bd64db71e0eba263f68136 Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:28:06 +0000 Subject: [PATCH 5/8] launcher: fix the cloud probe's StatusVersion comparison sky's StatusVersion enum defines only __ge__, so the '<' in _cloud_live_instances_probe raised and every cluster fell back to 'trusting sky.down' with a one-line warning; cloud verification had never actually run. Verified read-only against a stopped Nebius head: the probe now returns its instance id. --- tests/test_teardown_verify.py | 56 +++++++++++++++++++++++++++++++++++ yeto/launcher.py | 6 +++- 2 files changed, 61 insertions(+), 1 deletion(-) diff --git a/tests/test_teardown_verify.py b/tests/test_teardown_verify.py index 765a9661..326665a7 100644 --- a/tests/test_teardown_verify.py +++ b/tests/test_teardown_verify.py @@ -88,3 +88,59 @@ def probe(): raise RuntimeError("cloud API down") assert terminate_and_verify(sky, "c", probe=probe, sleep_fn=_no_sleep) is True + + +def test_cloud_probe_builds_against_sky_status_version_enum(monkeypatch): + """sky's StatusVersion defines only ``__ge__``; a ``<`` comparison raised + and the probe fell back to "trust sky.down" for every cloud, which is how + a stopped head and orphaned learners went unnoticed.""" + import enum + import sys + import types + + from yeto import launcher + + class StatusVersion(enum.Enum): + CLOUD_CLI = 1 + SKYPILOT = 2 + + def __ge__(self, other): + return self.value >= other.value + + class Cloud: + STATUS_VERSION = StatusVersion.SKYPILOT + + def __repr__(self): + return "Nebius" + + handle = types.SimpleNamespace( + launched_resources=types.SimpleNamespace(cloud=Cloud()), + cluster_name="c", + cluster_name_on_cloud="c-abc", + cluster_yaml="/tmp/c.yaml", + ) + queries = [] + + def query_instances(cloud_name, name, name_on_cloud, provider_config, non_terminated_only): + queries.append((cloud_name, name, name_on_cloud, provider_config, non_terminated_only)) + return {"i-live": ("RUNNING", None), "i-gone": (None, None)} + + sky_pkg = types.ModuleType("sky") + sky_pkg.clouds = types.SimpleNamespace(StatusVersion=StatusVersion) + sky_pkg.global_user_state = types.SimpleNamespace( + get_cluster_from_name=lambda cluster: {"handle": handle}, + get_cluster_yaml_dict=lambda path: {"provider": {"region": "eu-north1"}}, + ) + sky_pkg.provision = types.SimpleNamespace(query_instances=query_instances) + monkeypatch.setitem(sys.modules, "sky", sky_pkg) + monkeypatch.setitem(sys.modules, "sky.clouds", sky_pkg.clouds) + monkeypatch.setitem(sys.modules, "sky.global_user_state", sky_pkg.global_user_state) + monkeypatch.setitem(sys.modules, "sky.provision", sky_pkg.provision) + + probe = launcher._cloud_live_instances_probe("c") + assert probe is not None + assert probe() == ["i-live"] + assert queries == [("Nebius", "c", "c-abc", {"region": "eu-north1"}, True)] + + Cloud.STATUS_VERSION = StatusVersion.CLOUD_CLI + assert launcher._cloud_live_instances_probe("c") is None diff --git a/yeto/launcher.py b/yeto/launcher.py index f723d949..e8e151f2 100644 --- a/yeto/launcher.py +++ b/yeto/launcher.py @@ -3189,7 +3189,11 @@ def _cloud_live_instances_probe(cluster: str): return None handle = record["handle"] cloud = handle.launched_resources.cloud - if cloud is None or cloud.STATUS_VERSION < clouds.StatusVersion.SKYPILOT: + # sky's StatusVersion enum defines only ``>=``; ``<`` raises, which + # used to trip the except below and silently disable verification + # for every cloud ("cannot set up cloud verification ... '<' not + # supported"). + if cloud is None or not (cloud.STATUS_VERSION >= clouds.StatusVersion.SKYPILOT): return None cloud_name = repr(cloud) name = handle.cluster_name From af77cbd23ef83b7e1bf06227d7fc11547757333f Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Thu, 24 Sep 2026 14:35:45 +0000 Subject: [PATCH 6/8] modal: read app status from modal app list's real JSON keys; record the td1/td2 teardown runs --- docs/CLOUDS.md | 10 ++++++++ .../fix-head-mode-island-teardown/tasks.md | 4 ++-- tests/test_runs_cli.py | 24 +++++++++++++++++++ yeto/modal_runner.py | 8 ++++--- 4 files changed, 41 insertions(+), 5 deletions(-) diff --git a/docs/CLOUDS.md b/docs/CLOUDS.md index 971dc8f6..56514f4e 100644 --- a/docs/CLOUDS.md +++ b/docs/CLOUDS.md @@ -438,6 +438,16 @@ a rerun of `yeto down` continues from there. The order matters: cloud that cannot be queried is reported as "not cloud-verifiable; trusting sky", and then only a clean down or "does not exist" counts. +Verified 2026-09-24 on `yeto-td2` (Nebius head, one Modal `1xh100` SFT +island, torn down while training at outer step 3): the app stopped and +listed as `stopped, 0 tasks`, the head was downed and confirmed at the +cloud, exit 0, nothing left on Nebius or Modal. `yeto-td1` (whose Nebius +learner never provisioned: the tenant's public-IPv4 quota of 3 was full) +exercised the head-side path: the head reported the learner as never +existing, then the head itself was downed. The head-side teardown of a +provisioned Nebius learner, and the "head already gone" failure path, are +still to be run once the IPv4 quota has room for head + learner. + If the head is already gone (deleted by hand, or by an older `yeto down`), step 1 cannot run: the command exits 1 listing the learners, and the only way to find them is the cloud's own listing, e.g. diff --git a/openspec/changes/fix-head-mode-island-teardown/tasks.md b/openspec/changes/fix-head-mode-island-teardown/tasks.md index cc4dda1f..76df0305 100644 --- a/openspec/changes/fix-head-mode-island-teardown/tasks.md +++ b/openspec/changes/fix-head-mode-island-teardown/tasks.md @@ -20,5 +20,5 @@ ## 4. 文档与真机确认 - [x] 4.1 更新 `docs/CLOUDS.md`:把"`yeto down` 后手工看 `nebius compute instance list`"改为说明新行为与非零退出的含义;在 live-run-failures 第 36 条下记录处理。验证:文档改动与实现一致 -- [ ] 4.2 真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** -- [ ] 4.3 真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 +- [ ] 4.2 (2026-09-24 部分完成,见 CLOUDS.md:`yeto-td2` 用 Modal 岛做了运行中 `yeto down`,Modal app 确认 + head 云端核验通过;`yeto-td1` 的 Nebius 岛因租户公网 IPv4 配额(上限 3,被另一 run 占用)没开出来,只验证了经 head 的“learner 不存在”确认路径。Nebius learner 的经 head 拆除待配额空出后补做)真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** +- [ ] 4.3 (同上,待 IPv4 配额空出)真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 diff --git a/tests/test_runs_cli.py b/tests/test_runs_cli.py index 0c6584fa..72c3dead 100644 --- a/tests/test_runs_cli.py +++ b/tests/test_runs_cli.py @@ -642,3 +642,27 @@ def test_logs_follow_ends_when_worker_dead(capsys): def test_logs_unknown_run(capsys): assert cli.main(["logs", "ghost"]) == 1 assert "unknown run" in capsys.readouterr().err + + +def test_modal_ops_app_status_parses_modal_app_list_json(monkeypatch): + import json + import subprocess + from types import SimpleNamespace + + from yeto import modal_runner + + rows = [ + {"app_id": "ap-1", "description": "yeto-other", "state": "running", "tasks": "2"}, + {"app_id": "ap-2", "description": "yeto-td2", "state": "stopped", "tasks": "0"}, + ] + monkeypatch.setattr( + subprocess, "run", lambda *a, **k: SimpleNamespace(returncode=0, stdout=json.dumps(rows), stderr="") + ) + assert modal_runner.ModalOps("yeto-td2").app_status() == ("stopped", 0) + assert modal_runner.ModalOps("yeto-other").app_status() == ("running", 2) + assert modal_runner.ModalOps("yeto-none").app_status() is None + monkeypatch.setattr( + subprocess, "run", lambda *a, **k: SimpleNamespace(returncode=1, stdout="", stderr="token expired") + ) + with pytest.raises(RuntimeError, match="modal app list failed"): + modal_runner.ModalOps("yeto-td2").app_status() diff --git a/yeto/modal_runner.py b/yeto/modal_runner.py index 3aa8ad88..eadf3ed2 100644 --- a/yeto/modal_runner.py +++ b/yeto/modal_runner.py @@ -376,11 +376,13 @@ def app_status(self) -> tuple[str, int] | None: ) import json as _json + # `modal app list --json` rows: app_id, description, state, tasks, + # created_at, stopped_at (tasks is a string). for app in _json.loads(proc.stdout or "[]"): - if app.get("Description") == self.app_name or app.get("Name") == self.app_name: - state = str(app.get("State", "")).lower() + if app.get("description") == self.app_name: + state = str(app.get("state", "")).lower() try: - tasks = int(app.get("Tasks") or 0) + tasks = int(app.get("tasks") or 0) except (TypeError, ValueError): tasks = 0 return state, tasks From eaa9e42857e134b249540b32c72cecb79f3ae7f6 Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Fri, 25 Sep 2026 04:29:04 +0000 Subject: [PATCH 7/8] docs: record the td3/td4 Nebius teardown verification --- docs/CLOUDS.md | 20 ++++++++++++++++--- .../fix-head-mode-island-teardown/tasks.md | 4 ++-- 2 files changed, 19 insertions(+), 5 deletions(-) diff --git a/docs/CLOUDS.md b/docs/CLOUDS.md index 56514f4e..c0ddeb0f 100644 --- a/docs/CLOUDS.md +++ b/docs/CLOUDS.md @@ -444,9 +444,23 @@ listed as `stopped, 0 tasks`, the head was downed and confirmed at the cloud, exit 0, nothing left on Nebius or Modal. `yeto-td1` (whose Nebius learner never provisioned: the tenant's public-IPv4 quota of 3 was full) exercised the head-side path: the head reported the learner as never -existing, then the head itself was downed. The head-side teardown of a -provisioned Nebius learner, and the "head already gone" failure path, are -still to be run once the IPv4 quota has room for head + learner. +existing, then the head itself was downed. + +Verified 2026-09-25 with real Nebius learners once the IPv4 quota had room: + +- `yeto-td3` (Nebius head, Nebius `1xh100` SFT island, torn down at outer + step 3): the head confirmed `[head] yeto-td3-l0-eu-north1: down`, then + the head was downed; the cloud probe saw the head instance still live + three times while Nebius was deleting it and retried until it was gone, + exit 0, `nebius compute instance list` empty for the prefix. Before the + StatusVersion fix this probe had never run. +- `yeto-td4`, the failure case: the head was deleted by hand while the + island trained. `yeto down` cancelled nothing (head STOPPED), the ssh to + the head failed three times, and the command exited 1 with + `NOT tearing down yeto-td4-head: learner cluster(s) not confirmed down + from it: yeto-td4-l0-eu-north1`, run state `TEARDOWN_INCOMPLETE`. The + island was then deleted from the cloud by hand, which is the documented + recovery. If the head is already gone (deleted by hand, or by an older `yeto down`), step 1 cannot run: the command exits 1 listing the learners, and the only diff --git a/openspec/changes/fix-head-mode-island-teardown/tasks.md b/openspec/changes/fix-head-mode-island-teardown/tasks.md index 76df0305..59977a7d 100644 --- a/openspec/changes/fix-head-mode-island-teardown/tasks.md +++ b/openspec/changes/fix-head-mode-island-teardown/tasks.md @@ -20,5 +20,5 @@ ## 4. 文档与真机确认 - [x] 4.1 更新 `docs/CLOUDS.md`:把"`yeto down` 后手工看 `nebius compute instance list`"改为说明新行为与非零退出的含义;在 live-run-failures 第 36 条下记录处理。验证:文档改动与实现一致 -- [ ] 4.2 (2026-09-24 部分完成,见 CLOUDS.md:`yeto-td2` 用 Modal 岛做了运行中 `yeto down`,Modal app 确认 + head 云端核验通过;`yeto-td1` 的 Nebius 岛因租户公网 IPv4 配额(上限 3,被另一 run 占用)没开出来,只验证了经 head 的“learner 不存在”确认路径。Nebius learner 的经 head 拆除待配额空出后补做)真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** -- [ ] 4.3 (同上,待 IPv4 配额空出)真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 +- [x] 4.2 (2026-09-25 `yeto-td3` 完成:Nebius 岛经 head 确认 down、head 云端核验重试至空、退出 0、无残留。此前 2026-09-24 部分完成,见 CLOUDS.md:`yeto-td2` 用 Modal 岛做了运行中 `yeto down`,Modal app 确认 + head 云端核验通过;`yeto-td1` 的 Nebius 岛因租户公网 IPv4 配额(上限 3,被另一 run 占用)没开出来,只验证了经 head 的“learner 不存在”确认路径。Nebius learner 的经 head 拆除待配额空出后补做)真机:在 Nebius 开一个 head + Nebius 1 卡 SFT 岛(8.1 的配置,避开第 37 条的 docker 岛问题),岛运行中执行 `yeto down`,确认输出含每个 learner 的确认行、head 最后删除,`nebius compute instance list` 与 `sky status` 无该前缀残留。验证:结果记入 `docs/CLOUDS.md`;**不要动不属于本次运行的实例** +- [x] 4.3 (2026-09-25 `yeto-td4` 完成:手工删 head 后 `yeto down` 非零退出、列出 `yeto-td4-l0-eu-north1`、run 记 `TEARDOWN_INCOMPLETE`;岛随后手工删除。)真机反例:先手工删掉 head 再执行 `yeto down`,确认命令非零退出并列出未确认的 learner,随后手工删除该岛。验证:记录输出;同样只动本次运行的实例 From 512620e1b853f3761e1c92b6e222a6dd31873335 Mon Sep 17 00:00:00 2001 From: MichaelChung Date: Fri, 25 Sep 2026 04:32:39 +0000 Subject: [PATCH 8/8] openspec: archive fix-head-mode-island-teardown and publish the head-run-teardown spec --- .../.openspec.yaml | 0 .../design.md | 0 .../proposal.md | 0 .../specs/head-run-teardown/spec.md | 0 .../tasks.md | 0 openspec/specs/head-run-teardown/spec.md | 61 +++++++++++++++++++ 6 files changed, 61 insertions(+) rename openspec/changes/{fix-head-mode-island-teardown => archive/2026-09-25-fix-head-mode-island-teardown}/.openspec.yaml (100%) rename openspec/changes/{fix-head-mode-island-teardown => archive/2026-09-25-fix-head-mode-island-teardown}/design.md (100%) rename openspec/changes/{fix-head-mode-island-teardown => archive/2026-09-25-fix-head-mode-island-teardown}/proposal.md (100%) rename openspec/changes/{fix-head-mode-island-teardown => archive/2026-09-25-fix-head-mode-island-teardown}/specs/head-run-teardown/spec.md (100%) rename openspec/changes/{fix-head-mode-island-teardown => archive/2026-09-25-fix-head-mode-island-teardown}/tasks.md (100%) create mode 100644 openspec/specs/head-run-teardown/spec.md diff --git a/openspec/changes/fix-head-mode-island-teardown/.openspec.yaml b/openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/.openspec.yaml similarity index 100% rename from openspec/changes/fix-head-mode-island-teardown/.openspec.yaml rename to openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/.openspec.yaml diff --git a/openspec/changes/fix-head-mode-island-teardown/design.md b/openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/design.md similarity index 100% rename from openspec/changes/fix-head-mode-island-teardown/design.md rename to openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/design.md diff --git a/openspec/changes/fix-head-mode-island-teardown/proposal.md b/openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/proposal.md similarity index 100% rename from openspec/changes/fix-head-mode-island-teardown/proposal.md rename to openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/proposal.md diff --git a/openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md b/openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/specs/head-run-teardown/spec.md similarity index 100% rename from openspec/changes/fix-head-mode-island-teardown/specs/head-run-teardown/spec.md rename to openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/specs/head-run-teardown/spec.md diff --git a/openspec/changes/fix-head-mode-island-teardown/tasks.md b/openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/tasks.md similarity index 100% rename from openspec/changes/fix-head-mode-island-teardown/tasks.md rename to openspec/changes/archive/2026-09-25-fix-head-mode-island-teardown/tasks.md diff --git a/openspec/specs/head-run-teardown/spec.md b/openspec/specs/head-run-teardown/spec.md new file mode 100644 index 00000000..4306559a --- /dev/null +++ b/openspec/specs/head-run-teardown/spec.md @@ -0,0 +1,61 @@ +# head-run-teardown Specification + +## Purpose + +保证一个由 head 控制器托管的 run 在 `yeto down` 之后不会在云上留下任何计费中的机器:每个 learner 从唯一能看见它的地方拆除并逐个确认,head 只有在 learner 全部确认后才能删除,回收结束前向云核对,任何不确定都以显式失败而不是假成功结束。 + +## Requirements + +### Requirement: learner 必须从能看见它的 sky 拆除并逐个确认 + +对 head 模式的 run,`yeto down` SHALL 先在 head 上、用 head 的 sky 拆除每个 learner cluster,并为每个 learner 取得一条明确的"已拆除"或"本就不存在"的确认。本机 sky 对 learner 返回"不存在" MUST NOT 被视为该 learner 已拆除。 + +#### Scenario: 岛由 head 开出,本机 sky 不认识它 +- **WHEN** run 的元数据记录 `controller == "head"`,且某个 learner cluster 只存在于 head 的 sky 记录中 +- **THEN** `yeto down` 通过 head 执行该 learner 的拆除,并在输出中记录该 learner 的确认结果 +- **AND** 本机 sky 返回的 "does not exist" 不产生任何"已拆除"的判定 + +#### Scenario: 控制器 job 先于拆除被取消 +- **WHEN** `yeto down` 开始拆除 head 模式 run 的 learner +- **THEN** 它先取消 head 上仍在运行的控制器 job +- **AND** 拆除过程中不会有新的 learner 被重新拉起 + +### Requirement: head 只有在所有 learner 确认后才可删除 + +`yeto down` MUST NOT 删除 head,除非该 run 的每个 learner 都已取得拆除确认。存在未确认 learner 时,它 SHALL 保留 head、以非零退出码结束,并列出未确认的 learner 名字与下一步(重跑 `yeto down` 或到云控制台删除)。 + +#### Scenario: 某个 learner 无法确认拆除 +- **WHEN** 经 head 拆除后,重试用尽仍有 learner 没有确认行 +- **THEN** head 不被删除 +- **AND** 命令以非零退出码结束,错误信息列出未确认的 learner + +#### Scenario: head 暂时无法响应 +- **WHEN** 经 head 的拆除请求因 head 自身正在收尾(如其 sky 返回服务端错误)而失败 +- **THEN** `yeto down` 在有限次数内重试 +- **AND** 重试仍失败时按上一场景处理,而不是转而删除 head + +### Requirement: 回收结束前必须向云核对无残留 + +在宣布 run 已回收之前,`yeto down` SHALL 对每个能被云端查询的 cluster 向云(而不是 sky 的状态库)核对没有非终止状态的实例;发现残留 SHALL 重试拆除,重试后仍有残留 SHALL 以非零退出码结束并打印残留实例的标识。Modal learner SHALL 按其 app 已停止且无运行中任务来核对。 + +#### Scenario: 云上仍有实例存活 +- **WHEN** 拆除后云端查询仍返回该 cluster 的存活实例 +- **THEN** `yeto down` 重试拆除,重试后仍存活则非零退出并打印实例标识 +- **AND** 不打印 run 已回收的成功信息 + +#### Scenario: 云端无法核验 +- **WHEN** 某个 cluster 所在的云不支持云端实例查询,或查询本身出错 +- **THEN** `yeto down` 明确打印该 cluster 未经云端核验、按 sky 的结果信任 +- **AND** 这一情况不导致命令失败 + +### Requirement: 成功信息只在完全回收后出现 + +只有当所有 learner 已确认、head 已删除且云端核对无残留时,`yeto down` SHALL 打印 run 已回收并以零退出。任何一步未确认都 MUST NOT 产生零退出码。 + +#### Scenario: 全部回收成功 +- **WHEN** 所有 learner 确认拆除、head 删除成功、云端核对为空 +- **THEN** 命令打印 run 已回收并以零退出 + +#### Scenario: 本地模式的 run 不受影响 +- **WHEN** run 的元数据记录 `controller == "local"` +- **THEN** learner 由本机 sky 直接拆除,仍执行云端核对与成功信息的判定