Skip to content

Persist TS3 identity across container restarts - #17

Merged
queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt
May 1, 2026
Merged

queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt

Conversation

@queueeee

@queueeee queueeee commented May 1, 2026

Copy link
Copy Markdown
Owner

Summary

Every container start was minting a fresh hashcash-mined TS3 identity, so the TS6 server saw each restart as an unrelated cryptographic client. Stream slots from previous runs accumulated as zombies that the TS6 client UI still listed for joins; clicking a zombie routed the request to a long-gone clid and the live bot saw nothing.

This matches the symptom we just hit in network_mode: host: TS3 heartbeats flowed every second, the bot held a fresh setupstream slot, but notifyjoinstreamrequest never showed up in the log when a viewer clicked Join. The click was hitting one of the older zombie slots from previous OOM/revert restarts in this debug session, not the live one.

Changes

  • New src/ts6_stream_bot/ts3lib/identity_store.py with load_or_generate_identity(path, security_level) — atomic write via tempfile + os.replace, file mode 0600, parent dir created on demand.
  • Loader validates the stored public-key string against the private scalar (catches a tampered or truncated file before any signing happens). Corrupt JSON or invalid pairs fall back to fresh generation rather than locking the bot out.
  • controller._connect_ts6 swaps the unconditional generate_identity_async for the new helper.
  • config.py: IDENTITY_PATH (default /app/state/identity.json) and IDENTITY_SECURITY_LEVEL (default 8).
  • docker-compose.yml: new bot-state named volume mounted at /app/state. Existing operator setups need a one-time docker compose ... up -d to create it; data layer is otherwise unchanged.
  • .env.example: documents the two new keys.

Operator notes

After merging, on the live host:

cd /opt/ts6-stream-bot
docker compose -f docker-compose.yml -f docker-compose.host.yml down
docker compose -f docker-compose.yml -f docker-compose.host.yml up -d --build
docker compose -f docker-compose.yml -f docker-compose.host.yml logs -f

First start mines + writes a new identity (the hashcash work happens once). Every subsequent restart reuses the same identity, so the TS6 server cleans up the previous session via its normal client-disconnect path and there is exactly one stream slot per bot in the channel UI.

The bot-state volume holds the bot's TS6 credentials — anyone with read access to the host-side path can impersonate it, so treat that directory as sensitive.

Test plan

  • pytest — 228 passed, 1 skipped (full suite, including 9 new identity_store tests covering cold start, warm start, corrupt JSON, invalid keypair, file permissions, missing parent dir, tempfile cleanup, and concurrent first-call race).
  • mypy --strict clean (37 source files).
  • ruff check clean.
  • Live test on the host: docker compose ... down && up -d --build, click Join in TS6 client → expect ts3.command_received command=notifyjoinstreamrequest followed by stream_publisher.viewer_offer_sent in the bot log.

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at


Generated by Claude Code

Every container start was minting a fresh hashcash identity, so the
TS6 server saw each restart as an unrelated client. Stream slots
from previous runs accumulated as zombies that the client UI still
listed for joins; clicking one routed the request to a long-gone
clid and the live bot saw nothing - matches the symptom we hit in
host-network mode where heartbeats flowed but no
`notifyjoinstreamrequest` ever arrived.

Stores the identity JSON at /app/state/identity.json (mode 0600),
backed by a `bot-state` named volume in compose. Loads via a small
helper that validates the public-key string against the private
scalar before trusting the file, and atomically rewrites it via
tempfile + os.replace to keep partial writes from bricking startup.
@queueeee
queueeee merged commit 4ecc4c8 into main May 1, 2026
1 check passed
queueeee pushed a commit that referenced this pull request May 1, 2026
Two unrelated stability fixes that surfaced after the live test of
identity persistence (PR #17):

* `register_stream_notifications` was sending three
  `servernotifyregister` calls during stream setup; each got a
  `error id=516 invalid client type` back from the server because
  that command is for ServerQuery clients (type 1) only. As a voice
  client (type 0) we receive stream + channel notifications
  automatically, so the call is dead weight that polluted the log
  on every startup. Drop the call site and the method; keep a
  comment so the next reader doesn't re-add it from the upstream
  ts6-manager port (which is a ServerQuery client and does need it).

* Each viewer spins up its own VP8 encoder via the per-viewer
  `RTCPeerConnection` + MediaRelay subscription. Live measurement
  on 720p30 showed RSS jumping from ~110 MB (idle) to ~400 MB after
  the first viewer connected - the unlimited default
  (`STREAM_VIEWER_LIMIT=-1`) would silently OOM-kill the container
  on a small host once a fourth or fifth viewer joined. Default
  cap to 4; the operator can bump it once they've measured.
queueeee added a commit that referenced this pull request May 1, 2026
)

Two unrelated stability fixes that surfaced after the live test of
identity persistence (PR #17):

* `register_stream_notifications` was sending three
  `servernotifyregister` calls during stream setup; each got a
  `error id=516 invalid client type` back from the server because
  that command is for ServerQuery clients (type 1) only. As a voice
  client (type 0) we receive stream + channel notifications
  automatically, so the call is dead weight that polluted the log
  on every startup. Drop the call site and the method; keep a
  comment so the next reader doesn't re-add it from the upstream
  ts6-manager port (which is a ServerQuery client and does need it).

* Each viewer spins up its own VP8 encoder via the per-viewer
  `RTCPeerConnection` + MediaRelay subscription. Live measurement
  on 720p30 showed RSS jumping from ~110 MB (idle) to ~400 MB after
  the first viewer connected - the unlimited default
  (`STREAM_VIEWER_LIMIT=-1`) would silently OOM-kill the container
  on a small host once a fourth or fifth viewer joined. Default
  cap to 4; the operator can bump it once they've measured.

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants