Skip to content

Address bot crash on viewer kick: STUN + 720p defaults + robust SIGKILL recovery - #13

Merged
queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt
May 1, 2026
Merged

queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt

Conversation

@queueeee

@queueeee queueeee commented May 1, 2026

Copy link
Copy Markdown
Owner

Live trace summary

Operator clarified that the bot crashes FIRST → the viewer kick is a
consequence. The trace shows:

  • First viewer click: ice=checking for ~20 s, then the client gave up
    and sent cmd=reconnect — classic NAT-punch failure when one side
    has no STUN-derived srflx candidate.
  • After the reconnect, second attempt actually completed:
    ice=completed, state=connected, parec_audio.started.
  • ~6 s into active streaming: exit 137 (SIGKILL). Almost certainly
    OOM from the host — aiortc + PyAV doing 1080p30 VP8 encoding is
    memory-hungry and the container had no mem_limit.
  • On restart, PulseAudio bombs with "Daemon already running" because
    the SIGKILL skipped our cleanup trap and the pid file persisted in
    a path our explicit list didn't cover.

Fixes

  1. STUN in RTCPeerConnection config:
    stun:stun.l.google.com:19302. Same default the browser stacks use.
    Should make the first ICE attempt succeed without needing the
    reconnect dance.

  2. Default capture size 1280x720 (was 1920x1080). Cuts the per-peer
    encoder budget by ~2.25×. Operators with bigger hosts can override
    via SCREEN_WIDTH/SCREEN_HEIGHT in .env (and bump
    STREAM_BITRATE to match).

  3. Robust entrypoint cleanup against SIGKILL:

    • Up-front pkill -9 of any leftover pulseaudio / Xvfb from a
      prior crashed run.
    • find /run /var/run /root /tmp /home -name pid -path '*pulse*' -delete to catch XDG_RUNTIME_DIR variants the explicit path
      list missed.
    • Same find in the shutdown trap so a normal stop is also clean.
  4. Memory in heartbeat: ts3.heartbeat ... rss_mb=… lets us see
    the memory trajectory leading up to the next OOM (if there still is
    one).

Operator notes

If exit 137 still happens after this, on the host run:

dmesg | grep -i "killed process\|out of memory" | tail
docker stats ts6-stream-bot --no-stream

If dmesg confirms OOM and the host genuinely has no headroom, set a
container limit in docker-compose.yml:

services:
  bot:
    deploy:
      resources:
        limits:
          memory: 1500M

…and consider lowering further: SCREEN_WIDTH=854 SCREEN_HEIGHT=480 SCREEN_FPS=20 STREAM_BITRATE=1500 will encode comfortably in well
under 1 GB.

Test plan

  • ruff check src tests — clean
  • mypy --strict — clean
  • pytest — 218 passed
  • Live: redeploy. Expected:
    • First viewer click connects on the first try (no cmd=reconnect
      detour) thanks to STUN.
    • ts3.heartbeat rss_mb=… shows memory pattern; if it stays under
      ~600 MB the OOM should be gone.

Operator commands

cd /opt/ts6-stream-bot
git pull
docker compose down
docker compose build --no-cache
docker compose up -d
docker compose logs -f

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at


Generated by Claude Code

The previous live deploy showed the bot exit 137 (SIGKILL) ~6 seconds
into a successful viewer connection. Operator confirmed: the kick is
the SYMPTOM of the crash, not the cause - bot dies, viewer's session
follows. Likely root cause is OOM from the host: aiortc + PyAV doing
1080p30 VP8 encoding for even one peer is heavyweight, and the
container had no memory budget set.

Four changes:

1. Add STUN to RTCPeerConnection. The first-attempt ICE timeout we saw
   in the trace (`ice=checking` for 20 s before the client gave up and
   sent cmd=reconnect) is exactly the symptom of the bot lacking a
   server-reflexive candidate when both peers are behind NAT. Default
   to stun:stun.l.google.com:19302 - same conventional value the
   browser-side WebRTC stacks use out of the box.

2. Lower SCREEN_WIDTH/SCREEN_HEIGHT defaults from 1920x1080 to
   1280x720. 720p30 keeps the per-viewer encode comfortably under
   the OOM line on a 4 GB host. Bump in .env if your host has the
   budget; remember to bump STREAM_BITRATE alongside.

3. Robust entrypoint cleanup against SIGKILL. The trap doesn't run on
   SIGKILL so the pid file persists, AND a stray pulseaudio child can
   linger inside the same restarted container. Now we pkill -9
   pulseaudio + Xvfb up front and `find ... -name pid -path '*pulse*'
   -delete` to catch XDG_RUNTIME_DIR variants we missed in the
   explicit path list. Same on shutdown.

4. Memory in heartbeat: `ts3.heartbeat ... rss_mb=…` so the next
   exit-137 trace shows the trajectory before the kill.

ruff + mypy strict + 218 tests pass.

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at
@queueeee
queueeee merged commit fdab823 into main May 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants