Skip to content

Dedup duplicate join notifies + reject joins on dead broadcaster - #21

Merged
queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt
May 1, 2026
Merged

queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt

Conversation

@queueeee

@queueeee queueeee commented May 1, 2026

Copy link
Copy Markdown
Owner

Symptom

Live deploy at 1080p60 / 6 Mbit / viewer_limit=5: bot connects fine, stream allocates, then on the very first viewer click the bot goes into a tailspin. The log told the story:

22:53:01.780  notifyjoinstreamrequest clid=72   ← first
22:53:01.785  notifyjoinstreamrequest clid=72   ← +5ms
22:53:01.787  notifyjoinstreamrequest clid=72
... 9 retransmits within 16 ms
22:53:01.888  video_broadcaster.subscribed total_subscribers=1
22:53:02.276  video_broadcaster.subscribed total_subscribers=9   ← !!!
22:53:02.302  rss_mb=1343                       ← 1.3 GB RSS for one viewer
22:53:26.554  video_broadcaster.source_ended    ← ffmpeg/x11grab OOM-killed

The TS6 server retransmitted the same notifyjoinstreamrequest 9 times inside 16 ms. Our existing dedup checked self._viewers.pop(...) before the slow path (createOffer + ICE gathering = tens of ms), but the viewer slot only lands in _viewers after that slow path completes. All 9 retransmits raced past the dedup, each spawning a fresh peer connection + broadcaster subscription. The 9× concurrent ICE gatherings + encoder subscriptions blew RSS to 1.3 GB and OOM-killed the capture pipeline. After that, the broadcaster's pump task exited but subscribe() kept handing out queues nothing would ever fill — so subsequent viewers' UI hung on "connecting" forever (= "kommt nicht wieder").

Fixes

StreamPublisher._joining — a lock-protected set[int] populated before any expensive work (offer creation, ICE, broadcaster subscribe) and discarded in a finally block. Duplicate notifies for the same clid log stream_publisher.duplicate_join_dropped and bail out cheaply without touching aiortc or the broadcaster.

VideoBroadcaster.is_alive — False after the pump task exits for any reason (source MediaStreamError, encoder crash). When false, _handle_viewer_join sends respondjoinstreamrequest decision=0 and exits — the TS6 client UI fails the join cleanly instead of hanging on "connecting".

Test plan

  • pytest — 250 passed, 1 skipped.
    • New: test_duplicate_join_for_same_clid_only_creates_one_pc bursts 9 duplicate notifies and asserts exactly one PC, one broadcaster subscription, one audio relay subscription.
    • New: test_join_refused_when_broadcaster_source_dead flips is_alive=False and asserts decision=0 in the response with no PC, no subscriptions.
    • New: test_is_alive_* covers the broadcaster's lifecycle flag.
    • The test PC mock for createOffer / setLocalDescription now await asyncio.sleep(0)s — without it the in-flight dedup looked effective only because the mocked path ran to completion synchronously, but real aiortc's ICE-gathering yield is what lets parallel tasks observe each other's _joining state.
  • mypy --strict clean.
  • ruff check clean.
  • Live: at 1080p60 click should now produce a single peer with stream_publisher.duplicate_join_dropped logs for the 8 retransmits, no RSS spike, no OOM.

Operator notes

Even with this fix, 1080p60 is still asking a lot from libvpx-realtime on a small VPS. If the bot stays alive after this PR but the encoder still can't keep up (frames pile up, RSS climbs), drop to 1080p30:

SCREEN_FPS=30

Or 720p60:

SCREEN_WIDTH=1280
SCREEN_HEIGHT=720
SCREEN_FPS=60
STREAM_BITRATE=4500
cd /opt/ts6-stream-bot
git pull
docker compose -f docker-compose.yml -f docker-compose.host.yml down
docker compose -f docker-compose.yml -f docker-compose.host.yml up -d --build
docker compose -f docker-compose.yml -f docker-compose.host.yml logs -f

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at


Generated by Claude Code

Live deploy at 1080p60 hit two compounding bugs the moment a viewer
clicked Join:

1. The TS6 server retransmitted ``notifyjoinstreamrequest`` 9 times
   inside 16 ms (server quirk; we ACK at the TS3 layer but it
   resent anyway). The publisher's existing dedup looked at
   ``self._viewers.pop(...)`` *before* the slow path - but the slow
   path takes tens of milliseconds (createOffer + ICE gathering)
   and the viewer slot only lands in ``_viewers`` after that work
   completes. So all 9 retransmits raced past the dedup, each
   spawning a fresh peer connection + broadcaster subscription
   before any of them could populate the dict. Nine concurrent
   ICE gatherings + nine encoder subscriptions blew RSS to 1.3 GB
   and OOM-killed ffmpeg/x11grab.

2. Once the broadcaster's source raised MediaStreamError on the
   OOM-killed capture, the pump task exited but ``subscribe()``
   kept handing out queues that nothing would ever fill. Newly
   joining viewers got an offer that pointed at a dead media
   pipeline; their UI hung on "connecting" forever.

Fixes:

* ``StreamPublisher._joining: set[int]`` lock-protected, populated
  *before* any expensive work (offer creation, ICE, broadcaster
  subscribe) and discarded in a finally block. Duplicate notifies
  for the same clid log ``stream_publisher.duplicate_join_dropped``
  and bail out without touching aiortc or the broadcaster.

* ``VideoBroadcaster.is_alive`` (False after the pump task exits
  for any reason). When false, ``_handle_viewer_join`` sends
  ``respondjoinstreamrequest decision=0`` so the TS6 client UI
  fails the join cleanly rather than hanging.

The test mock for ``RTCPeerConnection.createOffer`` / ``setLocalDescription``
now actually yields via ``await asyncio.sleep(0)``. Without that,
the in-flight dedup looked effective only because the mocked path
ran to completion synchronously - in real aiortc the ICE-gathering
yield would let parallel tasks observe each other's ``_joining``
state, which is exactly what we want to assert.
@queueeee
queueeee merged commit cfc5d34 into main May 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants