Skip to content

Single-encoder video broadcast (one libvpx, N viewers) - #20

Merged
queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt
May 1, 2026
Merged

queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt

Conversation

@queueeee

@queueeee queueeee commented May 1, 2026

Copy link
Copy Markdown
Owner

Summary

Per-viewer VP8 encoders were the dominant memory cost: each RTCRtpSender ran its own libvpx through aiortc's Vp8Encoder, so per-viewer RSS scaled linearly with viewer count. Live measurement on 720p30 had one viewer push the bot's RSS past 750 MB and still climbing — at 1080p with 5+ viewers, no realistic host was big enough.

aiortc has a built-in pre-encoded path (rtcrtpsender.py:311-323):

if isinstance(data, Frame):
    payloads, timestamp = await loop.run_in_executor(None, encoder.encode, data, force_keyframe)
else:
    # Pack the pre-encoded data.
    payloads, timestamp = self.__encoder.pack(data)

When a track's recv() returns something other than an av.Frame, the sender calls encoder.pack(data) — cheap RTP packetization — instead of encoder.encode(frame) — expensive libvpx. This PR drops in a VideoBroadcaster that runs one libvpx context, encodes each captured frame once, and fans the resulting av.Packet out to per-viewer asyncio.Queues.

Architecture

              ┌─────────────────────────────────────────────┐
              │  VideoBroadcaster (one task, one libvpx)    │
  x11grab ───►│  recv() VideoFrame → encode → av.Packet     │
              │       └─ fan out to subscriber queues       │
              └────────┬───────────────┬──────────┬─────────┘
                       │               │          │
                  Queue#1          Queue#2     Queue#N
                       │               │          │
              ┌────────▼────┐ ┌────────▼────┐ ┌──▼──────────┐
              │BroadcastTrack│ │BroadcastTrack│ │BroadcastTrack│
              │ recv()→Packet│ │ recv()→Packet│ │ recv()→Packet│
              └────────┬────┘ └────────┬────┘ └──┬──────────┘
                       │               │          │
              ┌────────▼────┐ ┌────────▼────┐ ┌──▼──────────┐
              │RTCRtpSender │ │RTCRtpSender │ │RTCRtpSender │
              │ pack(packet)│ │ pack(packet)│ │ pack(packet)│
              │  (cheap RTP)│ │  (cheap RTP)│ │  (cheap RTP)│
              └─────────────┘ └─────────────┘ └─────────────┘

The expensive path (libvpx encode) runs once. Each viewer's per-sender Vp8Encoder only does RTP packetization (pack), which is just bytes manipulation.

Trade-offs accepted

  • Shared bitrate (no per-viewer REMB adaptation). We never used REMB-driven adaptation in practice — STREAM_BITRATE was static.
  • A PLI from any one viewer triggers a keyframe for all viewers. Acceptable: keyframes are infrequent (gop_size=3000 frames = 100 s at 30 fps).
  • Audio still flows through MediaRelay. Opus is cheap enough that the broadcaster pattern doesn't pay back; revisit if it becomes a bottleneck later.

Behaviour

  • One libvpx instance for the whole channel. Encoder settings mirror aiortc.codecs.vpx.Vp8Encoder (cpu-used=-6, deadline=realtime, gop=3000) so the wire format is unchanged.
  • Per-viewer asyncio.Queue with queue_size cap (default 30, ~1 s at 30 fps). On QueueFull we drop the oldest packet for that subscriber rather than block — one slow viewer can't stall the rest. Persistently broken peers are still evicted by the existing connectionstatechange callback.
  • New subscribe() sets force_keyframe = True so the joining viewer can start decoding without waiting for the next periodic I-frame.
  • Pump task survives encoder errors (logged, frame skipped) but exits on source MediaStreamError. On shutdown, all subscriber queues get a None sentinel that surfaces as MediaStreamError in recv() so each sender shuts down cleanly.

Test plan

  • pytest — 245 passed, 1 skipped (11 new in test_video_broadcaster.py):
    • subscription mechanics (subscribe returns video track, sets force_keyframe, unsubscribe removes)
    • recv() semantics (returns queued packet, raises MediaStreamError on None sentinel)
    • fanout (one packet → all subscribers, distinct queues)
    • slow-consumer drop (oldest evicted, drop counter increments, fast viewer not starved)
    • end-to-end smoke against real libvpx: synthetic frames → encode → fan out → Packet arrives at every recv()
    • stop() wakes pending recv() with sentinel
    • factory returning None raises at start()
  • mypy --strict clean (38 source files now).
  • ruff check clean.
  • Live: with the new defaults from PR Resource hardening for small hosts colocated with TS6 #19 plus this PR, test 5–7 viewers @ 1080p on the 4 GB host. Expected RSS plateau: ~250 MB base + roughly constant overhead per viewer (single encoder + N small RTP queues), targeting a total well under 1 GB even at 7 viewers.

Operator notes

After merge:

cd /opt/ts6-stream-bot
git pull
docker compose -f docker-compose.yml -f docker-compose.host.yml down
docker compose -f docker-compose.yml -f docker-compose.host.yml up -d --build
docker compose -f docker-compose.yml -f docker-compose.host.yml logs -f

Watch for:

  • video_broadcaster.started near startup (right after stream_publisher.started).
  • video_broadcaster.subscribed total_subscribers=N on each viewer join.
  • ts3.heartbeat rss_mb=… — should stay roughly flat as more viewers join, not climb linearly anymore.

If you want to push past 480p24 / 2500 kbps now that per-viewer cost is dramatically lower:

SCREEN_WIDTH=1920
SCREEN_HEIGHT=1080
SCREEN_FPS=30
STREAM_BITRATE=6000
STREAM_VIEWER_LIMIT=7

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at


Generated by Claude Code

Per-viewer VP8 encoders were the dominant memory cost: each
RTCRtpSender ran its own libvpx through aiortc's Vp8Encoder, so
per-viewer RSS scaled linearly with viewer count. Live measurement
on 720p30 had one viewer push the bot's RSS past 750 MB and still
climbing - 5+ viewers at 1080p was a non-starter on any reasonable
host.

aiortc has a built-in pre-encoded path (rtcrtpsender.py:311-323):
when a track's recv() returns something other than av.Frame, the
sender calls encoder.pack(data) (cheap RTP packetization) instead
of encoder.encode(frame) (expensive libvpx). This commit drops
in a VideoBroadcaster that runs a single libvpx context, encodes
each captured frame once, and fans the resulting av.Packet out to
per-viewer asyncio.Queues. Each BroadcastVideoTrack.recv() returns
those packets; aiortc takes the cheap path.

Concrete behaviour:

* One libvpx instance for the whole channel. Encoder settings
  mirror aiortc.codecs.vpx.Vp8Encoder (cpu-used=-6, deadline=
  realtime, gop=3000) so the wire format is unchanged.
* Per-viewer asyncio.Queue with a queue_size cap (default 30 ~
  1 s at 30 fps). On QueueFull we drop the oldest packet for
  that subscriber rather than block - one slow viewer can't
  stall the rest. The connection-state callback already evicts
  persistently broken peers.
* New subscribe() forces a keyframe so the joining viewer can
  start decoding without waiting for the next periodic I-frame.
* Audio still flows through MediaRelay; Opus is cheap enough
  that the broadcaster pattern doesn't pay back.

Trade-offs we accept:

* Shared bitrate (no per-viewer REMB adaptation). We never used
  REMB-driven adaptation in practice - bitrate was static.
* A PLI from any one viewer triggers a keyframe for all viewers.
  Acceptable: keyframes are infrequent (gop_size=3000 frames =
  100 s at 30 fps) and a few extra ones cost less than running
  N encoders.

Tests:
* 11 new tests in test_video_broadcaster.py covering subscription
  mechanics, fanout, slow-consumer drop, keyframe-on-join,
  shutdown sentinel, plus an end-to-end smoke test against real
  libvpx that confirms av.Packets reach every subscriber.
* test_stream_publisher.py updated for the new constructor +
  the per-viewer track-attachment expectation (broadcaster for
  video, relay for audio).
@queueeee
queueeee merged commit 7a7b493 into main May 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants