Single-encoder video broadcast (one libvpx, N viewers) - #20
Merged
Merged
Conversation
Per-viewer VP8 encoders were the dominant memory cost: each RTCRtpSender ran its own libvpx through aiortc's Vp8Encoder, so per-viewer RSS scaled linearly with viewer count. Live measurement on 720p30 had one viewer push the bot's RSS past 750 MB and still climbing - 5+ viewers at 1080p was a non-starter on any reasonable host. aiortc has a built-in pre-encoded path (rtcrtpsender.py:311-323): when a track's recv() returns something other than av.Frame, the sender calls encoder.pack(data) (cheap RTP packetization) instead of encoder.encode(frame) (expensive libvpx). This commit drops in a VideoBroadcaster that runs a single libvpx context, encodes each captured frame once, and fans the resulting av.Packet out to per-viewer asyncio.Queues. Each BroadcastVideoTrack.recv() returns those packets; aiortc takes the cheap path. Concrete behaviour: * One libvpx instance for the whole channel. Encoder settings mirror aiortc.codecs.vpx.Vp8Encoder (cpu-used=-6, deadline= realtime, gop=3000) so the wire format is unchanged. * Per-viewer asyncio.Queue with a queue_size cap (default 30 ~ 1 s at 30 fps). On QueueFull we drop the oldest packet for that subscriber rather than block - one slow viewer can't stall the rest. The connection-state callback already evicts persistently broken peers. * New subscribe() forces a keyframe so the joining viewer can start decoding without waiting for the next periodic I-frame. * Audio still flows through MediaRelay; Opus is cheap enough that the broadcaster pattern doesn't pay back. Trade-offs we accept: * Shared bitrate (no per-viewer REMB adaptation). We never used REMB-driven adaptation in practice - bitrate was static. * A PLI from any one viewer triggers a keyframe for all viewers. Acceptable: keyframes are infrequent (gop_size=3000 frames = 100 s at 30 fps) and a few extra ones cost less than running N encoders. Tests: * 11 new tests in test_video_broadcaster.py covering subscription mechanics, fanout, slow-consumer drop, keyframe-on-join, shutdown sentinel, plus an end-to-end smoke test against real libvpx that confirms av.Packets reach every subscriber. * test_stream_publisher.py updated for the new constructor + the per-viewer track-attachment expectation (broadcaster for video, relay for audio).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Per-viewer VP8 encoders were the dominant memory cost: each
RTCRtpSenderran its own libvpx through aiortc'sVp8Encoder, so per-viewer RSS scaled linearly with viewer count. Live measurement on 720p30 had one viewer push the bot's RSS past 750 MB and still climbing — at 1080p with 5+ viewers, no realistic host was big enough.aiortc has a built-in pre-encoded path (
rtcrtpsender.py:311-323):When a track's
recv()returns something other than anav.Frame, the sender callsencoder.pack(data)— cheap RTP packetization — instead ofencoder.encode(frame)— expensive libvpx. This PR drops in aVideoBroadcasterthat runs one libvpx context, encodes each captured frame once, and fans the resultingav.Packetout to per-viewerasyncio.Queues.Architecture
The expensive path (libvpx encode) runs once. Each viewer's per-sender
Vp8Encoderonly does RTP packetization (pack), which is just bytes manipulation.Trade-offs accepted
STREAM_BITRATEwas static.gop_size=3000frames = 100 s at 30 fps).Behaviour
aiortc.codecs.vpx.Vp8Encoder(cpu-used=-6,deadline=realtime,gop=3000) so the wire format is unchanged.asyncio.Queuewithqueue_sizecap (default 30, ~1 s at 30 fps). OnQueueFullwe drop the oldest packet for that subscriber rather than block — one slow viewer can't stall the rest. Persistently broken peers are still evicted by the existingconnectionstatechangecallback.subscribe()setsforce_keyframe = Trueso the joining viewer can start decoding without waiting for the next periodic I-frame.MediaStreamError. On shutdown, all subscriber queues get aNonesentinel that surfaces asMediaStreamErrorinrecv()so each sender shuts down cleanly.Test plan
pytest— 245 passed, 1 skipped (11 new intest_video_broadcaster.py):subscribereturns video track, sets force_keyframe,unsubscriberemoves)recv()semantics (returns queued packet, raisesMediaStreamErroronNonesentinel)Packetarrives at everyrecv()stop()wakes pendingrecv()with sentinelNoneraises atstart()mypy --strictclean (38 source files now).ruff checkclean.Operator notes
After merge:
Watch for:
video_broadcaster.startednear startup (right afterstream_publisher.started).video_broadcaster.subscribed total_subscribers=Non each viewer join.ts3.heartbeat rss_mb=…— should stay roughly flat as more viewers join, not climb linearly anymore.If you want to push past 480p24 / 2500 kbps now that per-viewer cost is dramatically lower:
https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at
Generated by Claude Code