Bound the source-side frame backlog: encode in executor + frame skip - #22
Merged
Merged
Conversation
Live deploy on the upgraded host (6 vCPU, 8 GB) leaked ~850 MB of
RSS in 16 seconds the moment a viewer tried to connect:
23:10:51 rss_mb=1163
23:11:07 rss_mb=2012 (+849 MB, +53 MB/s)
The leak isn't ours - it's the unbounded asyncio.Queue inside
aiortc's MediaPlayer track (contrib/media.py:229,
``self._queue: asyncio.Queue = asyncio.Queue()`` with no maxsize).
ffmpeg's worker thread pumps decoded VideoFrames into that queue;
our broadcaster pump consumes one per loop iteration. At 1080p
the raw frames are ~8 MB each (bgr0 from x11grab); even slightly
slower-than-realtime consumption piles up megabytes per second.
Two fixes that together cap the backlog at one frame:
1. ``codec.encode()`` now runs in ``loop.run_in_executor`` rather
than blocking the event loop. While libvpx works in a worker
thread, the loop can keep draining the source queue. This is
what aiortc itself does for per-sender encoders in
``RTCRtpSender._next_encoded_frame``.
2. ``_drain_to_latest`` peeks at the MediaPlayer's private
``_queue`` (stable in aiortc 1.x) and pulls every
already-buffered frame, returning only the newest. Older
frames in the backlog are counted into a new
``stale_frames_dropped`` metric and silently discarded -
far better than the alternative of a slowly-growing queue
that ends in OOM. The counter shows up in the
``video_broadcaster.stopped`` log so an operator can tell at
a glance whether their host is keeping up with the
configured resolution / framerate.
Tests:
* New ``test_drain_to_latest_skips_backlog_keeps_newest`` seeds
the source's queue with two stale frames + one current and
asserts only the latest is returned with the correct drop count.
* New ``test_drain_to_latest_no_backlog_returns_current``
confirms the empty-queue fast path is a no-op.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptom
Live on the upgraded host (6 vCPU, 8 GB) at 1080p / one viewer trying to connect:
This isn't encoder overhead — it's a backlog accumulating somewhere upstream.
Root cause
aiortc's
MediaPlayerexposes the decoded frames via a privateasyncio.Queue(contrib/media.py:229):No
maxsize. ffmpeg's worker thread pumps decoded VideoFrames in continuously; our broadcaster pump consumes one per loop iteration. At 1080p the raw frames are ~8 MB each (bgr0 from x11grab); even slightly slower-than-realtime consumption piles up megabytes per second. The math checks out: 850 MB / 16 s ≈ 53 MB/s ≈ ~7 frames/s of unconsumed 8-MB frames.Why the consumer fell behind:
codec.encode()is synchronous and ran on the event loop thread. Each call (10–30 ms at 1080p) blocked the entire loop, including the nextawait self._source.recv(). So even with libvpx running fine, the consumer was capped by event-loop occupancy.Fix
Two changes that together cap the backlog at a single frame:
1.
codec.encode()now runs inloop.run_in_executor. While libvpx works in a worker thread, the loop can keep draining the source queue. This mirrors what aiortc itself does for per-sender encoders inRTCRtpSender._next_encoded_frame.2.
_drain_to_latestpeeks at the MediaPlayer's_queue(stable in aiortc 1.x) and pulls every already-buffered frame, keeping only the newest. Older frames are counted intostale_frames_droppedand discarded — far better than the alternative of a slowly-growing queue that ends in OOM. The counter shows up invideo_broadcaster.stoppedso an operator can tell whether the host is keeping up.Test plan
pytest— 252 passed.test_drain_to_latest_skips_backlog_keeps_newestseeds the queue with two stale frames + one current and asserts only the newest is returned with the correct drop count.test_drain_to_latest_no_backlog_returns_currentconfirms the empty-queue fast path is a no-op.mypy --strictclean.ruff checkclean.stale_frames_droppedshould be small (a handful, not hundreds per second).Operator notes
After pulling and rebuilding, watch the
ts3.heartbeat rss_mb=line in the bot logs while one viewer connects. Expected behaviour:video_broadcaster.stopped stale_frames_dropped=N(when youdownthe stack) tells you whether you were keeping up:Nnear zero → host has plenty of headroom, you can push to 60 fps or higher resolutionNin the thousands → host is at its limit, lower fps or resolutionhttps://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at
Generated by Claude Code