Cap libvpx threads so the encoder doesn't starve a colocated TS6 server - #23
Merged
Merged
Conversation
Live on the 6-vCPU / 8 GB upgrade: bot's RSS plateau is healthy (~280 MB) but the moment a viewer clicks Join, the TS6 server next door goes silent and takes the bot's connection with it. Operator: "der Server [stirbt] vorher [...] das unterliegende software [reset]". No bot-side log of the join request, no `respondjoinstreamrequest` ever sent - because by the time the bot would write that to its socket, its peer (the TS6 process) has been killed by its own watchdog. Root cause: libvpx's thread_count heuristic in the broadcaster (a straight port of aiortc.codecs.vpx) returned ``min(cpu_count, 8)``, so 1080p on the 6-vCPU host pinned all 6 cores during each frame's encode. ffmpeg's reader thread, the bot's event loop, the TS6 server, and its watchdog all had to share whatever microseconds were left between encode bursts - which on TS6's beta build apparently isn't enough. * New ``_auto_thread_count`` caps at ``min(aiortc_pick, cpu_count - 2, 4)``. On the live host that drops 1080p from 6 threads to 4 and guarantees at least 2 cores are always free for everything else. * Two new settings (``ENCODER_THREAD_COUNT``, ``ENCODER_CPU_USED``) wire through to ``VideoBroadcasterConfig`` so an operator who's measured headroom can lift the cap, and one who can't sustain real-time can drop ``cpu_used`` to -8 or -10. * New unit test ``test_auto_thread_count_caps_at_four_and_leaves_headroom`` pins the contract: 1080p / 720p on 2-vCPU, 6-vCPU, and 8-vCPU hosts all return values that leave at least one core free.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptom
On the upgraded host (6 vCPU, 8 GB) at 1080p the broadcaster's memory plateau looks great (~280 MB stable), but the operator reports:
The bot is alive, the TS6 server is alive, until the moment a viewer clicks Join — at which point everything goes silent. No
respondjoinstreamrequestever leaves the bot. No log lines around the join.Diagnosis
This isn't memory anymore — that's fixed by #22. It's CPU starvation.
Our broadcaster's
_auto_thread_countwas a straight port ofaiortc.codecs.vpx.number_of_threads, which returnsmin(cpu_count, 8). On the 6-vCPU host, 1080p setthread_count = 6. So during each frame's encode, libvpx pinned all six cores at once. Inter-frame gaps (a few ms at 30 fps) were the only window for ffmpeg's reader thread, the asyncio event loop, and the TS6 server's watchdog to make progress.When the WebRTC handshake's added load tipped that balance, the TS6 server's watchdog couldn't get scheduled in time, decided the process was hung, and killed it. The bot's TS3 socket then disconnected — losing whatever log lines were buffered locally and the in-flight join response.
ts6-manager doesn't hit this because Pion encodes much faster (less wall-time pinned per frame) and because they run encode in a separate Go process.
Fix
_auto_thread_countnow caps atmin(aiortc_pick, cpu_count - 2, 4):Operators with measured headroom can override via:
And operators whose host can't sustain real-time at the configured resolution can speed the encoder up (at the cost of quality):
Test plan
pytest— 253 passed (1 new test pinning the heuristic across 2/6/8-vCPU and 720p/1080p combinations).mypy --strictclean.ruff checkclean.respondjoinstreamrequestactually arrives at the viewer.Operator notes
You don't need to change
.env— the new defaults pick the safe values automatically. If 1080p30 still looks rough (bot logsstale_frames_droppedin the thousands when youdownthe stack), drop the encoder speed:Or fall back to 720p30:
https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at
Generated by Claude Code