Skip to content

Cap libvpx threads so the encoder doesn't starve a colocated TS6 server - #23

Merged
queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt
May 1, 2026
Merged

queueeee merged 1 commit into
mainfrom
claude/review-claude-md-iOuUt

Conversation

@queueeee

@queueeee queueeee commented May 1, 2026

Copy link
Copy Markdown
Owner

Symptom

On the upgraded host (6 vCPU, 8 GB) at 1080p the broadcaster's memory plateau looks great (~280 MB stable), but the operator reports:

[Beim Join] sehe keinen Log, weil der Server vorher stirbt bzw das unterliegende software [reset]

The bot is alive, the TS6 server is alive, until the moment a viewer clicks Join — at which point everything goes silent. No respondjoinstreamrequest ever leaves the bot. No log lines around the join.

Diagnosis

This isn't memory anymore — that's fixed by #22. It's CPU starvation.

Our broadcaster's _auto_thread_count was a straight port of aiortc.codecs.vpx.number_of_threads, which returns min(cpu_count, 8). On the 6-vCPU host, 1080p set thread_count = 6. So during each frame's encode, libvpx pinned all six cores at once. Inter-frame gaps (a few ms at 30 fps) were the only window for ffmpeg's reader thread, the asyncio event loop, and the TS6 server's watchdog to make progress.

When the WebRTC handshake's added load tipped that balance, the TS6 server's watchdog couldn't get scheduled in time, decided the process was hung, and killed it. The bot's TS3 socket then disconnected — losing whatever log lines were buffered locally and the in-flight join response.

ts6-manager doesn't hit this because Pion encodes much faster (less wall-time pinned per frame) and because they run encode in a separate Go process.

Fix

_auto_thread_count now caps at min(aiortc_pick, cpu_count - 2, 4):

Host Resolution Old (aiortc) New
2 vCPU 1080p 2 1 (leaves 1 free)
6 vCPU 1080p 6 4 (leaves 2 free)
8 vCPU 1080p 8 4
6 vCPU 720p 4 4

Operators with measured headroom can override via:

ENCODER_THREAD_COUNT=6

And operators whose host can't sustain real-time at the configured resolution can speed the encoder up (at the cost of quality):

ENCODER_CPU_USED=-10

Test plan

  • pytest — 253 passed (1 new test pinning the heuristic across 2/6/8-vCPU and 720p/1080p combinations).
  • mypy --strict clean.
  • ruff check clean.
  • Live: at 1080p30 with the new cap, the TS6 server should no longer get watchdog-killed mid-join. RSS plateau still ~280 MB. respondjoinstreamrequest actually arrives at the viewer.

Operator notes

cd /opt/ts6-stream-bot
git pull
docker compose -f docker-compose.yml -f docker-compose.host.yml down
docker compose -f docker-compose.yml -f docker-compose.host.yml up -d --build
docker compose -f docker-compose.yml -f docker-compose.host.yml logs -f

You don't need to change .env — the new defaults pick the safe values automatically. If 1080p30 still looks rough (bot logs stale_frames_dropped in the thousands when you down the stack), drop the encoder speed:

ENCODER_CPU_USED=-10

Or fall back to 720p30:

SCREEN_WIDTH=1280
SCREEN_HEIGHT=720

https://claude.ai/code/session_016DuCjRJK995Tj9aDhhB9at


Generated by Claude Code

Live on the 6-vCPU / 8 GB upgrade: bot's RSS plateau is healthy (~280 MB)
but the moment a viewer clicks Join, the TS6 server next door goes
silent and takes the bot's connection with it. Operator: "der Server
[stirbt] vorher [...] das unterliegende software [reset]". No bot-side
log of the join request, no `respondjoinstreamrequest` ever sent -
because by the time the bot would write that to its socket, its peer
(the TS6 process) has been killed by its own watchdog.

Root cause: libvpx's thread_count heuristic in the broadcaster (a
straight port of aiortc.codecs.vpx) returned ``min(cpu_count, 8)``,
so 1080p on the 6-vCPU host pinned all 6 cores during each frame's
encode. ffmpeg's reader thread, the bot's event loop, the TS6
server, and its watchdog all had to share whatever microseconds
were left between encode bursts - which on TS6's beta build
apparently isn't enough.

* New ``_auto_thread_count`` caps at ``min(aiortc_pick, cpu_count - 2,
  4)``. On the live host that drops 1080p from 6 threads to 4 and
  guarantees at least 2 cores are always free for everything else.

* Two new settings (``ENCODER_THREAD_COUNT``, ``ENCODER_CPU_USED``)
  wire through to ``VideoBroadcasterConfig`` so an operator who's
  measured headroom can lift the cap, and one who can't sustain
  real-time can drop ``cpu_used`` to -8 or -10.

* New unit test ``test_auto_thread_count_caps_at_four_and_leaves_headroom``
  pins the contract: 1080p / 720p on 2-vCPU, 6-vCPU, and 8-vCPU
  hosts all return values that leave at least one core free.
@queueeee
queueeee merged commit 1e96a78 into main May 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants