We are evaluating MOSS-TTS-Realtime for a telephony voice agent and cannot
reproduce the published figures. We found that metric definition materially
changes the reported result. Because the published setup does not include the
fixture or benchmark script, we cannot determine how much of the remaining
difference comes from hardware, fixture, warmup coverage, or aggregation.
What the model card states
The model achieves a TTFB of 180ms (after warmup) with an RTF of 0.51 on a
single L20 GPU.
The commit that introduced these numbers does not include the benchmark script,
and the card does not state the fixture, sample count, or aggregation, so we
could not replicate the setup exactly.
What we measured
Your reference Transformers implementation, not a third-party serving stack:
|
|
| Implementation |
upstream MOSS-TTS Transformers path, source commit 58b20a0d5fcc6766658d50967a90a9d890009a46 |
| Checkpoint |
OpenMOSS-Team/MOSS-TTS-Realtime, revision 75682787d8e2fcc73faca37ba2931453ca9c4022 |
| Sampling |
card defaults: temperature=0.8, top_p=0.6, top_k=30, repetition_penalty=1.1, rep_window=50 |
| Reference audio |
official reference_en.m4a, SHA-256 ae644a5f287a812c58cb2a2257d9cf14278cf069594f28041db56c18b13c396c, normalized to PCM16 mono 24 kHz WAV, 3.84 s, SHA-256 dc4a44fd33d518030390e98ace984f2d40ed167044f45d26c2a31770fbe63943 |
| Prompt |
The quick brown fox jumps over the lazy dog., identical across runs |
| Seed |
unset |
| Hardware |
one dedicated RTX 5090, no other workload resident |
| Warmup |
preliminary requests paid the observed 172.9 s initial compile and 58.6 s shape compile. Neither delay is included in the reported 100-request distribution |
| Runs |
100 sequential requests, concurrency 1 |
Results, computed two ways over the same 100 responses:
| Metric |
Median |
p95 |
| RTF vs total returned audio duration |
0.896 |
0.972 |
| RTF vs speech duration (edge silence trimmed) |
1.139 |
1.655 |
The gap between the two rows is trailing and leading silence in the returned
audio. Leading silence alone had a p95 of 3.675 s.
Speech-span detection, so the second row is reproducible: 20 ms RMS frames, a
speech threshold of max(-60 dBFS, 2% of peak frame RMS), only leading and
trailing silence trimmed, internal pauses preserved.
Why we are asking
Reading your serving code, RTF appears to be computed against the full returned
audio duration. That is a perfectly reasonable throughput metric and we are not
suggesting it is wrong.
Because the hardware, fixture, and benchmark boundaries differ, our 0.896
neither reproduces nor contradicts 0.51. It only demonstrates that the metric
definition materially changes the conclusion on the same returned audio.
For telephony, however, two different metrics decide whether a model is usable,
and both differ from that definition:
- speech-span RTF — synthesis time against speech duration. Silence
padding inflates the denominator, so a run can be under 1.0 by the
total-duration definition while still falling behind during actual speech.
- time to first audible speech — a TTFB that ends at the first PCM chunk
is satisfied even when that chunk is silent. With a leading-silence p95 of
3.675 s, first-chunk latency and first-audible-speech latency can differ by
seconds.
These are our product metrics, not a claim that yours should change.
What would help
- The fixture behind the published numbers: prompt text(s), reference clip, and
the resulting audio-duration distribution.
- The hardware and warmup state, and whether
torch.compile was warm.
- The number of runs and the aggregation (single best run, mean, or median).
- Confirmation of whether the published RTF is computed against total returned
audio duration or against speech duration.
- Confirmation of what the 180 ms TTFB measures — first PCM chunk, or first
chunk containing speech.
- If possible, publishing the benchmark script so results are reproducible.
Happy to share our harness, raw WAVs, per-request JSONL and GPU telemetry if
that is useful for comparison.
We are evaluating MOSS-TTS-Realtime for a telephony voice agent and cannot
reproduce the published figures. We found that metric definition materially
changes the reported result. Because the published setup does not include the
fixture or benchmark script, we cannot determine how much of the remaining
difference comes from hardware, fixture, warmup coverage, or aggregation.
What the model card states
The commit that introduced these numbers does not include the benchmark script,
and the card does not state the fixture, sample count, or aggregation, so we
could not replicate the setup exactly.
What we measured
Your reference Transformers implementation, not a third-party serving stack:
MOSS-TTSTransformers path, source commit58b20a0d5fcc6766658d50967a90a9d890009a46OpenMOSS-Team/MOSS-TTS-Realtime, revision75682787d8e2fcc73faca37ba2931453ca9c4022temperature=0.8,top_p=0.6,top_k=30,repetition_penalty=1.1,rep_window=50reference_en.m4a, SHA-256ae644a5f287a812c58cb2a2257d9cf14278cf069594f28041db56c18b13c396c, normalized to PCM16 mono 24 kHz WAV, 3.84 s, SHA-256dc4a44fd33d518030390e98ace984f2d40ed167044f45d26c2a31770fbe63943The quick brown fox jumps over the lazy dog., identical across runsResults, computed two ways over the same 100 responses:
The gap between the two rows is trailing and leading silence in the returned
audio. Leading silence alone had a p95 of 3.675 s.
Speech-span detection, so the second row is reproducible: 20 ms RMS frames, a
speech threshold of
max(-60 dBFS, 2% of peak frame RMS), only leading andtrailing silence trimmed, internal pauses preserved.
Why we are asking
Reading your serving code, RTF appears to be computed against the full returned
audio duration. That is a perfectly reasonable throughput metric and we are not
suggesting it is wrong.
Because the hardware, fixture, and benchmark boundaries differ, our 0.896
neither reproduces nor contradicts 0.51. It only demonstrates that the metric
definition materially changes the conclusion on the same returned audio.
For telephony, however, two different metrics decide whether a model is usable,
and both differ from that definition:
padding inflates the denominator, so a run can be under 1.0 by the
total-duration definition while still falling behind during actual speech.
is satisfied even when that chunk is silent. With a leading-silence p95 of
3.675 s, first-chunk latency and first-audible-speech latency can differ by
seconds.
These are our product metrics, not a claim that yours should change.
What would help
the resulting audio-duration distribution.
torch.compilewas warm.audio duration or against speech duration.
chunk containing speech.
Happy to share our harness, raw WAVs, per-request JSONL and GPU telemetry if
that is useful for comparison.