← Byron Hsu

Page 64 improved this high-pressure Qwen run

One A/B pair · Qwen3-1.7B · 16 conversations × 6 turns · 96 successful requests each · both runs evicted GPU cache and restored CPU cache.

A: page 1 timeline + metricsB: page 64 timeline + metrics

MetricA · page 1B · page 64Change
Successful requests96 / 9696 / 96No errors
Late-turn TTFT p952,502.49 ms20.73 ms99.17% lower
Overall TTFT p95761.9 ms20.73 ms97.3% lower
Late-turn mean time per output token10.56 ms2.33 ms77.9% lower
Whole-run output throughput662.8 tokens/s821.1 tokens/s23.9% higher
Elapsed workload time18.54 s14.96 s
Last-scrape GPU evictions92,200 tokens66,112 tokensBoth evicted
Last-scrape CPU → GPU restores76,088 tokens49,984 tokensBoth restored

Late turns = 4–5 (zero-based). Single-run evidence, not a universal guarantee. Server-generated RNG seeds differed; 85/96 output hashes matched; input construction, context lengths, requested tool delays, and output lengths matched. Token timing is client-observed, not kernel timing. No trace-based claim about the underlying kernel or synchronization cost.

Time-based and matched-turn comparison of TTFT, token time, throughput, and CPU restored tokens