concurrency experiment measured: co-scheduling 13% slower than sequential — one worker per GPU confirmed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Hanashi
2026-08-06 14:27:32 -04:00
parent 24134b9653
commit fbc68bb749
2 changed files with 23 additions and 10 deletions
+12 -1
View File
@@ -77,7 +77,18 @@ both cards land then too.)
See `docs/CONCURRENCY.md` for the model. Experiment: two resident instances,
sequential pair vs concurrent pair of identical warm t2v jobs.
RESULTS_PENDING
Measured (pruned int8, t2v 864×480 5 s, both instances warm):
| Mode | Total for 2 videos | Per-video effective |
|-------------------------------|--------------------|---------------------|
| Sequential (one instance) | **128 s** | 64 s |
| Concurrent (two instances) | **145 s** | 72.5 s |
Concurrent was **13% slower** for the same work — each co-scheduled job ran
142 s vs 63 s alone (2.27×). There is no pipeline-overlap win; cache and
scheduler contention make co-scheduling strictly worse. **One worker per GPU,
queue in front, scale horizontally.** (This also removes any argument for
MPS/MIG partitioning for throughput.)
## Takeaways
+11 -9
View File
@@ -23,16 +23,18 @@ which is near-linear in time (2 outputs ≈ 2× one output's duration).
**Utilization caps (MPS / MIG):** the B200 supports both — MPS SM partitioning
(`CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`) and MIG (e.g. 2 × 90 GB instances, each
big enough for the pruned-int8 stack). But partitioning divides FLOPs, it
doesn't create them: two half-GPU workers each generate ~2× slower. The only
real gain from two workers on one GPU is **pipeline overlap** — worker A's
text-encode/VAE-decode running under worker B's denoise. Expected benefit:
**+10–20% throughput at the cost of ~2× per-video latency**. Planned
experiment on this pod (second ComfyUI instance, concurrent vs sequential
pairs) will measure the real number.
doesn't create them: two half-GPU workers each generate ~2× slower.
**Bottom line:** concurrency = number of GPUs (or GPU partitions, at pro-rata
speed). Scale horizontally with one resident worker per GPU; don't co-schedule
for throughput unless the overlap experiment surprises us.
**MEASURED (2026-08-06, two ComfyUI instances on the B200):** a sequential
pair of warm 5 s t2v jobs took **128 s**; the same pair run concurrently took
**145 s** — co-scheduling was **13% slower** for identical work (each
concurrent job: 142 s vs 63 s alone). The hypothesized 10–20% pipeline-overlap
gain does not exist in practice; contention eats it and more. Details in
`docs/BENCHMARKS.md`.
**Bottom line (now empirical):** concurrency = number of GPUs. One resident
worker per GPU, queue in front, scale horizontally. Never co-schedule; skip
MPS/MIG for throughput purposes.
## 2. Cost per video across GPU tiers