From fbc68bb749e2404350a0b2a563df5f1cd713d46d Mon Sep 17 00:00:00 2001 From: Hanashi Date: Thu, 6 Aug 2026 14:27:32 -0400 Subject: [PATCH] =?UTF-8?q?concurrency=20experiment=20measured:=20co-sched?= =?UTF-8?q?uling=2013%=20slower=20than=20sequential=20=E2=80=94=20one=20wo?= =?UTF-8?q?rker=20per=20GPU=20confirmed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- docs/BENCHMARKS.md | 13 ++++++++++++- docs/CONCURRENCY.md | 20 +++++++++++--------- 2 files changed, 23 insertions(+), 10 deletions(-) diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index 6c5322b..3ccd3ca 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -77,7 +77,18 @@ both cards land then too.) See `docs/CONCURRENCY.md` for the model. Experiment: two resident instances, sequential pair vs concurrent pair of identical warm t2v jobs. -RESULTS_PENDING +Measured (pruned int8, t2v 864×480 5 s, both instances warm): + +| Mode | Total for 2 videos | Per-video effective | +|-------------------------------|--------------------|---------------------| +| Sequential (one instance) | **128 s** | 64 s | +| Concurrent (two instances) | **145 s** | 72.5 s | + +Concurrent was **13% slower** for the same work — each co-scheduled job ran +142 s vs 63 s alone (2.27×). There is no pipeline-overlap win; cache and +scheduler contention make co-scheduling strictly worse. **One worker per GPU, +queue in front, scale horizontally.** (This also removes any argument for +MPS/MIG partitioning for throughput.) ## Takeaways diff --git a/docs/CONCURRENCY.md b/docs/CONCURRENCY.md index 5e3d0c9..7f9a808 100644 --- a/docs/CONCURRENCY.md +++ b/docs/CONCURRENCY.md @@ -23,16 +23,18 @@ which is near-linear in time (2 outputs ≈ 2× one output's duration). **Utilization caps (MPS / MIG):** the B200 supports both — MPS SM partitioning (`CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`) and MIG (e.g. 2 × 90 GB instances, each big enough for the pruned-int8 stack). But partitioning divides FLOPs, it -doesn't create them: two half-GPU workers each generate ~2× slower. The only -real gain from two workers on one GPU is **pipeline overlap** — worker A's -text-encode/VAE-decode running under worker B's denoise. Expected benefit: -**+10–20% throughput at the cost of ~2× per-video latency**. Planned -experiment on this pod (second ComfyUI instance, concurrent vs sequential -pairs) will measure the real number. +doesn't create them: two half-GPU workers each generate ~2× slower. -**Bottom line:** concurrency = number of GPUs (or GPU partitions, at pro-rata -speed). Scale horizontally with one resident worker per GPU; don't co-schedule -for throughput unless the overlap experiment surprises us. +**MEASURED (2026-08-06, two ComfyUI instances on the B200):** a sequential +pair of warm 5 s t2v jobs took **128 s**; the same pair run concurrently took +**145 s** — co-scheduling was **13% slower** for identical work (each +concurrent job: 142 s vs 63 s alone). The hypothesized 10–20% pipeline-overlap +gain does not exist in practice; contention eats it and more. Details in +`docs/BENCHMARKS.md`. + +**Bottom line (now empirical):** concurrency = number of GPUs. One resident +worker per GPU, queue in front, scale horizontally. Never co-schedule; skip +MPS/MIG for throughput purposes. ## 2. Cost per video across GPU tiers