diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index 6c5322b..3ccd3ca 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -77,7 +77,18 @@ both cards land then too.) See `docs/CONCURRENCY.md` for the model. Experiment: two resident instances, sequential pair vs concurrent pair of identical warm t2v jobs. -RESULTS_PENDING +Measured (pruned int8, t2v 864×480 5 s, both instances warm): + +| Mode | Total for 2 videos | Per-video effective | +|-------------------------------|--------------------|---------------------| +| Sequential (one instance) | **128 s** | 64 s | +| Concurrent (two instances) | **145 s** | 72.5 s | + +Concurrent was **13% slower** for the same work — each co-scheduled job ran +142 s vs 63 s alone (2.27×). There is no pipeline-overlap win; cache and +scheduler contention make co-scheduling strictly worse. **One worker per GPU, +queue in front, scale horizontally.** (This also removes any argument for +MPS/MIG partitioning for throughput.) ## Takeaways diff --git a/docs/CONCURRENCY.md b/docs/CONCURRENCY.md index 5e3d0c9..7f9a808 100644 --- a/docs/CONCURRENCY.md +++ b/docs/CONCURRENCY.md @@ -23,16 +23,18 @@ which is near-linear in time (2 outputs ≈ 2× one output's duration). **Utilization caps (MPS / MIG):** the B200 supports both — MPS SM partitioning (`CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`) and MIG (e.g. 2 × 90 GB instances, each big enough for the pruned-int8 stack). But partitioning divides FLOPs, it -doesn't create them: two half-GPU workers each generate ~2× slower. The only -real gain from two workers on one GPU is **pipeline overlap** — worker A's -text-encode/VAE-decode running under worker B's denoise. Expected benefit: -**+10–20% throughput at the cost of ~2× per-video latency**. Planned -experiment on this pod (second ComfyUI instance, concurrent vs sequential -pairs) will measure the real number. +doesn't create them: two half-GPU workers each generate ~2× slower. -**Bottom line:** concurrency = number of GPUs (or GPU partitions, at pro-rata -speed). Scale horizontally with one resident worker per GPU; don't co-schedule -for throughput unless the overlap experiment surprises us. +**MEASURED (2026-08-06, two ComfyUI instances on the B200):** a sequential +pair of warm 5 s t2v jobs took **128 s**; the same pair run concurrently took +**145 s** — co-scheduling was **13% slower** for identical work (each +concurrent job: 142 s vs 63 s alone). The hypothesized 10–20% pipeline-overlap +gain does not exist in practice; contention eats it and more. Details in +`docs/BENCHMARKS.md`. + +**Bottom line (now empirical):** concurrency = number of GPUs. One resident +worker per GPU, queue in front, scale horizontally. Never co-schedule; skip +MPS/MIG for throughput purposes. ## 2. Cost per video across GPU tiers