mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
concurrency experiment measured: co-scheduling 13% slower than sequential — one worker per GPU confirmed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
+12
-1
@@ -77,7 +77,18 @@ both cards land then too.)
|
||||
See `docs/CONCURRENCY.md` for the model. Experiment: two resident instances,
|
||||
sequential pair vs concurrent pair of identical warm t2v jobs.
|
||||
|
||||
RESULTS_PENDING
|
||||
Measured (pruned int8, t2v 864×480 5 s, both instances warm):
|
||||
|
||||
| Mode | Total for 2 videos | Per-video effective |
|
||||
|-------------------------------|--------------------|---------------------|
|
||||
| Sequential (one instance) | **128 s** | 64 s |
|
||||
| Concurrent (two instances) | **145 s** | 72.5 s |
|
||||
|
||||
Concurrent was **13% slower** for the same work — each co-scheduled job ran
|
||||
142 s vs 63 s alone (2.27×). There is no pipeline-overlap win; cache and
|
||||
scheduler contention make co-scheduling strictly worse. **One worker per GPU,
|
||||
queue in front, scale horizontally.** (This also removes any argument for
|
||||
MPS/MIG partitioning for throughput.)
|
||||
|
||||
## Takeaways
|
||||
|
||||
|
||||
+11
-9
@@ -23,16 +23,18 @@ which is near-linear in time (2 outputs ≈ 2× one output's duration).
|
||||
**Utilization caps (MPS / MIG):** the B200 supports both — MPS SM partitioning
|
||||
(`CUDA_MPS_ACTIVE_THREAD_PERCENTAGE`) and MIG (e.g. 2 × 90 GB instances, each
|
||||
big enough for the pruned-int8 stack). But partitioning divides FLOPs, it
|
||||
doesn't create them: two half-GPU workers each generate ~2× slower. The only
|
||||
real gain from two workers on one GPU is **pipeline overlap** — worker A's
|
||||
text-encode/VAE-decode running under worker B's denoise. Expected benefit:
|
||||
**+10–20% throughput at the cost of ~2× per-video latency**. Planned
|
||||
experiment on this pod (second ComfyUI instance, concurrent vs sequential
|
||||
pairs) will measure the real number.
|
||||
doesn't create them: two half-GPU workers each generate ~2× slower.
|
||||
|
||||
**Bottom line:** concurrency = number of GPUs (or GPU partitions, at pro-rata
|
||||
speed). Scale horizontally with one resident worker per GPU; don't co-schedule
|
||||
for throughput unless the overlap experiment surprises us.
|
||||
**MEASURED (2026-08-06, two ComfyUI instances on the B200):** a sequential
|
||||
pair of warm 5 s t2v jobs took **128 s**; the same pair run concurrently took
|
||||
**145 s** — co-scheduling was **13% slower** for identical work (each
|
||||
concurrent job: 142 s vs 63 s alone). The hypothesized 10–20% pipeline-overlap
|
||||
gain does not exist in practice; contention eats it and more. Details in
|
||||
`docs/BENCHMARKS.md`.
|
||||
|
||||
**Bottom line (now empirical):** concurrency = number of GPUs. One resident
|
||||
worker per GPU, queue in front, scale horizontally. Never co-schedule; skip
|
||||
MPS/MIG for throughput purposes.
|
||||
|
||||
## 2. Cost per video across GPU tiers
|
||||
|
||||
|
||||
Reference in New Issue
Block a user