-GPU_IDS="0,1,2,3,4,5,6,7"
+| Model Name | Tasks | Ckpts Tested | Min. GPU
VRAM | Comments |
+|:------------:|:---------------------:|:------------------------------------------------------------------------------------------------------------------------------------------------------:|:----------------------:|:--------------------------------------------------------------:|
+| [LTX-Video](https://huggingface.co/docs/diffusers/main/api/pipelines/ltx_video) |
| [Lightricks/LTX-Video](https://huggingface.co/Lightricks/LTX-Video) | 11 GB | Fast to train |
+| [HunyuanVideo](https://huggingface.co/docs/diffusers/main/api/pipelines/hunyuan_video) |
| [tencent/HunyuanVideo](https://huggingface.co/tencent/HunyuanVideo) | 42 GB | - |
+| [CogVideoX](https://huggingface.co/docs/diffusers/main/api/pipelines/cogvideox) |
|
- [THUDM/CogVideoX1.5-5B](https://huggingface.co/THUDM/CogVideoX1.5-5B)
- [THUDM/CogVideoX-5b](https://huggingface.co/THUDM/CogVideoX-5b)
- [THUDM/CogVideoX-2b](https://huggingface.co/THUDM/CogVideoX-2b)
| - GB | Training with multi-bucket, multi-resolution frames is supported. |
-DATA_ROOT="/path/to/dataset"
-CAPTION_COLUMN="prompts.txt"
-VIDEO_COLUMN="videos.txt"
-OUTPUT_DIR="/path/to/models/hunyuan-video/hunyuan-video-loras/hunyuan-video_cakify_500_3e-5_constant_with_warmup"
+
-ID_TOKEN="afkx"
+Note that the memory consumption in the table is reported with most of the options, discussed in [docs/training/optimizations](./docs/training/optimization.md), enabled.
-# Model arguments
-model_cmd="--model_name hunyuan_video \
- --pretrained_model_name_or_path hunyuanvideo-community/HunyuanVideo"
+If you would like to use a custom dataset, refer to the dataset preparation guide [here](./docs/dataset/README.md).
-# Dataset arguments
-dataset_cmd="--data_root $DATA_ROOT \
- --video_column $VIDEO_COLUMN \
- --caption_column $CAPTION_COLUMN \
- --id_token $ID_TOKEN \
- --video_resolution_buckets 17x512x768 49x512x768 61x512x768 \
- --caption_dropout_p 0.05"
+## Acknowledgements
-# Dataloader arguments
-dataloader_cmd="--dataloader_num_workers 0"
-
-# Diffusion arguments
-diffusion_cmd=""
-
-# Training arguments
-training_cmd="--training_type lora \
- --seed 42 \
- --mixed_precision bf16 \
- --batch_size 1 \
- --train_steps 500 \
- --rank 128 \
- --lora_alpha 128 \
- --target_modules to_q to_k to_v to_out.0 \
- --gradient_accumulation_steps 1 \
- --gradient_checkpointing \
- --checkpointing_steps 500 \
- --checkpointing_limit 2 \
- --enable_slicing \
- --enable_tiling"
-
-# Optimizer arguments
-optimizer_cmd="--optimizer adamw \
- --lr 2e-5 \
- --lr_scheduler constant_with_warmup \
- --lr_warmup_steps 100 \
- --lr_num_cycles 1 \
- --beta1 0.9 \
- --beta2 0.95 \
- --weight_decay 1e-4 \
- --epsilon 1e-8 \
- --max_grad_norm 1.0"
-
-# Validation arguments
-validation_cmd="--validation_prompts \"$ID_TOKEN A baker carefully cuts a green bell pepper cake on a white plate against a bright yellow background, followed by a strawberry cake with a similar slice of cake being cut before the interior of the bell pepper cake is revealed with the surrounding cake-to-object sequence.@@@49x512x768:::$ID_TOKEN A cake shaped like a Nutella container is carefully sliced, revealing a light interior, amidst a Nutella-themed setup, showcasing deliberate cutting and preserved details for an appetizing dessert presentation on a white base with accompanying jello and cutlery, highlighting culinary skills and creative cake designs.@@@49x512x768:::$ID_TOKEN A cake shaped like a Nutella container is carefully sliced, revealing a light interior, amidst a Nutella-themed setup, showcasing deliberate cutting and preserved details for an appetizing dessert presentation on a white base with accompanying jello and cutlery, highlighting culinary skills and creative cake designs.@@@61x512x768:::$ID_TOKEN A vibrant orange cake disguised as a Nike packaging box sits on a dark surface, meticulous in its detail and design, complete with a white swoosh and 'NIKE' logo. A person's hands, holding a knife, hover over the cake, ready to make a precise cut, amidst a simple and clean background.@@@61x512x768:::$ID_TOKEN A vibrant orange cake disguised as a Nike packaging box sits on a dark surface, meticulous in its detail and design, complete with a white swoosh and 'NIKE' logo. A person's hands, holding a knife, hover over the cake, ready to make a precise cut, amidst a simple and clean background.@@@97x512x768:::$ID_TOKEN A vibrant orange cake disguised as a Nike packaging box sits on a dark surface, meticulous in its detail and design, complete with a white swoosh and 'NIKE' logo. A person's hands, holding a knife, hover over the cake, ready to make a precise cut, amidst a simple and clean background.@@@129x512x768:::$ID_TOKEN A person with gloved hands carefully cuts a cake shaped like a Skittles bottle, beginning with a precise incision at the lid, followed by careful sequential cuts around the neck, eventually detaching the lid from the body, revealing the chocolate interior of the cake while showcasing the layered design's detail.@@@61x512x768:::$ID_TOKEN A woman with long brown hair and light skin smiles at another woman with long blonde hair. The woman with brown hair wears a black jacket and has a small, barely noticeable mole on her right cheek. The camera angle is a close-up, focused on the woman with brown hair's face. The lighting is warm and natural, likely from the setting sun, casting a soft glow on the scene. The scene appears to be real-life footage@@@61x512x768\" \
- --num_validation_videos 1 \
- --validation_steps 100"
-
-# Miscellaneous arguments
-miscellaneous_cmd="--tracker_name finetrainers-hunyuan-video \
- --output_dir $OUTPUT_DIR \
- --nccl_timeout 1800 \
- --report_to wandb"
-
-cmd="accelerate launch --config_file accelerate_configs/uncompiled_8.yaml --gpu_ids $GPU_IDS train.py \
- $model_cmd \
- $dataset_cmd \
- $dataloader_cmd \
- $diffusion_cmd \
- $training_cmd \
- $optimizer_cmd \
- $validation_cmd \
- $miscellaneous_cmd"
-
-echo "Running command: $cmd"
-eval $cmd
-echo -ne "-------------------- Finished executing script --------------------\n\n"
-```
-
-### Inference:
-
-Assuming your LoRA is saved and pushed to the HF Hub, and named `my-awesome-name/my-awesome-lora`, we can now use the finetuned model for inference:
-
-```py
-import torch
-from diffusers import HunyuanVideoPipeline
-
-import torch
-from diffusers import HunyuanVideoPipeline, HunyuanVideoTransformer3DModel
-from diffusers.utils import export_to_video
-
-model_id = "hunyuanvideo-community/HunyuanVideo"
-transformer = HunyuanVideoTransformer3DModel.from_pretrained(
- model_id, subfolder="transformer", torch_dtype=torch.bfloat16
-)
-pipe = HunyuanVideoPipeline.from_pretrained(model_id, transformer=transformer, torch_dtype=torch.float16)
-pipe.load_lora_weights("my-awesome-name/my-awesome-lora", adapter_name="hunyuanvideo-lora")
-pipe.set_adapters(["hunyuanvideo-lora"], [0.6])
-pipe.vae.enable_tiling()
-pipe.to("cuda")
-
-output = pipe(
- prompt="A cat walks on the grass, realistic",
- height=320,
- width=512,
- num_frames=61,
- num_inference_steps=30,
-).frames[0]
-export_to_video(output, "output.mp4", fps=15)
-```
-
-### Memory Usage
-
-LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolutions, **without precomputation**:
-
-```
-Training configuration: {
- "trainable parameters": 163577856,
- "total samples": 69,
- "train epochs": 1,
- "train steps": 10,
- "batches per device": 1,
- "total batches observed per epoch": 69,
- "train batch size": 1,
- "gradient accumulation steps": 1
-}
-```
-
-| stage | memory_allocated | max_memory_reserved |
-|:-----------------------:|:----------------:|:-------------------:|
-| before training start | 38.889 | 39.020 |
-| before validation start | 39.747 | 56.266 |
-| after validation end | 39.748 | 58.385 |
-| after epoch 1 | 39.748 | 40.910 |
-| after training end | 25.288 | 40.910 |
-
-Note: requires about `59` GB of VRAM when validation is performed.
-
-LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolutions, **with precomputation**:
-
-```
-Training configuration: {
- "trainable parameters": 163577856,
- "total samples": 1,
- "train epochs": 10,
- "train steps": 10,
- "batches per device": 1,
- "total batches observed per epoch": 1,
- "train batch size": 1,
- "gradient accumulation steps": 1
-}
-```
-
-| stage | memory_allocated | max_memory_reserved |
-|:-----------------------------:|:----------------:|:-------------------:|
-| after precomputing conditions | 14.232 | 14.461 |
-| after precomputing latents | 14.717 | 17.244 |
-| before training start | 24.195 | 26.039 |
-| after epoch 1 | 24.83 | 42.387 |
-| before validation start | 24.842 | 42.387 |
-| after validation end | 39.558 | 46.947 |
-| after training end | 24.842 | 41.039 |
-
-Note: requires about `47` GB of VRAM with validation. If validation is not performed, the memory usage is reduced to about `42` GB.
-
-