mirror of
https://github.com/storytold/FineTrainers-Conditioning.git
synced 2026-10-09 00:09:45 +00:00
[docs] refactor docs for easier info parsing (#175)
* refactor docs for easier info parsing * refactor readme * anonym paths * updates * remove notes. * add a toc. * minor typos. * move cog.md -> cogvideox.md * add a note about the model-specific docs in training readme. * add memory usage for CogVideoX. Co-authored-by: a-r-r-o-w <contact.aryanvs@gmail.com> * change to 5b from 2b for CogVideoX. Co-authored-by: a-r-r-o-w <contact.aryanvs@gmail.com> * more appropriate names. * add headers to the model docs. * fix adapter name * minor * updates * fix cog training command example --------- Co-authored-by: a-r-r-o-w <contact.aryanvs@gmail.com>
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
## Dataset Format
|
||||
|
||||
### Prompt Dataset Requirements
|
||||
|
||||
Create a `prompt.txt` file, which should contain prompts separated by lines. Please note that the prompts must be in English, and it is recommended to use the [prompt refinement script](https://github.com/THUDM/CogVideo/blob/main/inference/convert_demo.py) for better prompts. Alternatively, you can use [CogVideo-caption](https://huggingface.co/THUDM/cogvlm2-llama3-caption) for data annotation:
|
||||
|
||||
```
|
||||
A black and white animated sequence featuring a rabbit, named Rabbity Ribfried, and an anthropomorphic goat in a musical, playful environment, showcasing their evolving interaction.
|
||||
A black and white animated sequence on a ship’s deck features a bulldog character, named Bully Bulldoger, showcasing exaggerated facial expressions and body language...
|
||||
...
|
||||
```
|
||||
|
||||
### Video Dataset Requirements
|
||||
|
||||
The framework supports resolutions and frame counts that meet the following conditions:
|
||||
|
||||
- **Supported Resolutions (Width * Height)**:
|
||||
- Any resolution as long as it is divisible by 32. For example, `720 * 480`, `1920 * 1020`, etc.
|
||||
|
||||
- **Supported Frame Counts (Frames)**:
|
||||
- Must be `4 * k` or `4 * k + 1` (example: 16, 32, 49, 81)
|
||||
|
||||
It is recommended to place all videos in a single folder.
|
||||
|
||||
Next, create a `videos.txt` file. The `videos.txt` file should contain the video file paths, separated by lines. Please note that the paths must be relative to the `--data_root` directory. The format is as follows:
|
||||
|
||||
```
|
||||
videos/00000.mp4
|
||||
videos/00001.mp4
|
||||
...
|
||||
```
|
||||
|
||||
For developers interested in more details, you can refer to the relevant `BucketSampler` code.
|
||||
|
||||
### Dataset Structure
|
||||
|
||||
Your dataset structure should look like this. Running the `tree` command, you should see:
|
||||
|
||||
```
|
||||
dataset
|
||||
├── prompt.txt
|
||||
├── videos.txt
|
||||
├── videos
|
||||
├── videos/00000.mp4
|
||||
├── videos/00001.mp4
|
||||
├── ...
|
||||
```
|
||||
|
||||
### Using the Dataset
|
||||
|
||||
When using this format, the `--caption_column` should be set to `prompt.txt`, and the `--video_column` should be set to `videos.txt`. If your data is stored in a CSV file, you can also specify `--dataset_file` as the path to the CSV file, with `--caption_column` and `--video_column` set to the actual column names in the CSV. Please refer to the [test_dataset](../tests/test_dataset.py) file for some simple examples.
|
||||
|
||||
For instance, you can fine-tune using [this](https://huggingface.co/datasets/Wild-Heart/Disney-VideoGeneration-Dataset) Disney dataset. The download can be done via the 🤗 Hugging Face CLI:
|
||||
|
||||
```
|
||||
huggingface-cli download --repo-type dataset Wild-Heart/Disney-VideoGeneration-Dataset --local-dir video-dataset-disney
|
||||
```
|
||||
|
||||
This dataset has been prepared in the expected format and can be used directly. However, directly using the video dataset may cause Out of Memory (OOM) issues on GPUs with smaller VRAM because it requires loading the [VAE](https://huggingface.co/THUDM/CogVideoX-5b/tree/main/vae) (which encodes videos into latent space) and the large [T5-XXL](https://huggingface.co/google/t5-v1_1-xxl/) text encoder. To reduce memory usage, you can use the `training/prepare_dataset.py` script to precompute latents and embeddings.
|
||||
|
||||
Fill or modify the parameters in `prepare_dataset.sh` and execute it to get precomputed latents and embeddings (make sure to specify `--save_latents_and_embeddings` to save the precomputed artifacts). If preparing for image-to-video training, make sure to pass `--save_image_latents`, which encodes and saves image latents along with videos. When using these artifacts during training, ensure that you specify the `--load_tensors` flag, or else the videos will be used directly, requiring the text encoder and VAE to be loaded. The script also supports PyTorch DDP so that large datasets can be encoded in parallel across multiple GPUs (modify the `NUM_GPUS` parameter).
|
||||
@@ -0,0 +1,19 @@
|
||||
This directory contains the training-related specifications for all the models we support in `finetrainers`. Each model page has:
|
||||
|
||||
* an example training command
|
||||
* inference example
|
||||
* numbers on memory consumption
|
||||
|
||||
By default, we don't include any validation-related arguments in the example training commands. To enable validation inference, one can pass:
|
||||
|
||||
```diff
|
||||
+ --validation_prompts "$ID_TOKEN A black and white animated scene unfolds with an anthropomorphic goat surrounded by musical notes and symbols, suggesting a playful environment. Mickey Mouse appears, leaning forward in curiosity as the goat remains still. The goat then engages with Mickey, who bends down to converse or react. The dynamics shift as Mickey grabs the goat, potentially in surprise or playfulness, amidst a minimalistic background. The scene captures the evolving relationship between the two characters in a whimsical, animated setting, emphasizing their interactions and emotions.@@@49x512x768:::$ID_TOKEN A woman with long brown hair and light skin smiles at another woman with long blonde hair. The woman with brown hair wears a black jacket and has a small, barely noticeable mole on her right cheek. The camera angle is a close-up, focused on the woman with brown hair's face. The lighting is warm and natural, likely from the setting sun, casting a soft glow on the scene. The scene appears to be real-life footage@@@49x512x768" \
|
||||
+ --num_validation_videos 1 \
|
||||
+ --validation_steps 100
|
||||
```
|
||||
|
||||
## Model-specific docs
|
||||
|
||||
* [CogVideoX](./cogvideox.md)
|
||||
* [LTX-Video](./ltx_video.md)
|
||||
* [HunyuanVideo](./hunyuan_video.md)
|
||||
@@ -0,0 +1,134 @@
|
||||
# CogVideoX
|
||||
|
||||
## Training
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
export WANDB_MODE="offline"
|
||||
export NCCL_P2P_DISABLE=1
|
||||
export TORCH_NCCL_ENABLE_MONITORING=0
|
||||
export FINETRAINERS_LOG_LEVEL=DEBUG
|
||||
|
||||
GPU_IDS="0,1"
|
||||
|
||||
DATA_ROOT="/path/to/dataset"
|
||||
CAPTION_COLUMN="prompt.txt"
|
||||
VIDEO_COLUMN="videos.txt"
|
||||
OUTPUT_DIR="/path/to/models/cog/"
|
||||
ID_TOKEN="BW_STYLE"
|
||||
|
||||
# Model arguments
|
||||
model_cmd="--model_name cogvideox \
|
||||
--pretrained_model_name_or_path THUDM/CogVideoX-5b"
|
||||
|
||||
# Dataset arguments
|
||||
dataset_cmd="--data_root $DATA_ROOT \
|
||||
--video_column $VIDEO_COLUMN \
|
||||
--caption_column $CAPTION_COLUMN \
|
||||
--id_token $ID_TOKEN \
|
||||
--video_resolution_buckets 49x480x720 \
|
||||
--caption_dropout_p 0.05"
|
||||
|
||||
# Dataloader arguments
|
||||
dataloader_cmd="--dataloader_num_workers 4"
|
||||
|
||||
# Training arguments
|
||||
training_cmd="--training_type lora \
|
||||
--seed 42 \
|
||||
--mixed_precision bf16 \
|
||||
--batch_size 1 \
|
||||
--precompute_conditions \
|
||||
--train_steps 1000 \
|
||||
--rank 128 \
|
||||
--lora_alpha 128 \
|
||||
--target_modules to_q to_k to_v to_out.0 \
|
||||
--gradient_accumulation_steps 1 \
|
||||
--gradient_checkpointing \
|
||||
--checkpointing_steps 200 \
|
||||
--checkpointing_limit 2 \
|
||||
--resume_from_checkpoint=latest \
|
||||
--enable_slicing \
|
||||
--enable_tiling"
|
||||
|
||||
# Optimizer arguments
|
||||
optimizer_cmd="--optimizer adamw \
|
||||
--use_8bit_bnb \
|
||||
--lr 3e-5 \
|
||||
--lr_scheduler constant_with_warmup \
|
||||
--lr_warmup_steps 100 \
|
||||
--lr_num_cycles 1 \
|
||||
--beta1 0.9 \
|
||||
--beta2 0.95 \
|
||||
--weight_decay 1e-4 \
|
||||
--epsilon 1e-8 \
|
||||
--max_grad_norm 1.0"
|
||||
|
||||
# Miscellaneous arguments
|
||||
miscellaneous_cmd="--tracker_name finetrainers-cog \
|
||||
--output_dir $OUTPUT_DIR \
|
||||
--nccl_timeout 1800 \
|
||||
--report_to wandb"
|
||||
|
||||
cmd="accelerate launch --config_file accelerate_configs/deepspeed.yaml --gpu_ids $GPU_IDS train.py \
|
||||
$model_cmd \
|
||||
$dataset_cmd \
|
||||
$dataloader_cmd \
|
||||
$training_cmd \
|
||||
$optimizer_cmd \
|
||||
$miscellaneous_cmd"
|
||||
|
||||
echo "Running command: $cmd"
|
||||
eval $cmd
|
||||
echo -ne "-------------------- Finished executing script --------------------\n\n"
|
||||
```
|
||||
|
||||
## Memory Usage
|
||||
|
||||
LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x480x720` resolutions, **with precomputation**:
|
||||
|
||||
```
|
||||
Training configuration: {
|
||||
"trainable parameters": 132120576,
|
||||
"total samples": 69,
|
||||
"train epochs": 1,
|
||||
"train steps": 10,
|
||||
"batches per device": 1,
|
||||
"total batches observed per epoch": 69,
|
||||
"train batch size": 1,
|
||||
"gradient accumulation steps": 1
|
||||
}
|
||||
```
|
||||
|
||||
| stage | memory_allocated | max_memory_reserved |
|
||||
|:-----------------------------:|:-----------------:|:-------------------:|
|
||||
| after precomputing conditions | 8.880 | 8.941 |
|
||||
| after precomputing latents | 9.300 | 12.441 |
|
||||
| before training start | 10.622 | 20.701 |
|
||||
| after epoch 1 | 11.145 | 20.701 |
|
||||
| before validation start | 11.145 | 20.702 |
|
||||
| after validation end | 11.145 | 28.324 |
|
||||
| after training end | 11.144 | 11.592 |
|
||||
|
||||
## Inference
|
||||
|
||||
Assuming your LoRA is saved and pushed to the HF Hub, and named `my-awesome-name/my-awesome-lora`, we can now use the finetuned model for inference:
|
||||
|
||||
```diff
|
||||
import torch
|
||||
from diffusers import CogVideoXPipeline
|
||||
from diffusers.utils import export_to_video
|
||||
|
||||
pipe = CogVideoXPipeline.from_pretrained(
|
||||
"THUDM/CogVideoX-5b", torch_dtype=torch.bfloat16
|
||||
).to("cuda")
|
||||
+ pipe.load_lora_weights("my-awesome-name/my-awesome-lora", adapter_name="cogvideox-lora")
|
||||
+ pipe.set_adapters(["cogvideox-lora"], [0.75])
|
||||
|
||||
video = pipe("<my-awesome-prompt>").frames[0]
|
||||
export_to_video(video, "output.mp4")
|
||||
```
|
||||
|
||||
You can refer to the following guides to know more about performing LoRA inference in `diffusers`:
|
||||
|
||||
* [Load LoRAs for inference](https://huggingface.co/docs/diffusers/main/en/tutorials/using_peft_for_inference)
|
||||
* [Merge LoRAs](https://huggingface.co/docs/diffusers/main/en/using-diffusers/merge_loras)
|
||||
@@ -0,0 +1,177 @@
|
||||
# HunyuanVideo
|
||||
|
||||
## Training
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
|
||||
export WANDB_MODE="offline"
|
||||
export NCCL_P2P_DISABLE=1
|
||||
export TORCH_NCCL_ENABLE_MONITORING=0
|
||||
export FINETRAINERS_LOG_LEVEL=DEBUG
|
||||
|
||||
GPU_IDS="0,1"
|
||||
|
||||
DATA_ROOT="/path/to/dataset"
|
||||
CAPTION_COLUMN="prompts.txt"
|
||||
VIDEO_COLUMN="videos.txt"
|
||||
OUTPUT_DIR="/path/to/models/hunyuan-video/"
|
||||
|
||||
ID_TOKEN="afkx"
|
||||
|
||||
# Model arguments
|
||||
model_cmd="--model_name hunyuan_video \
|
||||
--pretrained_model_name_or_path hunyuanvideo-community/HunyuanVideo"
|
||||
|
||||
# Dataset arguments
|
||||
dataset_cmd="--data_root $DATA_ROOT \
|
||||
--video_column $VIDEO_COLUMN \
|
||||
--caption_column $CAPTION_COLUMN \
|
||||
--id_token $ID_TOKEN \
|
||||
--video_resolution_buckets 17x512x768 49x512x768 61x512x768 \
|
||||
--caption_dropout_p 0.05"
|
||||
|
||||
# Dataloader arguments
|
||||
dataloader_cmd="--dataloader_num_workers 0"
|
||||
|
||||
# Diffusion arguments
|
||||
diffusion_cmd=""
|
||||
|
||||
# Training arguments
|
||||
training_cmd="--training_type lora \
|
||||
--seed 42 \
|
||||
--mixed_precision bf16 \
|
||||
--batch_size 1 \
|
||||
--train_steps 500 \
|
||||
--rank 128 \
|
||||
--lora_alpha 128 \
|
||||
--target_modules to_q to_k to_v to_out.0 \
|
||||
--gradient_accumulation_steps 1 \
|
||||
--gradient_checkpointing \
|
||||
--checkpointing_steps 500 \
|
||||
--checkpointing_limit 2 \
|
||||
--enable_slicing \
|
||||
--enable_tiling"
|
||||
|
||||
# Optimizer arguments
|
||||
optimizer_cmd="--optimizer adamw \
|
||||
--lr 2e-5 \
|
||||
--lr_scheduler constant_with_warmup \
|
||||
--lr_warmup_steps 100 \
|
||||
--lr_num_cycles 1 \
|
||||
--beta1 0.9 \
|
||||
--beta2 0.95 \
|
||||
--weight_decay 1e-4 \
|
||||
--epsilon 1e-8 \
|
||||
--max_grad_norm 1.0"
|
||||
|
||||
# Miscellaneous arguments
|
||||
miscellaneous_cmd="--tracker_name finetrainers-hunyuan-video \
|
||||
--output_dir $OUTPUT_DIR \
|
||||
--nccl_timeout 1800 \
|
||||
--report_to wandb"
|
||||
|
||||
cmd="accelerate launch --config_file accelerate_configs/uncompiled_8.yaml --gpu_ids $GPU_IDS train.py \
|
||||
$model_cmd \
|
||||
$dataset_cmd \
|
||||
$dataloader_cmd \
|
||||
$diffusion_cmd \
|
||||
$training_cmd \
|
||||
$optimizer_cmd \
|
||||
$miscellaneous_cmd"
|
||||
|
||||
echo "Running command: $cmd"
|
||||
eval $cmd
|
||||
echo -ne "-------------------- Finished executing script --------------------\n\n"
|
||||
```
|
||||
|
||||
## Memory Usage
|
||||
|
||||
LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolutions, **without precomputation**:
|
||||
|
||||
```
|
||||
Training configuration: {
|
||||
"trainable parameters": 163577856,
|
||||
"total samples": 69,
|
||||
"train epochs": 1,
|
||||
"train steps": 10,
|
||||
"batches per device": 1,
|
||||
"total batches observed per epoch": 69,
|
||||
"train batch size": 1,
|
||||
"gradient accumulation steps": 1
|
||||
}
|
||||
```
|
||||
|
||||
| stage | memory_allocated | max_memory_reserved |
|
||||
|:-----------------------:|:----------------:|:-------------------:|
|
||||
| before training start | 38.889 | 39.020 |
|
||||
| before validation start | 39.747 | 56.266 |
|
||||
| after validation end | 39.748 | 58.385 |
|
||||
| after epoch 1 | 39.748 | 40.910 |
|
||||
| after training end | 25.288 | 40.910 |
|
||||
|
||||
Note: requires about `59` GB of VRAM when validation is performed.
|
||||
|
||||
LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolutions, **with precomputation**:
|
||||
|
||||
```
|
||||
Training configuration: {
|
||||
"trainable parameters": 163577856,
|
||||
"total samples": 1,
|
||||
"train epochs": 10,
|
||||
"train steps": 10,
|
||||
"batches per device": 1,
|
||||
"total batches observed per epoch": 1,
|
||||
"train batch size": 1,
|
||||
"gradient accumulation steps": 1
|
||||
}
|
||||
```
|
||||
|
||||
| stage | memory_allocated | max_memory_reserved |
|
||||
|:-----------------------------:|:----------------:|:-------------------:|
|
||||
| after precomputing conditions | 14.232 | 14.461 |
|
||||
| after precomputing latents | 14.717 | 17.244 |
|
||||
| before training start | 24.195 | 26.039 |
|
||||
| after epoch 1 | 24.83 | 42.387 |
|
||||
| before validation start | 24.842 | 42.387 |
|
||||
| after validation end | 39.558 | 46.947 |
|
||||
| after training end | 24.842 | 41.039 |
|
||||
|
||||
Note: requires about `47` GB of VRAM with validation. If validation is not performed, the memory usage is reduced to about `42` GB.
|
||||
|
||||
## Inference
|
||||
|
||||
Assuming your LoRA is saved and pushed to the HF Hub, and named `my-awesome-name/my-awesome-lora`, we can now use the finetuned model for inference:
|
||||
|
||||
```py
|
||||
import torch
|
||||
from diffusers import HunyuanVideoPipeline
|
||||
|
||||
import torch
|
||||
from diffusers import HunyuanVideoPipeline, HunyuanVideoTransformer3DModel
|
||||
from diffusers.utils import export_to_video
|
||||
|
||||
model_id = "hunyuanvideo-community/HunyuanVideo"
|
||||
transformer = HunyuanVideoTransformer3DModel.from_pretrained(
|
||||
model_id, subfolder="transformer", torch_dtype=torch.bfloat16
|
||||
)
|
||||
pipe = HunyuanVideoPipeline.from_pretrained(model_id, transformer=transformer, torch_dtype=torch.float16)
|
||||
pipe.load_lora_weights("my-awesome-name/my-awesome-lora", adapter_name="hunyuanvideo-lora")
|
||||
pipe.set_adapters(["hunyuanvideo-lora"], [0.6])
|
||||
pipe.vae.enable_tiling()
|
||||
pipe.to("cuda")
|
||||
|
||||
output = pipe(
|
||||
prompt="A cat walks on the grass, realistic",
|
||||
height=320,
|
||||
width=512,
|
||||
num_frames=61,
|
||||
num_inference_steps=30,
|
||||
).frames[0]
|
||||
export_to_video(output, "output.mp4", fps=15)
|
||||
```
|
||||
|
||||
You can refer to the following guides to know more about performing LoRA inference in `diffusers`:
|
||||
|
||||
* [Load LoRAs for inference](https://huggingface.co/docs/diffusers/main/en/tutorials/using_peft_for_inference)
|
||||
* [Merge LoRAs](https://huggingface.co/docs/diffusers/main/en/using-diffusers/merge_loras)
|
||||
@@ -0,0 +1,165 @@
|
||||
# LTX-Video
|
||||
|
||||
## Training
|
||||
|
||||
Provided you have a dataset:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
export WANDB_MODE="offline"
|
||||
export NCCL_P2P_DISABLE=1
|
||||
export TORCH_NCCL_ENABLE_MONITORING=0
|
||||
export FINETRAINERS_LOG_LEVEL=DEBUG
|
||||
|
||||
GPU_IDS="0,1"
|
||||
|
||||
DATA_ROOT="/path/to/dataset"
|
||||
CAPTION_COLUMN="prompts.txt"
|
||||
VIDEO_COLUMN="videos.txt"
|
||||
OUTPUT_DIR="/path/to/models/ltx-video/"
|
||||
|
||||
ID_TOKEN="BW_STYLE"
|
||||
|
||||
# Model arguments
|
||||
model_cmd="--model_name ltx_video \
|
||||
--pretrained_model_name_or_path Lightricks/LTX-Video"
|
||||
|
||||
# Dataset arguments
|
||||
dataset_cmd="--data_root $DATA_ROOT \
|
||||
--video_column $VIDEO_COLUMN \
|
||||
--caption_column $CAPTION_COLUMN \
|
||||
--id_token $ID_TOKEN \
|
||||
--video_resolution_buckets 49x512x768 \
|
||||
--caption_dropout_p 0.05"
|
||||
|
||||
# Dataloader arguments
|
||||
dataloader_cmd="--dataloader_num_workers 0"
|
||||
|
||||
# Diffusion arguments
|
||||
diffusion_cmd="--flow_resolution_shifting"
|
||||
|
||||
# Training arguments
|
||||
training_cmd="--training_type lora \
|
||||
--seed 42 \
|
||||
--mixed_precision bf16 \
|
||||
--batch_size 1 \
|
||||
--train_steps 1200 \
|
||||
--rank 128 \
|
||||
--lora_alpha 128 \
|
||||
--target_modules to_q to_k to_v to_out.0 \
|
||||
--gradient_accumulation_steps 1 \
|
||||
--gradient_checkpointing \
|
||||
--checkpointing_steps 500 \
|
||||
--checkpointing_limit 2 \
|
||||
--enable_slicing \
|
||||
--enable_tiling"
|
||||
|
||||
# Optimizer arguments
|
||||
optimizer_cmd="--optimizer adamw \
|
||||
--lr 3e-5 \
|
||||
--lr_scheduler constant_with_warmup \
|
||||
--lr_warmup_steps 100 \
|
||||
--lr_num_cycles 1 \
|
||||
--beta1 0.9 \
|
||||
--beta2 0.95 \
|
||||
--weight_decay 1e-4 \
|
||||
--epsilon 1e-8 \
|
||||
--max_grad_norm 1.0"
|
||||
|
||||
# Miscellaneous arguments
|
||||
miscellaneous_cmd="--tracker_name finetrainers-ltxv \
|
||||
--output_dir $OUTPUT_DIR \
|
||||
--nccl_timeout 1800 \
|
||||
--report_to wandb"
|
||||
|
||||
cmd="accelerate launch --config_file accelerate_configs/uncompiled_2.yaml --gpu_ids $GPU_IDS train.py \
|
||||
$model_cmd \
|
||||
$dataset_cmd \
|
||||
$dataloader_cmd \
|
||||
$diffusion_cmd \
|
||||
$training_cmd \
|
||||
$optimizer_cmd \
|
||||
$miscellaneous_cmd"
|
||||
|
||||
echo "Running command: $cmd"
|
||||
eval $cmd
|
||||
echo -ne "-------------------- Finished executing script --------------------\n\n"
|
||||
```
|
||||
|
||||
## Memory Usage
|
||||
|
||||
LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolution, **without precomputation**:
|
||||
|
||||
```
|
||||
Training configuration: {
|
||||
"trainable parameters": 117440512,
|
||||
"total samples": 69,
|
||||
"train epochs": 1,
|
||||
"train steps": 10,
|
||||
"batches per device": 1,
|
||||
"total batches observed per epoch": 69,
|
||||
"train batch size": 1,
|
||||
"gradient accumulation steps": 1
|
||||
}
|
||||
```
|
||||
|
||||
| stage | memory_allocated | max_memory_reserved |
|
||||
|:-----------------------:|:----------------:|:-------------------:|
|
||||
| before training start | 13.486 | 13.879 |
|
||||
| before validation start | 14.146 | 17.623 |
|
||||
| after validation end | 14.146 | 17.623 |
|
||||
| after epoch 1 | 14.146 | 17.623 |
|
||||
| after training end | 4.461 | 17.623 |
|
||||
|
||||
Note: requires about `18` GB of VRAM without precomputation.
|
||||
|
||||
LoRA with rank 128, batch size 1, gradient checkpointing, optimizer adamw, `49x512x768` resolution, **with precomputation**:
|
||||
|
||||
```
|
||||
Training configuration: {
|
||||
"trainable parameters": 117440512,
|
||||
"total samples": 1,
|
||||
"train epochs": 10,
|
||||
"train steps": 10,
|
||||
"batches per device": 1,
|
||||
"total batches observed per epoch": 1,
|
||||
"train batch size": 1,
|
||||
"gradient accumulation steps": 1
|
||||
}
|
||||
```
|
||||
|
||||
| stage | memory_allocated | max_memory_reserved |
|
||||
|:-----------------------------:|:----------------:|:-------------------:|
|
||||
| after precomputing conditions | 8.88 | 8.920 |
|
||||
| after precomputing latents | 9.684 | 11.613 |
|
||||
| before training start | 3.809 | 10.010 |
|
||||
| after epoch 1 | 4.26 | 10.916 |
|
||||
| before validation start | 4.26 | 10.916 |
|
||||
| after validation end | 13.924 | 17.262 |
|
||||
| after training end | 4.26 | 14.314 |
|
||||
|
||||
Note: requires about `17.5` GB of VRAM with precomputation. If validation is not performed, the memory usage is reduced to `11` GB.
|
||||
|
||||
## Inference
|
||||
|
||||
Assuming your LoRA is saved and pushed to the HF Hub, and named `my-awesome-name/my-awesome-lora`, we can now use the finetuned model for inference:
|
||||
|
||||
```diff
|
||||
import torch
|
||||
from diffusers import LTXPipeline
|
||||
from diffusers.utils import export_to_video
|
||||
|
||||
pipe = LTXPipeline.from_pretrained(
|
||||
"Lightricks/LTX-Video", torch_dtype=torch.bfloat16
|
||||
).to("cuda")
|
||||
+ pipe.load_lora_weights("my-awesome-name/my-awesome-lora", adapter_name="ltxv-lora")
|
||||
+ pipe.set_adapters(["ltxv-lora"], [0.75])
|
||||
|
||||
video = pipe("<my-awesome-prompt>").frames[0]
|
||||
export_to_video(video, "output.mp4", fps=8)
|
||||
```
|
||||
|
||||
You can refer to the following guides to know more about performing LoRA inference in `diffusers`:
|
||||
|
||||
* [Load LoRAs for inference](https://huggingface.co/docs/diffusers/main/en/tutorials/using_peft_for_inference)
|
||||
* [Merge LoRAs](https://huggingface.co/docs/diffusers/main/en/using-diffusers/merge_loras)
|
||||
@@ -0,0 +1,9 @@
|
||||
To lower memory requirements during training:
|
||||
|
||||
- Use a DeepSpeed config to launch training (refer to [`accelerate_configs/deepspeed.yaml`](./accelerate_configs/deepspeed.yaml) as an example).
|
||||
- Pass `--precompute_conditions` when launching training.
|
||||
- Pass `--gradient_checkpointing` when launching training.
|
||||
- Pass `--use_8bit_bnb` when launching training. Note that this is only applicable to Adam and AdamW optimizers.
|
||||
- Do not perform validation/testing. This saves a significant amount of memory, which can be used to focus solely on training if you're on smaller VRAM GPUs.
|
||||
|
||||
We will continue to add more features that help to reduce memory consumption.
|
||||
Reference in New Issue
Block a user