Commit Graph

117 Commits

Author SHA1 Message Date
sayakpaul 1d5e0566fb private 2024-12-26 08:01:16 +05:30
Sayak Paul 0309f3bf67 Apply suggestions from code review
Co-authored-by: Aryan <aryan@huggingface.co>
2024-12-26 08:00:04 +05:30
sayakpaul 0c0a0a5e4a scheduler 2024-12-26 07:55:07 +05:30
sayakpaul 28ecfe6b65 torch.is_tensor check 2024-12-26 07:49:51 +05:30
Aryan 5757b1b886 fix grad norm related logging 2024-12-25 19:25:22 +01:00
sayakpaul e94c3c71d1 revert configs. 2024-12-25 21:36:22 +05:30
sayakpaul 6b54867c3c fixes 2024-12-25 21:34:56 +05:30
Sayak Paul 5aef4a6c80 Merge pull request #148 from a-r-r-o-w/deepspeed-changes
Grad Norm tracking in DeepSpeed
2024-12-25 20:08:09 +05:30
Aryan 47e4aaac2f track grad norm for deepspeed 2024-12-25 15:04:16 +01:00
Aryan e2178deea0 fix deadlock; make style 2024-12-25 14:16:31 +01:00
sayakpaul d7aa50917f fixes 2024-12-25 12:07:50 +05:30
sayakpaul 69b585a992 fixes 2024-12-25 11:35:09 +05:30
sayakpaul 875cb72f63 address review feedback. 2024-12-24 21:20:37 +05:30
sayakpaul 6329e9ed07 support DeepSpeed. 2024-12-24 10:59:21 +05:30
Sayak Paul cf9be17819 Merge pull request #129 from a-r-r-o-w/condition-precomputation
Precomputation of conditions and latents
2024-12-23 20:00:34 +05:30
Aryan 2858346b27 address review 2024-12-23 15:11:03 +01:00
Aryan 1c8f33967c add note about grad ckpt in readme 2024-12-23 14:59:05 +01:00
Aryan 8c0c28c971 create markdown tables 2024-12-23 14:16:47 +01:00
Aryan 78712f60a5 update comment 2024-12-23 14:01:39 +01:00
Aryan 75a854f9df add coauthor
Co-Authored-By: Sayak Paul <spsayakpaul@gmail.com>
2024-12-23 13:59:45 +01:00
Aryan 5261bed284 remove pipeline hooks 2024-12-23 13:58:37 +01:00
Aryan 86e1d1888a other condtions -> text conditions 2024-12-23 13:48:58 +01:00
Aryan 363841ae3f address some reviews; change hunyuan checkpoint repo-id 2024-12-23 12:54:52 +01:00
Aryan b045037850 make style 2024-12-21 15:19:03 +01:00
Aryan 0db7bf28aa update 2024-12-21 15:18:46 +01:00
Aryan 2da8cb5798 update 2024-12-21 12:24:33 +01:00
Aryan 66189c804b update 2024-12-20 10:03:39 +01:00
Aryan 223add1a59 Hunyuan Video LoRA (#126)
* add hunyuan-video lora support

* minor fixes; make style

* update readme

* update

* update

* Update README.md

* Update README.md

Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>

* update

* update

* change move train scripts to internal directory

* update

---------

Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>
2024-12-20 08:02:08 +05:30
Aryan 9ef58e2f3a LTX Video (#123)
* rename files

* ltx finetuning

* update

* update

* improvements

* make style

* gradient clipping

* update

* fix distributed inference

* update

* update
2024-12-19 04:27:36 +05:30
jiashenggu 80d1150a0e add I2V sft and fix an error (#97)
* adaption for CogVideoX1.5

* add patch_size_t in full finetuning of T2V and lora finetuning of I2V

* add I2V sft and add ofs to adapt cogvideox 1.5
2024-12-06 01:59:57 +05:30
Sayak Paul 1c7d796629 Merge pull request #112 from a-r-r-o-w/sayakpaul-patch-1
Update README.md to include mochi-1 trainer
2024-12-05 15:11:31 +05:30
Sayak Paul aa8612fc06 Update README.md to include mochi-1 trainer 2024-12-05 12:59:13 +05:30
Aryan 6c41984d8a RoPE fixes for 1.5, bfloat16 support in prepare_dataset, gradient_accumulation grad norm undefined fix (#107)
* rope scaling fixes for cog 1.5

* grad norm log fix

* bf16 prepare dataset fix

* update

* style
2024-12-02 15:12:12 +05:30
Sayak Paul 76a9f2b911 Merge pull request #106 from a-r-r-o-w/resume-mochi-1
feat: support checkpointing saving and loading
2024-12-01 20:43:25 +05:30
sayakpaul cca97bfd40 readme 2024-12-01 20:35:51 +05:30
sayakpaul 79b546b264 feat: support checkpointing. 2024-12-01 20:34:35 +05:30
glide-the ec3ec47891 Unbound fix (#105)
* is some error
```
                                [rank1]:   File "/mnt/ceph/develop/jiawei/cogvideox-distillation/training/cogvideox_image_to_video_lora.py", line 1005, in <module>                                                                                                             [rank1]:     main(args)                                                                                                 [rank1]:   File "/mnt/ceph/develop/jiawei/cogvideox-distillation/training/cogvideox_image_to_video_lora.py", line 884, in main                                                                                                                  [rank1]:     "gradient_norm_before_clip": gradient_norm_before_clip,                                                    [rank1]: UnboundLocalError: local variable 'gradient_norm_before_clip' referenced before assignment
```
ref1 :https://github.com/a-r-r-o-w/cogvideox-factory/pull/84
ref2: https://github.com/a-r-r-o-w/cogvideox-factory/pull/100
It seems that there is a bug in the accelerator.is_main_process task and accelerator.distributed_type. Is_main_process has scheduling problems during the training initiation and training stages.

* fix
2024-12-01 01:58:27 +05:30
Sayak Paul 099c9c35c4 Merge pull request #104 from a-r-r-o-w/fix-format-specifier
fix format specifier.
2024-11-29 16:29:58 +05:30
sayakpaul c8a118cefd fix format specifier. 2024-11-29 16:23:14 +05:30
Sayak Paul d5cc7c66f1 Merge pull request #100 from a-r-r-o-w/wandb-tracker-fix
wandb tracker in scheduling problems during the training initiation and training stages
2024-11-29 15:08:19 +05:30
Sayak Paul d10963f707 Merge pull request #90 from a-r-r-o-w/mochi-1-tuner
[feat] add Mochi-1 trainer
2024-11-29 14:56:20 +05:30
Sayak Paul 7090bcb9de Update README.md 2024-11-29 13:47:26 +05:30
sayakpaul dceded0a94 add a preview table 2024-11-29 13:40:02 +05:30
sayakpaul 0011fa1406 better command 2024-11-29 13:34:05 +05:30
sayakpaul 95775ba5d1 pyav pining. 2024-11-29 13:27:23 +05:30
sayakpaul 8a32b139af pin moviepy. 2024-11-29 13:22:23 +05:30
Sayak Paul 38f157c6d0 Apply suggestions from code review
Co-authored-by: Aryan <aryan@huggingface.co>
2024-11-29 13:11:08 +05:30
sayakpaul 9c86706cbb fix help message 2024-11-29 13:08:55 +05:30
sayakpaul e1866d844d better example code. 2024-11-29 13:07:26 +05:30
sayakpaul 4e3bb7ab5f updates 2024-11-29 10:52:12 +05:30