gigagan pytorchダウンロード - gigagan pytorchソースコードのダウンロード

Adobe の新しい SOTA GAN、GigaGAN (プロジェクトページ) の実装。

また、より高速な収束 (スキップ層励起) とより優れた安定性 (弁別器での再構成補助損失) のために、軽量 gan からの発見をいくつか追加します。

また、1k ～ 4k アップサンプラーのコードも含まれており、これがこのペーパーのハイライトであると思います。

LAION コミュニティでレプリケーションを支援することに興味がある場合は、ぜひご参加ください。

$ pip install gigagan-pytorch

 import torch

from gigagan_pytorch import (
    GigaGAN ,
    ImageDataset
)

gan = GigaGAN (
    generator = dict (
        dim_capacity = 8 ,
        style_network = dict (
            dim = 64 ,
            depth = 4
        ),
        image_size = 256 ,
        dim_max = 512 ,
        num_skip_layers_excite = 4 ,
        unconditional = True
    ),
    discriminator = dict (
        dim_capacity = 16 ,
        dim_max = 512 ,
        image_size = 256 ,
        num_skip_layers_excite = 4 ,
        unconditional = True
    ),
    amp = True
). cuda ()

# dataset

dataset = ImageDataset (
    folder = '/path/to/your/data' ,
    image_size = 256
)

dataloader = dataset . get_dataloader ( batch_size = 1 )

# you must then set the dataloader for the GAN before training

gan . set_dataloader ( dataloader )

# training the discriminator and generator alternating
# for 100 steps in this example, batch size 1, gradient accumulated 8 times

gan (
    steps = 100 ,
    grad_accum_every = 8
)

# after much training

images = gan . generate ( batch_size = 4 ) # (4, 3, 256, 256)

 import torch
from gigagan_pytorch import (
    GigaGAN ,
    ImageDataset
)

gan = GigaGAN (
    train_upsampler = True ,     # set this to True
    generator = dict (
        style_network = dict (
            dim = 64 ,
            depth = 4
        ),
        dim = 32 ,
        image_size = 256 ,
        input_image_size = 64 ,
        unconditional = True
    ),
    discriminator = dict (
        dim_capacity = 16 ,
        dim_max = 512 ,
        image_size = 256 ,
        num_skip_layers_excite = 4 ,
        multiscale_input_resolutions = ( 128 ,),
        unconditional = True
    ),
    amp = True
). cuda ()

dataset = ImageDataset (
    folder = '/path/to/your/data' ,
    image_size = 256
)

dataloader = dataset . get_dataloader ( batch_size = 1 )

gan . set_dataloader ( dataloader )

# training the discriminator and generator alternating
# for 100 steps in this example, batch size 1, gradient accumulated 8 times

gan (
    steps = 100 ,
    grad_accum_every = 8
)

# after much training

lowres = torch . randn ( 1 , 3 , 64 , 64 ). cuda ()

images = gan . generate ( lowres ) # (1, 3, 256, 256)

正常な実行では、 G 、 MSG 、 D 、 MSD値は0から10の間で変動し、通常はほぼ一定に保たれます。 1,000 トレーニングステップ後のいずれかの時点で、これらの値が 3 桁を維持する場合、それは何かが間違っていることを意味します。ジェネレーターとディスクリミネーターの値が時折マイナスに下がっても問題ありませんが、上記の範囲まで戻るはずです。

GPとSSL 0に向かってプッシュする必要があります。 GP時折スパイクすることがあります。ネットワークが何らかのひらめきを経験しているように想像してみたいと思います

GigaGANクラスに?が搭載されました。アクセル。 accelerate CLI を使用すると、2 つのステップでマルチ GPU トレーニングを簡単に実行できます。

トレーニングスクリプトがあるプロジェクトのルートディレクトリで、次のコマンドを実行します。

$ accelerate launch train . py

 @misc { https://doi.org/10.48550/arxiv.2303.05511 ,
    url     = { https://arxiv.org/abs/2303.05511 } ,
    author  = { Kang, Minguk and Zhu, Jun-Yan and Zhang, Richard and Park, Jaesik and Shechtman, Eli and Paris, Sylvain and Park, Taesung } ,  
    title   = { Scaling up GANs for Text-to-Image Synthesis } ,
    publisher = { arXiv } ,
    year    = { 2023 } ,
    copyright = { arXiv.org perpetual, non-exclusive license }
}

 @article { Liu2021TowardsFA ,
    title   = { Towards Faster and Stabilized GAN Training for High-fidelity Few-shot Image Synthesis } ,
    author  = { Bingchen Liu and Yizhe Zhu and Kunpeng Song and A. Elgammal } ,
    journal = { ArXiv } ,
    year    = { 2021 } ,
    volume  = { abs/2101.04775 }
}

 @inproceedings { dao2022flashattention ,
    title   = { Flash{A}ttention: Fast and Memory-Efficient Exact Attention with {IO}-Awareness } ,
    author  = { Dao, Tri and Fu, Daniel Y. and Ermon, Stefano and Rudra, Atri and R{'e}, Christopher } ,
    booktitle = { Advances in Neural Information Processing Systems } ,
    year    = { 2022 }
}

 @inproceedings { Karras2020ada ,
    title     = { Training Generative Adversarial Networks with Limited Data } ,
    author    = { Tero Karras and Miika Aittala and Janne Hellsten and Samuli Laine and Jaakko Lehtinen and Timo Aila } ,
    booktitle = { Proc. NeurIPS } ,
    year      = { 2020 }
}

 @article { Xu2024VideoGigaGANTD ,
    title   = { VideoGigaGAN: Towards Detail-rich Video Super-Resolution } ,
    author  = { Yiran Xu and Taesung Park and Richard Zhang and Yang Zhou and Eli Shechtman and Feng Liu and Jia-Bin Huang and Difan Liu } ,
    journal = { ArXiv } ,
    year    = { 2024 } ,
    volume  = { abs/2404.12388 } ,
    url     = { https://api.semanticscholar.org/CorpusID:269214195 }
}