Download mmdit - download do código-fonte mmdit

mmdit

Código-Fonte de IA

0.2.1

Baixar

MMDiT

Implementação de uma única camada do MMDiT, proposta por Esser et al. em Difusão Estável 3, em Pytorch

Além de uma reprodução direta, também generalizarei para > 2 modalidades, pois posso imaginar um MMDiT para imagens, áudio e texto.

Também oferecerá uma variante improvisada de autoatenção que seleciona de forma adaptativa os pesos a serem usados por meio do controle aprendido. Esta ideia veio de convoluções adaptativas aplicadas por Kang et al. para GigaGAN.

Instalar

$ pip install mmdit

Uso

 import torch
from mmdit import MMDiTBlock

# define mm dit block

block = MMDiTBlock (
    dim_joint_attn = 512 ,
    dim_cond = 256 ,
    dim_text = 768 ,
    dim_image = 512 ,
    qk_rmsnorm = True
)

# mock inputs

time_cond = torch . randn ( 2 , 256 )

text_tokens = torch . randn ( 2 , 512 , 768 )
text_mask = torch . ones (( 2 , 512 )). bool ()

image_tokens = torch . randn ( 2 , 1024 , 512 )

# single block forward

text_tokens_next , image_tokens_next = block (
    time_cond = time_cond ,
    text_tokens = text_tokens ,
    text_mask = text_mask ,
    image_tokens = image_tokens
)

Uma versão generalizada pode ser usada como tal

 import torch
from mmdit . mmdit_generalized_pytorch import MMDiT

mmdit = MMDiT (
    depth = 2 , 
    dim_modalities = ( 768 , 512 , 384 ),
    dim_joint_attn = 512 ,
    dim_cond = 256 ,
    qk_rmsnorm = True
)

# mock inputs

time_cond = torch . randn ( 2 , 256 )

text_tokens = torch . randn ( 2 , 512 , 768 )
text_mask = torch . ones (( 2 , 512 )). bool ()

video_tokens = torch . randn ( 2 , 1024 , 512 )

audio_tokens = torch . randn ( 2 , 256 , 384 )

# forward

text_tokens , video_tokens , audio_tokens = mmdit (
    modality_tokens = ( text_tokens , video_tokens , audio_tokens ),
    modality_masks = ( text_mask , None , None ),
    time_cond = time_cond ,
)

Citações

 @article { Esser2024ScalingRF ,
    title   = { Scaling Rectified Flow Transformers for High-Resolution Image Synthesis } ,
    author  = { Patrick Esser and Sumith Kulal and A. Blattmann and Rahim Entezari and Jonas Muller and Harry Saini and Yam Levi and Dominik Lorenz and Axel Sauer and Frederic Boesel and Dustin Podell and Tim Dockhorn and Zion English and Kyle Lacey and Alex Goodwin and Yannik Marek and Robin Rombach } ,
    journal = { ArXiv } ,
    year    = { 2024 } ,
    volume  = { abs/2403.03206 } ,
    url     = { https://api.semanticscholar.org/CorpusID:268247980 }
}

 @inproceedings { Darcet2023VisionTN ,
    title   = { Vision Transformers Need Registers } ,
    author  = { Timoth'ee Darcet and Maxime Oquab and Julien Mairal and Piotr Bojanowski } ,
    year    = { 2023 } ,
    url     = { https://api.semanticscholar.org/CorpusID:263134283 }
}

 @article { Zhu2024HyperConnections ,
    title   = { Hyper-Connections } ,
    author  = { Defa Zhu and Hongzhi Huang and Zihao Huang and Yutao Zeng and Yunyao Mao and Banggu Wu and Qiyang Min and Xun Zhou } ,
    journal = { ArXiv } ,
    year    = { 2024 } ,
    volume  = { abs/2409.19606 } ,
    url     = { https://api.semanticscholar.org/CorpusID:272987528 }
}

Expandir

Informações adicionais

Versão 0.2.1
Tipo Código-Fonte de IA
Data da Última Atualização 2025-01-16
tamanho 147.7KB
Vindo de Github

Aplicativos Relacionados

node telegram bot api

2024-12-14
typebot.io

2024-12-14
python wechaty getting started

2024-12-14
TranscriberBot

2024-12-14
genal chat

2024-12-14
Facemoji

2024-12-14

Recomendado para você

chat.petals.dev

Outro código-fonte

1.0.0
GPT Prompt Templates

Outro código-fonte

1.0.0
GPTyped

Outro código-fonte

GPTyped 1.0.5
node telegram bot api

Código-Fonte de IA

v0.50.0
typebot.io

Código-Fonte de IA

v3.1.2
python wechaty getting started

Código-Fonte de IA

1.0.0
waymo open dataset

Outro código-fonte

December 2023 Update
termwind

Outras categorias

v2.3.0
wp functions

Outras categorias

1.0.0

Informações Relacionadas Todos