q transformer下載 - q transformer原始碼下載

q transformer

Ai源碼

0.3.0

下載

Q-變壓器

Q-Transformer 的實現，透過自回歸 Q-Functions 進行可擴展離線強化學習，來自 Google Deepmind

我將保留單一動作 Q 學習的邏輯，只是為了與提議的多個動作自回歸 Q 學習進行最終比較。也作為對我自己和公眾的教育。

Kotb 等人複製了自迴歸 Q 學習公式。

安裝

$ pip install q-transformer

用法

 import torch

from q_transformer import (
    QRoboticTransformer ,
    QLearner ,
    Agent ,
    ReplayMemoryDataset
)

# the attention model

model = QRoboticTransformer (
    vit = dict (
        num_classes = 1000 ,
        dim_conv_stem = 64 ,
        dim = 64 ,
        dim_head = 64 ,
        depth = ( 2 , 2 , 5 , 2 ),
        window_size = 7 ,
        mbconv_expansion_rate = 4 ,
        mbconv_shrinkage_rate = 0.25 ,
        dropout = 0.1
    ),
    num_actions = 8 ,
    action_bins = 256 ,
    depth = 1 ,
    heads = 8 ,
    dim_head = 64 ,
    cond_drop_prob = 0.2 ,
    dueling = True
)

# you need to supply your own environment, by overriding BaseEnvironment

from q_transformer . mocks import MockEnvironment

env = MockEnvironment (
    state_shape = ( 3 , 6 , 224 , 224 ),
    text_embed_shape = ( 768 ,)
)

# env.init()     should return instructions and initial state: Tuple[str, Tensor[*state_shape]]
# env(actions)   should return rewards, next state, and done flag: Tuple[Tensor[()], Tensor[*state_shape], Tensor[()]]

# agent is a class that allows the q-model to interact with the environment to generate a replay memory dataset for learning

agent = Agent (
    model ,
    environment = env ,
    num_episodes = 1000 ,
    max_num_steps_per_episode = 100 ,
)

agent ()

# Q learning on the replay memory dataset on the model

q_learner = QLearner (
    model ,
    dataset = ReplayMemoryDataset (),
    num_train_steps = 10000 ,
    learning_rate = 3e-4 ,
    batch_size = 4 ,
    grad_accum_every = 16 ,
)

q_learner ()

# after much learning
# your robot should be better at selecting optimal actions

video = torch . randn ( 2 , 3 , 6 , 224 , 224 )

instructions = [
    'bring me that apple sitting on the table' ,
    'please pass the butter'
]

actions = model . get_optimal_actions ( video , instructions )

欣賞

StabilityAI、A16Z 開源 AI 資助計劃，以及？感謝慷慨的贊助，以及我的其他贊助商，為我提供了開源當前人工智慧研究的獨立性

托多

引文

 @inproceedings { qtransformer ,
    title   = { Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions } ,
    authors = { Yevgen Chebotar and Quan Vuong and Alex Irpan and Karol Hausman and Fei Xia and Yao Lu and Aviral Kumar and Tianhe Yu and Alexander Herzog and Karl Pertsch and Keerthana Gopalakrishnan and Julian Ibarz and Ofir Nachum and Sumedh Sontakke and Grecia Salazar and Huong T Tran and Jodilyn Peralta and Clayton Tan and Deeksha Manjunath and Jaspiar Singht and Brianna Zitkovich and Tomas Jackson and Kanishka Rao and Chelsea Finn and Sergey Levine } ,
    booktitle = { 7th Annual Conference on Robot Learning } ,
    year   = { 2023 }
}

 @inproceedings { dao2022flashattention ,
    title   = { Flash{A}ttention: Fast and Memory-Efficient Exact Attention with {IO}-Awareness } ,
    author  = { Dao, Tri and Fu, Daniel Y. and Ermon, Stefano and Rudra, Atri and R{'e}, Christopher } ,
    booktitle = { Advances in Neural Information Processing Systems } ,
    year    = { 2022 }
}

 @inproceedings { Kumar2023MaintainingPI ,
    title   = { Maintaining Plasticity in Continual Learning via Regenerative Regularization } ,
    author  = { Saurabh Kumar and Henrik Marklund and Benjamin Van Roy } ,
    year    = { 2023 } ,
    url     = { https://api.semanticscholar.org/CorpusID:261076021 }
}

展開

附加信息

版本 0.3.0
類型 Ai源碼
更新時間 2025-01-14
大小 1.42MB
來自於 Github

相關應用

Q房網

2024-09-08
Monster Transformer手機版

2023-09-07
QCFUN應用程式

2023-08-28
芭比Q app

2023-06-27
心掛Q

2022-08-29
Q-目錄

2009-06-22

爲您推薦

chat.petals.dev

其他源碼

1.0.0
GPT Prompt Templates

其他源碼

1.0.0
GPTyped

其他源碼

GPTyped 1.0.5
node telegram bot api

Ai源碼

v0.50.0
typebot.io

Ai源碼

v3.1.2
python wechaty getting started

Ai源碼

1.0.0
waymo open dataset

其他源碼

December 2023 Update
termwind

其他類別

v2.3.0
wp functions

其他類別

1.0.0

相關資訊全部