合集-PHD.ZJU

摘要:Transformer-xl: Attentive language models beyond a fixed-length context.ACL 2019 其是对Transformer架构的改造。 Transformer-XL 使学习依赖性超过固定长度而不破坏时间连贯性(450% longer 阅读全文
posted @ 2025-01-17 10:43 霜尘FrostDust 阅读(181) 评论(0) 推荐(0)
摘要:GROOT:Learning to Follow Instructions by Watching Gameplay Viedos.作者为北京大学梁一韬所在的Team CraftJarvis,发表时间为2023 Background 在开放世界下开发类人级别的具身智能体以解决开放式任务一直是人工智能 阅读全文
posted @ 2025-01-17 11:15 霜尘FrostDust 阅读(226) 评论(0) 推荐(0)
摘要:What can rl bring to vla generalization? an empirical study. arxiv 在vla模型的最后一层外接MLP来得到Q-value,从而可以使用PPO等强化学习算法进行微调 PPO表现优于DPO、GRPO等 RL微调vla使其泛化性提高 Sho 阅读全文
posted @ 2025-09-03 21:52 霜尘FrostDust 阅读(77) 评论(0) 推荐(0)
摘要:Mastering the game of Go with deep neural networks and tree search AlphaGo 2016 人类数据训练网络 —— 自我对弈强化学习 —— MCTS(PUCT) Mastering the game of Go without hu 阅读全文
posted @ 2025-10-09 11:06 霜尘FrostDust 阅读(44) 评论(0) 推荐(1)