[置顶] 自我博弈偏好优化(Self-Play Preference Optimization,SPO)能否奖励模型?
posted @ 2025-08-22 11:07 limingqi 阅读(43) 评论(0) 推荐(0)
posted @ 2025-08-22 11:07 limingqi 阅读(43) 评论(0) 推荐(0)
posted @ 2025-07-26 12:48 limingqi 阅读(45) 评论(0) 推荐(0)
posted @ 2025-07-26 12:47 limingqi 阅读(72) 评论(0) 推荐(0)
posted @ 2025-10-24 15:28 limingqi 阅读(15) 评论(0) 推荐(0)
posted @ 2025-10-24 15:24 limingqi 阅读(21) 评论(0) 推荐(0)
posted @ 2025-10-22 23:00 limingqi 阅读(18) 评论(0) 推荐(0)
posted @ 2025-10-22 23:00 limingqi 阅读(27) 评论(0) 推荐(0)
posted @ 2025-10-18 20:10 limingqi 阅读(31) 评论(0) 推荐(0)
posted @ 2025-10-17 20:58 limingqi 阅读(20) 评论(0) 推荐(0)
posted @ 2025-10-11 11:01 limingqi 阅读(85) 评论(0) 推荐(0)
posted @ 2025-09-11 16:30 limingqi 阅读(413) 评论(0) 推荐(0)
posted @ 2025-09-10 11:20 limingqi 阅读(15) 评论(0) 推荐(0)
posted @ 2025-09-04 11:48 limingqi 阅读(104) 评论(0) 推荐(0)