生成时间: 2026-08-03 19:24:07 (UTC+8); Arxiv 发布时间: 2026-08-03 20:00 EDT (2026-08-04 08:00 UTC+8)

今天共有 20 篇相关文章

Keyword: reinforcement learning

Learning Stateful Predictive Knowledge From Experience

从经验中学习有状态的预测知识

ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

ThinkReset:可学习的有界上下文长视野推理中间界面构建

TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter

TAPR:通过任务感知提示重写器提升LLM性能

NeuroSynth: A Biologically Inspired Continual Reinforcement Learning Architecture for Mitigating Catastrophic Forgetting

NeuroSynth:一种生物启发的持续强化学习架构,用于缓解灾难性遗忘

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

将大型语言模型中的知识提炼为轻量级强化学习代理,用于自主网络操作

Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity

基于超梯度的双级强化学习,提升样本复杂度

Gated Q-learning: Add Off-Policy Bias to Taste

门控Q学习:为品味添加非政策偏见

Think2Go: Generative Next POI Recommendation with LLM Reasoning

Think2Go:基于LLM推理的生成性下一个POI推荐

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

对一致性多参考图像编辑的评估-验证奖励

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

CLIFT:通过非侵入式闭环迭代微调,将Gemini机器人设备转变为类人型专家

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

端到端的标量奖励模型学习潜在推理痕迹

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

SAF-OPD:稳定优势融合用于政策提炼

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

《不要混合奖励,混合策略:多奖励强化学习的策略分解与优化》

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

带思想的翻译:通过强化学习实现多领域机器翻译的难度自适应推理

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

通过最佳策略识别的高效分层强化学习示例

Explore Beyond the Boundary Using Entropic Information

利用熵信息探索边界之外

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

LEMUR:学习与多目标强化学习对齐偏好反馈

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

扩散环境中多武装匪徒政策梯度的趋同与遗憾

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

WCM:视觉-语言-行动强化学习的世界批判模型

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

CodeShrink:自适应视觉压缩,实现高效的多模态代码理解

Keyword: diffusion policy

There is no result