生成时间: 2026-07-21 18:28:21 (UTC+8); Arxiv 发布时间: 2026-07-21 20:00 EDT (2026-07-22 08:00 UTC+8)
今天共有 57 篇相关文章
Keyword: reinforcement learning
Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization
增强学习引导的NSGA-II加灰色关系系数增强多目标优化:在纳斯达克投资组合优化中的应用
- Authors: Zhiyuan Wang, Qinxu Ding, Ding Ding, Siying Zhu, Jing Ren, Yue Wang, Chong Hui Tan
- Subjects: Subjects:
Machine Learning (cs.LG); General Economics (econ.GN)
- Arxiv link: https://arxiv.org/abs/2607.16194
- Pdf link: https://arxiv.org/pdf/2607.16194
- Abstract
In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives. This paper addresses constrained multi-objective optimization (MOO) with an application to portfolio optimization for minimizing risk and maximizing return. To address existing gaps, we propose a novel reinforcement learning (RL)-guided non-dominated sorting genetic algorithm II (NSGA-II) enhanced with gray relational coefficients (GRC), termed RL-NSGA-II-GRC, which combines an RL agent controller and GRC-based selection to improve convergence and diversity of Pareto fronts. The agent adapts evolutionary parameters online using metrics of hypervolume, feasibility, and diversity, while the GRC tournament operator ranks parents via a unified score considering dominance rank, crowding distance, and proximity to ideal reference. We evaluate the framework on the Kursawe and CONSTR benchmarks and a NASDAQ portfolio application. On the benchmarks, RL-NSGA-II-GRC achieves convergence improvements of about 5.8% and 4.4% over NSGA-II, while preserving well-distributed non-dominated solutions. In the portfolio application, it produces a smooth, densely populated efficient frontier supporting identification of the maximum Sharpe ratio portfolio (annualized Sharpe =1.92) and utility-optimal portfolios for different risk-aversion levels. The main contributions are three-fold: 1) we propose an RL-NSGA-II-GRC method integrating an RL agent into the evolutionary framework to adaptively control parameters via generational feedback; 2) we design a GRC-enhanced binary tournament operator providing a comprehensive indicator to guide the search toward the Pareto front; 3) we demonstrate, on benchmark MOO and a NASDAQ case study, that the method delivers improved convergence and well-populated frontiers supporting actionable insights.
- 中文摘要
在现代金融市场中,决策者越来越依赖定量方法来应对多个且常常相互冲突的复杂目标之间的权衡。本文探讨了受限多目标优化(MOO)及其在投资组合优化中的应用,以实现风险最小化和回报最大化。为弥补现有空白,我们提出了一种新型强化学习(RL)引导的非支配排序遗传算法II(NSGA-II),并辅以灰色关系系数(GRC),称为RL-NSGA-II-GRC,它结合了RL代理控制器和基于GRC的选择,以提升帕累托前沿的收敛性和多样性。代理通过超量、可行性和多样性等指标在线调整进化参数,而GRC锦标赛运营者则通过统一评分对父母进行排名,考虑优势排名、拥挤距离和理想参考距离。我们评估了该框架,基于Kursawe和CONSTR基准以及纳斯达克投资组合应用。在基准测试中,RL-NSGA-II-GRC 在收敛率提升方面分别比 NSGA-II 分别提升了 5.8% 和 4.4%,同时保持了分布良好的非支配解。在投资组合应用中,它生成了一个平滑且人口密集的高效前沿,支持识别最大夏普比率组合(年化夏普=1.92)和不同风险厌恶水平下的效用最优投资组合。主要贡献有三方面:1)我们提出了一种RL-NSGA-II-GRC方法,将强化学习代理整合进进化框架,通过代际反馈自适应地控制参数;2)我们设计了一个增强GRC的二元锦标赛操作员,提供一个全面的指示器,引导搜索向帕累托前端方向;3)我们通过基准MOO和纳斯达克案例研究证明,该方法带来了更好的收敛性和丰富的前沿,支持可操作的洞见。
Rater State Bias in RLHF Preference Data: An Audit Framework
RLHF偏好数据中的评级者状态偏差:审计框架
- Authors: Elena Kopteva, Vitaliy Hlynianyi-Zhuk
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16195
- Pdf link: https://arxiv.org/pdf/2607.16195
- Abstract
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time. As a result, preference data can encode rater state alongside judgments about response quality. These shifts differ from ordinary disagreement or random label noise. They are state dependent, can be shared across annotators working under similar conditions, and can propagate through reward modeling and policy optimization. We therefore propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also define survival level emotional authenticity as a measurable response pattern using lexical, pragmatic, discourse, and safety related features. We analyze how correlated rater state bias can survive aggregation and enter learned reward signals. We derive five falsifiable predictions and effect size thresholds for an initial audit. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model. Our goal is to isolate a plausible and testable source of structured bias in RLHF preference data.
- 中文摘要
我们识别出一个结构化混淆因素,体现在来自人类反馈的强化学习(RLHF)中。成对偏好标签旨在反映比较输出,但也可能反映评分者在注释时的状态。在持续的压力或痛苦条件下,评级者的偏好可能会随时间发生变化。因此,偏好数据可以将评级状态与对回答质量的判断一并编码。这些变化不同于普通的争议或随机标签噪声。它们依赖于状态,可以在类似条件下工作的标注者之间共享,并通过奖励建模和策略优化传播。因此,我们提出评级者状态转移作为RLHF偏好数据中结构化偏倚的合理且可检验来源。本文提出了一个假设和审计框架,用于研究这一偏见来源。我们定义了评级者状态转移、评级者状态混淆因素以及相关评级者状态偏差。我们还将生存层面的情感真实性定义为一种可测量的反应模式,利用词汇、语用、话语和安全相关特征。我们分析了相关评审者状态偏差如何通过聚合存活并进入学习到的奖励信号。我们推导出五个可证伪的预测和初步审计的效应量阈值。最后,我们提出了可应用于公开教学调优模型的审计协议和试点研究计划。我们不推断任何具体部署型号的训练历史。我们的目标是在RLHF偏好数据中,分离出一个合理且可检验的结构化偏倚来源。
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
蒙面扩散语言模型是针对代理强化学习的强大且可引导的文本世界模型
- Authors: Darshan Deshpande
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.16204
- Pdf link: https://arxiv.org/pdf/2607.16204
- Abstract
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial state, task context, tool schemas, domain rules, and steering directives, and (ii) curate 239,403 grounded state-action trajectories spanning nine open-source environments and twelve frontier model families. We compare AR LMs and masked diffusion language models (MDLMs), showing MDLMs, via bidirectional anchor-aware denoising, achieve better coherence, groundedness, and empirically validated rollout diversity than LLMs over 4x their parameter size, at comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three OOD environments (ScienceWorld, ALFWorld, AppWorld) across three 1.2B-7B agent backbones (LFM2.5, Qwen3, Mistral), achieving up to 47% absolute gains over baselines without environment-specific fine-tuning. We further conduct behavioral analysis of failure modes under adversarial scenarios and human evaluation on realism, outcome correctness, and training utility. We open-source our work to encourage research in this direction.
- 中文摘要
强化学习(RL)的近期增长凸显了对多样化、专业培训环境的需求。手工策划的环境,任务和奖励难度固定,随着模型性能提升,会成为无效信号;而长期奖励稀疏会导致特定工作流或工具结构的模式崩溃。模拟环境状态的世界模型已与纯粹的推广性能匹配,使其在按需扩展多样性方面具有前景。然而,自回归(AR)世界模型存在从左到右的偏向,无法对全局相互依赖的状态锚点(如工具模式、前置转向和预期结果)进行条件化。我们(i)将基于文本的世界建模形式化为可引导的过渡动态问题,分解为初始状态、任务上下文、工具模式、领域规则和引导指令;(ii)整理了涵盖九个开源环境和十二个前沿模型家族的239,403条基准状态-动作轨迹。我们比较了AR LM和掩蔽扩散语言模型(MDLMs),显示MDLM通过双向锚点感知去噪,在参数规模为其4倍的情况下,在推断延迟下实现了更好的相干性、接地性和经实证验证的扩展多样性,而推断延迟相当。我们引入了即插即用的GRPO训练框架,支持确定性状态检查,并在三个1.2B-7B代理骨干(LFM2.5、Qwen3、Mistral)上对三种OOD环境(ScienceWorld、ALFWorld、AppWorld)进行零样本传输消融,实现了在不针对特定环境微调的情况下,实现了高达47%的绝对提升。我们还进一步对对抗情境下的失效模式进行行为分析,并对现实性、结果正确性和训练效用进行人类评估。我们开源我们的工作,以鼓励在这方面的研究。
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
它需要8个令牌:弱到强的策略外强强化学习,通过辅助分支
- Authors: Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Liwei Qian, Xin Pei, Jizhou Huang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16205
- Pdf link: https://arxiv.org/pdf/2607.16205
- Abstract
Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottleneck in this paradigm: on challenging reasoning tasks, the target model's samples often exhibit semantic redundancy, converging into the same erroneous "reasoning basins" that offer negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak to strong learning paradigm, where a policy's exploration is informed by a weaker but computationally efficient auxiliary model. We introduce W2SPO, an off policy RL method that injects short auxiliary segments often as brief as 8 tokens into intermediate target model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior performance among evaluated 4B scale models on mathematical reasoning benchmarks, outperforming evaluated post trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55 times training speedup. These results suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support.
- 中文摘要
带有可验证奖励的强化学习已成为提升大型语言模型推理能力的标准方法,通常通过对比多个自生成的推广来优化策略。然而,我们发现该范式中一个关键的支持限制瓶颈:在具有挑战性的推理任务中,目标模型的样本常表现出语义冗余,收敛到同样错误的“推理槽”,这些“推理槽”对政策更新的奖励对比度微乎其微。本文提出通过弱到强学习范式克服这一局限,即策略探索由较弱但计算效率高的辅助模型指导。我们引入W2SPO,一种非策略强化学习方法,向中间目标模型轨迹注入通常短达8个符号的短辅助段,目标模型随后从这些偏转状态完成推理路径。政策更新仅限于这些基于最终可验证奖励的短小插入片段。从实证角度看,W2SPO在数学推理基准测试中表现优于已评估的4B尺度模型,优于训练后评估基线。与原版GRPO在相同采样预算下相比,W2SPO Pass@1提升率从62.3%提升至64.2%,同时训练加速3.55倍。这些结果表明,弱辅助分支可以通过扩大局部勘探支持来诱导更强的目标推理政策。
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
PPO-HSC:基于广域保单覆盖优化的探索性强化学习框架
- Authors: Yujie Shen, Haowen Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16206
- Pdf link: https://arxiv.org/pdf/2607.16206
- Abstract
This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the "Invisible Shackles" of mode collapse in Large Language Model (LLM) fine-tuning. While standard Reinforcement Learning from Verifiable Rewards (RLVR) effectively reinforces high-reward trajectories, it often leads models to over-optimize known solutions, sacrificing curiosity and the ability to explore broader solution manifolds. To overcome this, PPO-HSC incorporates a High-order Sampling Coverage (HSC) reward that incentivizes the discovery of "low-similarity yet high-validity" reasoning patterns. By maintaining a dynamic trajectory library of verified unique solutions, the framework provides a differentiable signal that rewards semantic novelty while ensuring structural rationality through a plausibility constraint. Empirical evaluations on mathematical reasoning (GSM8K, SVAMP) and code generation tasks demonstrate that PPO-HSC significantly enhances solution diversity and state-space coverage while maintaining or surpassing the accuracy and syntax integrity of state-of-the-art RL baselines.
- 中文摘要
本文介绍了PPO-HSC(近端策略优化,具高阶采样覆盖),这是一种探索性强化学习框架,旨在解决大型语言模型(LLM)微调中模式崩溃的“隐形枷锁”问题。虽然标准的可验证奖励强化学习(RLVR)有效强化了高奖励轨迹,但常常导致模型过度优化已知解,牺牲好奇心和探索更广泛解流形的能力。为克服这一问题,PPO-HSC引入了高阶抽样覆盖率(HSC)奖励,激励发现“低相似性但高效度”的推理模式。通过保持一个动态的经过验证的唯一解轨迹库,该框架提供了一个可微信号,奖励语义新颖性,同时通过合理性约束确保结构合理性。数学推理(GSM8K、SVAMP)和代码生成任务的实证评估表明,PPO-HSC在保持或超越最先进强化学习基线的准确性和语法完整性的同时,显著提升了解的多样性和状态空间覆盖。
A Survey on the Verification of Reinforcement Learning Policies
关于强化学习政策验证的调查
- Authors: Luca Marzari, Ezio Bartocci, Enrico Marchesini
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16210
- Pdf link: https://arxiv.org/pdf/2607.16210
- Abstract
Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for neural network-based policies remains a major barrier to deployment. Recent advances in policy expressiveness and scale have intensified this challenge, leading to a rapidly growing but conceptually fragmented body of work on RL policy verification. This survey provides a unifying perspective on RL verification methods. We introduce a taxonomy that clarifies relationships among existing approaches along three axes: verification paradigm (formal versus probabilistic), temporal scope (step-wise versus multi-step), and guarantees strength. Beyond taxonomy, we unify underlying theoretical foundations, make implicit assumptions and limitations explicit, and identify emerging directions.
- 中文摘要
强化学习(RL)越来越多地应用于复杂且安全关键的领域,但基于神经网络的策略缺乏严格的行为保障仍然是部署的主要障碍。政策表达性和规模的近期进步加剧了这一挑战,导致在强化学习政策验证领域迅速增长但概念上分散。本调查为强化学习验证方法提供了统一的视角。我们引入了一种分类法,明确了现有方法之间的三条轴关系:验证范式(形式范式与概率范式)、时间范围(分步骤与多步式),以及保证强度。超越分类学,我们统一理论基础,明确隐含假设和局限,识别新兴方向。
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
CIGPO:多回合证据读取LLM代理的上下文信息获取策略优化
- Authors: Hao Dou
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.16244
- Pdf link: https://arxiv.org/pdf/2607.16244
- Abstract
Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit. In HotpotQA experiments with Qwen2.5-3B-Instruct, GRPO initially improves (standard F1 0.430) but subsequently collapses to 100% format-violating outputs. Training-log diagnosis reveals a zero-advantage lock-in mechanism: all sampled trajectories receive the minimum format penalty (-2.0), group-relative advantages vanish, and the policy-gradient loss becomes zero--an optimization deadlock. We propose a variance-injection strategy: by assigning per-turn rewards to intermediate evidence-reading turns, we prevent the group reward distribution from collapsing to a single value--preserving the variation that GRPO's group-relative advantage requires. Contextual Information-Gain Policy Optimization (CIGPO) implements this strategy using the marginal increase in the frozen reference model's log-likelihood of the ground-truth answer as the per-turn signal. With separate normalization of IG and F1 rewards and an IG-weight curriculum, CIGPO reaches a standard F1 of 0.518 on HotpotQA at the 3B scale (from 0.252 base; +105%), compared with 0.430 for the best GRPO checkpoint and 0.000 for the final GRPO checkpoint. CIGPO maintains meaningful reward variance and avoids zero-advantage lock-in throughout training. These results identify reward-variance collapse as a concrete failure mode of outcome-only GRPO and show that turn-level IG rewards can prevent it in this HotpotQA setting.
- 中文摘要
用仅结果强化学习训练多回合证据阅读代理不稳定,因为中间回合获得的直接认可很少。在 Qwen2.5-3B-Ininstruction 的 HotpotQA 实验中,GRPO 最初有所改善(标准 F1 0.430),但随后崩溃为 100% 格式违规输出。训练日志诊断揭示了一种零优势锁定机制:所有采样轨迹都受到最小格式惩罚(-2.0),群体相对优势消失,策略梯度损失变为零——即优化死锁。我们提出了方差注入策略:通过为中间证据阅读回合分配每回合奖励,防止群体奖励分布崩溃为单一值——保持GRPO群体相对优势所需的变异。情境信息增益政策优化(CIGPO)利用固定参考模型中对真实答案对数似然的边际增加作为每次转向信号来实现这一策略。通过IG和F1奖励的独立规范化以及IG权重课程,CIGPO在3B等级的HotpotQA标准F1(从0.252基础起;+105%)达到0.518,而最佳GRPO检查点为0.430,最终GRPO检查点为0.000。CIGPO保持有意义的奖励方差,避免培训过程中零优势锁定。这些结果将奖励-方差崩溃识别为仅结果GRPO的具体失败模式,并表明回合级的投资方差奖励可以在HotpotQA环境中防止其发生。
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
从结果到行动:利用事后诸葛亮进行长期语言代理培训
- Authors: Zishang Jiang, Tingyun Li, Jinyi Han, Xinyi Wang, Sihang Jiang, Yizhou Ying, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.16257
- Pdf link: https://arxiv.org/pdf/2607.16257
- Abstract
Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longer-horizon interactions. One major bottleneck is distinguishing the contribution of different actions in long-horizon interaction, leading to high optimization variance. To address this, we introduce a novel policy gradient method, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them. We theoretically and empirically show that aggregating semantically similar states and actions in the intent space yields a bounded-variance estimator and improves policy performance stably. Our code is available online.
- 中文摘要
强化学习(RL)已成为改进大型语言模型(LLMs)复杂任务的广泛技术。尽管取得了这些进展,现有强化学习方法在训练具有更长视野相互作用的智能体方面仍面临挑战。一个主要瓶颈是区分不同动作在长视野交互中的贡献,导致优化方差较大。为此,我们引入了一种新颖的政策梯度方法——事后诸葛政策优化(Hindsight Policy Optimization,HPO),该方法将当前政策分布和事后诸葛亮分布都投射到一个意图空间,并从它们之间的瓦瑟斯坦距离中提取低方差学习信号。我们理论和实证地证明,在意图空间中聚合语义相似的状态和动作可以得到有界方差估计器,并稳定提升策略性能。我们的代码在线提供。
Interactive Task Alignment as a POMDP
作为POMDP的交互式任务对齐
- Authors: Andy Dai, Zexue He, Zhenyu Zhang, Alex Pentland, Jiaxin Pei
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16412
- Pdf link: https://arxiv.org/pdf/2607.16412
- Abstract
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, exploratory, or even inconsistent goals, requiring the assistant to first determine the intended task before carrying it out. We study this problem as task alignment: the ability to align with a user on their intended task. We introduce a general framework for converting specified tasks into underspecified interactions, formalized as a POMDP in which the model must infer a latent task from partial and evolving user intent. We validate our user simulator post hoc with a human user study. Across shopping, coding, and professional work settings, we find that while models often perform well once the task is specified, models still struggle with task alignment: current models act prematurely, interact ineffectively, and fail to resolve ambiguous requests. Models on average recover the user's intended task only 22-32% of the time under ambiguity. In a human study in the same setting, humans reach 48%, outperforming all evaluated models. We show that post-training with supervised fine-tuning and reinforcement learning improves task alignment, but models still lag behind humans in resolving uncertainty through interaction. Together, our results suggest that current models still lack key interaction abilities required for reliable agency.
- 中文摘要
当前语言模型的基准测试主要评估对完全指定任务的执行。然而,真实用户任务往往存在模糊性。用户通常带着不完整、探索性甚至不一致的目标进入,助理必须先确定预期任务后才能执行。我们将此问题研究为任务对齐:即与用户在其预期任务上保持一致的能力。我们引入了一个通用框架,用于将特定任务转化为欠具体的交互,形式化为POMDP,模型必须从部分且不断演变的用户意图推断出潜在任务。我们事后通过人体用户研究验证了用户模拟器。在购物、编码和专业工作环境中,我们发现虽然模型在任务指定后通常表现良好,但模型在任务对齐方面仍存在困难:当前模型行动过早,交互效果不佳,且未能解决模糊请求。模型平均仅有22-32%的时间恢复用户预期任务。在同一环境下的人类研究中,人类的表现达到了48%,超过了所有评估模型。我们表明,训练后通过监督微调和强化学习改善任务对齐,但模型在通过互动解决不确定性方面仍落后于人类。综合结果表明,当前模型仍缺乏实现可靠代理所需的关键交互能力。
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
何时制定计划:学会在反应控制与深思熟虑规划之间做出选择
- Authors: Adam Labiosa, Josiah P. Hanna
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16421
- Pdf link: https://arxiv.org/pdf/2607.16421
- Abstract
It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning. In this paper, we study the question of how to learn this ability, known as meta-reasoning, in artificial agents. We model reactive decision-making as a policy that directly maps state observations to actions. Such policies can be trained with reinforcement learning (RL) or imitation learning, but may generalize poorly outside of their training distribution. Alternatively, model-based decision-time planning is more likely to produce good actions across a broader set of states but requires additional computation time, which delays acting. In this work, we introduce an RL method for training a meta-reasoning policy that allocates computation by conditioning on a reactive-policy uncertainty score. This score enables it to predict when the reactive policy is likely to perform poorly and when planning is needed. We conduct an empirical study on motion planning and navigation environments, showing that this design enables the meta-reasoning policy to learn when the reactive policy provides a good-enough action versus when decision-time planning is needed. Additionally, we show that our design enables the meta-agent to shift toward fully reactive control as the reactive policy improves.
- 中文摘要
人类早已被认可,能够在快速、被动的决策和更慢、深思熟虑的规划之间切换。本文探讨了如何在人工代理中学习这一被称为元推理的能力的问题。我们将反应性决策建模为一种政策,直接将状态观察映射到行动。这类策略可以通过强化学习(RL)或模仿学习进行训练,但在训练分布之外可能难以推广。另一种是基于模型的决策时间规划更有可能在更广泛的状态中产生良好行动,但需要额外的计算时间,从而延迟行动。在本研究中,我们引入了一种强化学习方法,用于训练元推理策略,该策略通过以反应策略不确定性评分为条件来分配计算。该评分使其能够预测反应式政策何时可能表现不佳,以及何时需要规划。我们对运动规划和导航环境进行了实证研究,表明该设计使元推理策略能够学习何时反应策略提供足够有效的行动,何时需要决策时间规划。此外,我们展示了我们的设计使元智能体能够随着反应策略的改进,向完全反应式控制转变。
Certifiable Safe Model-Based Reinforcement Learning with Control-Affine Dynamics Approximation
可认证的安全基于模型的强化学习,含控制-仿射动力学近似
- Authors: Hao Zhou, Yanze Zhang, Cameron Reid, Wenhao Luo
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.16501
- Pdf link: https://arxiv.org/pdf/2607.16501
- Abstract
Safe model-based reinforcement learning (RL) often bridges control-theoretic analysis and RL for robots to safely explore (partially) unknown system dynamics while deriving control actions for task efficiency. The control performance and safety assurance typically rely on prior knowledge of partially modeled nominal system dynamics and the data-driven models that compensate for residual model uncertainties. However, existing methods often overlook the structure of residual model uncertainties (e.g., components affine in control), which could lead to overly conservative robot behaviors or invalid safety guarantees under the safe learning-based controllers. This paper proposes a safe reinforcement learning framework that learns control-affine dynamics with a certifiable data-driven safe policy using control barrier functions (CBF). Specifically, we first use Control-Affine Random Fourier Features (ARFF) to model robot dynamics in a control-affine form, which offers computational efficiency that scales with dataset size and reduces potential model bias for model-based reinforcement learning. Then, a model-free, efficient uncertainty quantification method using adaptive conformal prediction (ACP) is applied to quantify the uncertainty in the safety constraint arising from the learned control-affine dynamics. This allows for data-driven safety assurance amenable to principled and efficient controller synthesis with CBF. Simulation results on the cartpole and the 3D quadrotor platforms demonstrate the effectiveness of the proposed framework.
- 中文摘要
基于模型的安全强化学习(RL)通常连接控制理论分析与强化学习,使机器人能够安全地探索(部分)未知的系统动力学,同时推导控制动作以提升任务效率。控制性能和安全保障通常依赖于部分建模的系统名义动力学的先验知识,以及补偿残余模型不确定性的数据驱动模型。然而,现有方法常常忽视残余模型不确定性的结构(例如控制中的仿射组件),这可能导致机器人行为过于保守,或在基于安全学习的控制器下获得无效的安全保障。本文提出了一个安全强化学习框架,利用控制障碍函数(CBF)通过可认证的数据驱动安全策略学习控制-仿射动力学。具体来说,我们首先使用控制仿射随机傅里叶特征(ARFF)以控制仿射形式建模机器人动力学,这种方法提供了随数据集规模扩展的计算效率,并减少基于模型的强化学习潜在的模型偏差。然后,应用一种无模型、高效的不确定性量化方法,利用自适应共形预测(ACP)来量化由学习到的控制-仿射动力学引起的安全性约束的不确定性。这使得基于数据的安全保障能够与CBF进行有原则且高效的控制器综合。在车杆和三维四旋翼平台上的模拟结果展示了该框架的有效性。
Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models
用于视觉-语言-动作模型的远程视野机器人操作的前瞻性残留强化学习
- Authors: Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.16506
- Pdf link: https://arxiv.org/pdf/2607.16506
- Abstract
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometrically successful for the current skill can be brittle for downstream skills. We show this failure mode in residual reinforcement learning (RL) over a frozen VLA base policy: constant sparse success rewards improve each subtask in isolation yet yield little or no gain when skills are chained, because terminal state quality is uncontrolled. We propose Foresight Residual RL, which optimizes handoff quality by augmenting each subtask's sparse success reward with an offline-estimated foresight value -- the probability of future subtask success conditioned on the terminal state of the current subtask. Concretely, we (i) train a visual foresight predictor from images of terminal states of the base policy, labeled using downstream rollout statistics, and (ii) train residual policies via backward foresight induction, using the predictor output as a reward multiplier. On a three-phase wrench-based nut-tightening assembly task in Isaac Gym (grasp, move-insert, rotate), our method achieves 85.6% full-task success, outperforming standard subtask residual RL (54.5%) and VLA baselines, while leaving per-subtask success unchanged. These results highlight that improving long-horizon performance requires shaping which successful states are produced at each sub-task, not only whether success occurs.
- 中文摘要
视觉-语言-动作(VLA)策略提供了强的通用操作先验,但在严格公差、接触丰富组装中常因长视野的功劳分配和子任务耦合而失败:当前技能几何上成功的状态对后续技能可能变得脆弱。我们在冻结的VLA基础策略上展示了残余强化学习(RL)中的失败模式:恒定稀疏的成功奖励在单独时提升每个子任务,但技能链式时几乎没有收益,因为终端状态质量未受控。我们提出了前瞻性残差强化学习(Foresight Residual RL),通过在每个子任务的稀疏成功奖励中加入一个离线估计的前瞻性值来优化交接质量——即基于当前子任务终端状态的未来子任务成功概率。具体来说,我们(i)从基础策略的终端状态图像中训练视觉前瞻预测器,这些图像使用下游推广统计,(ii)通过逆向前瞻归纳训练残余策略,将预测输出作为奖励乘数。在Isaac Gym中基于三相扳手的螺母拧紧任务(握持、移动-插入、旋转)中,我们的方法实现了85.6%的全任务成功率,优于标准子任务残余强化逻辑(54.5%)和VLA基线,同时每个子任务的成功率保持不变。这些结果表明,提升长期视野性能需要塑造每个子任务中产生哪些成功状态,而不仅仅是是否成功。
Differentiable Reinforcement Learning for Path Tracking by an Agile Fish-Like Robot
敏捷鱼类机器人用于路径跟踪的可微强化学习
- Authors: Prashanth Chivkula, Kartik Loya, Venkata Ravindhra Reddy Varikuti, Phanindra Tallapragada
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2607.16508
- Pdf link: https://arxiv.org/pdf/2607.16508
- Abstract
Fish-like swimming has inspired the design of several dozens if not hundreds of bioinspired robots in the last few decades. But the control and motion planning of such robots has been challenging due to the poorly modeled fluid-structure interaction and the nonlinear underactuated dynamics of such robots. While reinforcement learning has allowed significant advances in the context of ground and aerial robots, the lack of a suitable simulation environment with appropriate computational speed and accuracy have prevented similar progress for fish-like robots. We address this two-fold problem by developing a simulation platform that approximates the motion of our fish-like robot with computational efficiency. Then the motion control and path tracking by the robot is performed using PID control where the (variable) gains are learned using back propagation through time and training on a curriculum. The policy learned in the simulation is then applied on the physical platform, demonstrating an excellent match.
- 中文摘要
鱼类游泳激发了过去几十年里数十甚至数百个仿生机器人的设计。但由于流体-结构相互作用建模不充分以及其非线性欠致动动力学,这些机器人的控制和运动规划一直具有挑战性。虽然强化学习在地面和空中机器人领域取得了显著进展,但缺乏具备适当计算速度和精度的合适模拟环境,阻碍了鱼类机器人的类似进展。我们通过开发一个模拟平台,以计算效率近似鱼类机器人的运动来解决这一双重问题。随后,机器人通过PID控制进行运动控制和路径跟踪,通过时间的逆向传播和课程训练来学习(可变的)增益。模拟中学到的策略随后应用到物理平台上,证明了极佳的匹配。
Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
Building2Building:一个面向可推广现实世界强化学习的大规模基准
- Authors: Vincent Taboga, Justin Veilleux, Doseok Jang, Anushree Rankawat, Pierre-Luc Bacon
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16534
- Pdf link: https://arxiv.org/pdf/2607.16534
- Abstract
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
- 中文摘要
强化学习(RL)在控制方面取得了显著成效,但已学习的策略仍难以适应动态、行动空间、观察空间或目标的变化,这对现实应用来说是关键限制。现有基准测试的多样性和复杂性有限,使得在强化学习中严谨研究迁移、多任务学习和元学习变得困难。我们介绍了Building2Building(B2B),这是一套基于EnergyPlus(先进的建筑模拟器)构建的大型真实供暖、通风和空调(HVAC)控制环境套件。B2B完全兼容体育馆接口,并具备参数化建筑生成器,能够系统生成具有异构观察和动作空间的多样化建筑配置。基于该套件,我们定义了针对强化学习关键开放挑战的基准任务,包括目标适应、动态适应、行动空间转移和跨域转移。通过提供一个大规模、多样化且物理基础的测试平台,并采用标准化评估协议,B2B实现了对持续控制中泛化和转移的系统性研究。除了推动强化学习泛化研究外,这一新基准还通过实现大规模改进暖通空调控制(HVAC)而具有重大社会意义,而暖通空调是建筑中最耗能的系统之一。
From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology
从最优政策到个体差异:重新思考生物学中的强化学习
- Authors: Patrick Govoni, Palina Bartashevich, Clémence Bergerot, Valerii Chirkov, Valentin Lecheval, Pawel Romanczuk
- Subjects: Subjects:
Neural and Evolutionary Computing (cs.NE)
- Arxiv link: https://arxiv.org/abs/2607.16542
- Pdf link: https://arxiv.org/pdf/2607.16542
- Abstract
Reinforcement learning (RL) is primarily known as a computational method for optimizing control tasks, but it is increasingly used to explain biological behavior. While RL successfully captures key aspects of biology, a major gap remains: between-agent behavioral variability. Consistent individual differences naturally permeate biological populations, yet RL models typically present only the single best individual or the population average. Addressing this gap requires moving beyond current practices to generate behavioral diversity using biologically plausible mechanisms. Here, we examine approaches from various subfields of RL and outline potential paths forward to close the gap between biology and simulation.
- 中文摘要
强化学习(RL)主要作为一种优化控制任务的计算方法而闻名,但它也越来越多地被用于解释生物行为。虽然强化学习成功捕捉了生物学的关键方面,但仍存在一个重大空白:代理间行为变异性。生物群体中自然存在一致的个体差异,但强化学习模型通常只呈现单一最佳个体或群体平均值。解决这一差距需要超越现有做法,利用生物学上合理的机制来实现行为多样性。本文探讨了强化学习各个子领域的方法,并概述了缩小生物学与模拟之间差距的潜在路径。
PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration
PAVXploreRL:物理-动作-视觉世界模型强化学习结合行动探索
- Authors: Han Wang, Zijun Wang, Shuoshuo Xue, Rui Cao, Fenjiao Cheng, Xiaodang Liang, Roy Ka-Wei Lee
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.16602
- Pdf link: https://arxiv.org/pdf/2607.16602
- Abstract
Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-world rollouts. To accurately capture diverse action-induced dynamics, such models should satisfy three key objectives-Physical Plausibility (P), Action Adherence (A), and Visual Fidelity (V), collectively referred to as PAV-while remaining robust to both in-distribution (ID) expert demonstrations and out-of-distribution (OOD) actions. However, existing methods primarily rely on ID action-video pairs and pixel-level reconstruction losses, which do not explicitly optimize PAV objectives and generalize poorly beyond expert data. To address this, we propose PAVXploreRL, a reinforcement learning framework built on a pretrained latent world model that explicitly optimizes PAV objectives through reward-driven training. To improve action generalization, our method jointly leverages ID trajectories and noise-driven OOD action exploration, without paired video supervision. Experiments show that PAVXploreRL consistently outperforms pretrained baselines, achieving a 5.6% average gain across benchmarks and producing higher-quality PAV properties. As a policy evaluator, it also yields more reliable performance estimates and reduces the overestimation bias of prior expert-only world models such as Ctrl-World. Code: this https URL
- 中文摘要
行动条件世界模型是具身人工智能的关键组成部分,作为可扩展的政策评估工具,减少对昂贵的现实世界推广的依赖。为了准确捕捉多样的动作诱导动态,这些模型应满足三个关键目标——物理合理性(P)、动作依从性(A)和视觉真实度(V),统称为PAV——同时对分布内(ID)专家演示和外分布(OOD)动作保持鲁棒性。然而,现有方法主要依赖ID动作-视频对和像素级重建损失,这些方法并未明确优化PAV目标,且在专家数据之外的泛化能力较差。为此,我们提出了PAVXploreRL,这是一个基于预训练潜在世界模型的强化学习框架,通过奖励驱动训练明确优化PAV目标。为了提升动作泛化,我们的方法结合了ID轨迹和噪声驱动的OOD动作探索,无需配对视频监督。实验显示,PAVXploreRL持续优于预训练基线,在各基准测试中平均获得5.6%的提升,并产生更高质量的PAV属性。作为策略评估工具,它还能提供更可靠的性能估计,并减少之前仅专家世界模型(如Ctrl-World)的高估偏差。代码:这个 https URL
FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images
FUSAR-R1:用于智能解读SAR图像的大规模推理模型
- Authors: Yi Yang, Xiaokun Zhang, Yuxuan Li, Ruyi Zhang, Xinpeng Zhou, Haipeng Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16819
- Pdf link: https://arxiv.org/pdf/2607.16819
- Abstract
In recent years, large-scale vision-language models have been driving a paradigm shift in intelligent remote sensing image interpretation. By incorporating textual semantic information, the cognitive expression, semantic understanding, and human-computer interaction capabilities of interpretation models have been significantly improved, achieving initial progress in the field of Synthetic Aperture Radar (SAR) image interpretation. However, SAR images are affected by factors such as coherent imaging mechanisms, complex scattering characteristics, speckle noise interference, and target-background coupling, resulting in complex and variable image features with significant uncertainties and specializations. Existing SAR vision-language models do not yet possess the step-by-step analysis, logical judgment, and self-correction capabilities of human experts, making it difficult to support reliable intelligent interpretation in complex scenarios. To address this issue, this paper proposes a large-scale reasoning model, FUSAR-R1, for intelligent interpretation of SAR images. The model first constructs explicit chain-of-thought reasoning data by simulating the interpretation process of human experts and uses this data to guide instruction learning, thereby endowing the model with basic reasoning capabilities. Subsequently, a reinforcement learning strategy is introduced to optimize the model's outputs based on inference results, enabling self-correction and more reliable reasoning. Experimental results demonstrate that FUSAR-R1 consistently outperforms existing multimodal large-scale models across various SAR interpretation tasks, including target detection, target counting and classification, and land-cover category recognition.
- 中文摘要
近年来,大规模视觉语言模型正在推动智能遥感图像解读的范式转变。通过整合文本语义信息,解释模型的认知表达、语义理解和人机交互能力得到了显著提升,在合成孔径雷达(SAR)图像解读领域取得了初步进展。然而,SAR图像受相干成像机制、复杂散射特性、斑点噪声干扰以及目标-背景耦合等因素的影响,导致图像特征复杂且可变,具有显著的不确定性和专业化。现有的SAR视觉语言模型尚未具备人类专家的逐步分析、逻辑判断和自我纠正能力,这使得在复杂场景中支持可靠智能解读变得困难。为解决这一问题,本文提出了一个大规模推理模型FUSAR-R1,用于智能解读SAR图像。该模型首先通过模拟人类专家的解读过程构建显式的思维链推理数据,并利用这些数据指导教学学习,从而赋予模型基本的推理能力。随后,引入了强化学习策略,基于推理结果优化模型输出,实现自我纠正和更可靠的推理。实验结果表明,FUSAR-R1在包括目标检测、目标计数与分类以及土地覆盖类别识别在内的多种SAR解读任务中,始终优于现有的多模态大尺度模型。
Group Entropy-Controlled Policy Optimization
群熵控制策略优化
- Authors: Guangran Cheng, Chengqi Lyu, Songyang Gao, Wenwei Zhang, Kai Chen
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.16850
- Pdf link: https://arxiv.org/pdf/2607.16850
- Abstract
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce distinct entropy regimes under the same policy, making global or token-level entropy regulation insufficient to corresponding heterogeneous needs of exploration. This heterogeneity further makes GRPO-style normalized advantages induce an entropy-dependent bias, making advantage signals across prompt groups statistically non-comparable. To address this issue, we propose Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy, estimated from existing grouped samples to perform entropy-conditioned asymmetric advantage shaping. GEPO attenuates positive advantages in low-entropy groups to reduce over-exploitation, and negative advantages in high-entropy groups to preserve exploration, with adaptive thresholds derived from historical entropy statistics. Extensive experiments on two base models across thirteen benchmarks spanning mathematics, physics, science, code generation, and instruction following show that GEPO consistently outperforms GRPO and recent entropy-controlled methods, delivering balanced cross-task improvements while preserving task-specific exploration levels throughout training.
- 中文摘要
熵控制已成为大型语言模型(LLM)强化学习(RL)中的有效工具,有助于在对齐过程中平衡探索与利用之间的权衡。这种强化学习范式通常在多种异质任务的混合上进行,这些任务在同一策略下诱导不同的熵状态,使得全局或代币级熵调控无法满足相应的异质探索需求。这种异质性进一步使得GRPO风格的归一化优势产生熵依赖性偏差,使得提示组间的优势信号在统计上不可比较。为解决这一问题,我们提出了群熵控制策略优化(GEPO),这是一种轻量级的GRPO扩展,利用现有分组样本估算群熵,进行熵条件的非对称优势塑形。GEPO通过自适应阈值根据历史熵统计量,削弱低熵组的正向优势以减少过度开发,在高熵组中减少负面优势以保持探索。对涵盖数学、物理、科学、代码生成和指令跟踪的13个基准测试的两个基础模型进行的广泛实验表明,GEPO始终优于GRPO及近期熵控制方法,在实现跨任务的平衡提升的同时,在整个训练过程中保持任务特定的探索水平。
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
通过无模型的认知自由能估计量实现有原则的无方向内在动机
- Authors: Alireza Furutanpey, Schahram Dustdar
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16858
- Pdf link: https://arxiv.org/pdf/2607.16858
- Abstract
Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space.
- 中文摘要
在不确定因素混合的环境中,无监督强化学习需要内在动机,不预先确定某个惊喜方向。突袭最小化的设计范围是针对“不稳定”环境。预测误差好奇心奖励总预期惊喜,包括不可约噪声。强盗或混合在惊喜最小化和惊喜最大化奖励之间切换,通过结构重新引入了非平稳性。我们提出一个单一的内在奖励,固定在每个窗口内,源自无偏好期望自由能目标的新颖贡献,以奖励最大化形式表达。我们的主张是参数信息增益,即下一态的预期惊喜减去其不可约部分,是状态空间高熵和低熵分量中的适当内在信号。最大化它正是在寻找模型能够解释的惊喜。在未解决的动态领域,这一认识术语推动了探索。随着动力学的解决,认识论术语消失,而偶然惩罚则有利于方差较低的转变,且没有拟合显式的下一状态预测变量。伪计数提供认识价值,基于探针的惩罚捕捉偶然性方差,短视野门保护信息继承者。对所有定义奖励对象进行基于窗口的冻结,得到一个平稳的Bellman算子、学习目标的显式界限,以及在混合、平滑性、带宽和容量假设下非参数估计量的条件均匀集中结果。在主动推断的术语中,代理人在保留新颖性、标准似然模糊性消失、在可完全可观测性下消失、添加非标准转移熵惩罚,且在状态空间的已解析区域出现惊喜最小化时,该代理是无偏好的。
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
基于跟踪的策略上蒸馏,用于掩蔽扩散语言模型
- Authors: Haolin Ren, Ziyang Huang, Chenhao Yuan, Jun Zhao, Kang Liu
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.16872
- Pdf link: https://arxiv.org/pdf/2607.16872
- Abstract
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without reward estimation. The key idea is to supervise a dLLM on its own denoising trajectory, focusing on the trace-aligned token decisions that form the final response. Specifically, TOPD samples on-policy diffusion trajectories from the target dLLM, obtains teacher token distributions from a teacher model on the corresponding partially denoised states, and updates the target dLLM with a token-level Reverse Kullback-Leibler (Reverse-KL) objective. This design preserves dense teacher supervision while aligning training with the model's own denoising states. On mathematical reasoning benchmarks, TOPD enables SDAR-4B-Chat to match the MATH500 accuracy of its RL-trained counterpart TraDo-4B-Instruct, with gains of +5.7 under static evaluation and +4.5 under dynamic evaluation. Compared with the RL-trained counterpart, TOPD achieves this with 4$\times$ fewer rollout rounds, corresponding to an estimated 96.0$\times$ to-accuracy model-compute speedup.
- 中文摘要
扩散大型语言模型(dLLMs)是自回归生成的有前景替代方案。然而,面向推理的dLLM后期培训仍然具有挑战性。dLLM的监督微调(SFT)需要密集但常常偏离策略的掩蔽状态,而强化学习(RL)则依赖稀疏奖励或价值建模。本文提出了 \textbf{基于追踪的策略蒸馏(TOPD)},这是一个教师监督的框架,能够将推理能力转移到目标 dLLM 中,而无需奖励估计。关键思想是监督dLLM自身的去噪轨迹,关注形成最终响应的跟踪对齐令牌决策。具体来说,TOPD从目标dLLM中采样策略扩散轨迹,从教师模型中获得对应部分去噪状态的教师令牌分布,并用令牌级反向Kullback-Leibler(Reverse-KL)目标更新目标dLLM。这种设计既保持了密集的教师监督,又将培训与模型自身的去噪状态对齐。在数学推理基准测试中,TOPD使SDAR-4B-Chat能够与其强化学习对应的TraDo-4B-Ininstruction匹配MATH500精度,静态评估时提升+5.7,动态评估时提升+4.5。与强化学习训练的对应版本相比,TOPD以少4$\times$的滚动轮数实现这一目标,约对应模型计算精度提升96.0$\times$。
Enhancing Personalized Bladder Cancer Treatment Through Reinforcement Learning: A Recurrent Patient State Transition Decision Support Framework
通过强化学习提升个性化膀胱癌治疗:反复患者状态转换决策支持框架
- Authors: Divyansh Chawla, Anshu Garg, Isshaan Singh
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.16916
- Pdf link: https://arxiv.org/pdf/2607.16916
- Abstract
Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes. Conventional clinical decision support systems typically rely on static treatment guidelines or single-step predictive models, limiting their ability to capture disease progression over time. This paper presents a recurrent patient state-transition simulation framework for bladder cancer treatment planning that integrates predictive state-transition modeling with a Markov Decision Process (MDP) and a Deep Q-Network (DQN) reinforcement learning environment. The predictive module estimates changes in tumor characteristics following treatment, while the reinforcement learning agent sequentially optimizes treatment decisions by interacting with simulated patient trajectories. This framework enables dynamic, patient-specific treatment planning by continuously adapting recommendations to evolving clinical states. It also generates interpretable treatment trajectories and detailed simulation logs to improve transparency and support clinical decision-making. The proposed framework was evaluated against existing reinforcement learning-based treatment planning approaches. It achieved a cumulative reward of 63,918.87, an average training loss per episode of 0.0056, and a policy improvement score of 6.62%, demonstrating effective sequential learning and robust treatment optimization in a simulated recurrent treatment environment. These findings highlight the potential of recurrent patient state-transition simulation with reinforcement learning as a flexible decision-support framework for personalized bladder cancer treatment planning and AI-assisted precision oncology.
- 中文摘要
膀胱癌治疗需要个性化和适应性决策,尤其是针对复发性疾病,治疗效果在连续临床发作中变化。传统的临床决策支持系统通常依赖静态治疗指南或单步预测模型,限制了其随时间追踪疾病进展的能力。本文提出了一种用于膀胱癌治疗规划的循环患者状态转换模拟框架,该框架将预测性状态转变建模与马尔可夫决策过程(MDP)和深度Q网络(DQN)强化学习环境整合。预测模块估计治疗后肿瘤特征的变化,而强化学习代理则通过与模拟患者轨迹交互,顺序优化治疗决策。该框架通过不断调整建议以适应不断演变的临床状态,实现动态且针对患者的治疗规划。它还生成可解读的治疗轨迹和详细的模拟日志,以提高透明度并支持临床决策。该框架与现有基于强化学习的治疗计划方法进行了比较。其累计奖励为63,918.87,平均每集训练损失为0.0056,策略改进得分为6.62%,在模拟循环治疗环境中展现了有效的顺序学习和稳健的治疗优化。这些发现凸显了反复患者状态转换模拟的潜力,强化学习作为个性化膀胱癌治疗计划和AI辅助精准肿瘤学的灵活决策支持框架。
Counterfactual Shapley Credit Assignment
反事实沙普利信用分配
- Authors: Mingxuan Li, Kaizhan-Lee, Elias Bareinboim
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.16999
- Pdf link: https://arxiv.org/pdf/2607.16999
- Abstract
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to CAP must isolate the true causal drivers of observed outcomes from spurious correlations and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel framework grounded in causal theory that attributes credit and blame via the Counterfactual Shapley Value ($\phi$-value). By redistributing environmental rewards, $\phi$-values enhance temporal credit assignment across three critical dimensions: sparse causality, high stochasticity, and delayed rewards, all while preserving the optimal policy. We derive a consistent estimator that computes $\phi$-values efficiently, enabling a new class of policy gradient methods, $\phi$-PPO, combined with Prioritized Trajectory Replay (PTR). Empirical results demonstrate that $\phi$-values align precisely to the ground truth causes of task rewards with superior sample efficiency in challenging environments where prior state-of-the-art methods fail to converge.
- 中文摘要
学分分配问题(CAP)是开发高效且可解释强化学习(RL)代理的基础。现有框架,无论是依赖时间连续性还是事后条件奖励加权,常常未能正确区分主体策略(技能)与环境随机性(运气)之间的关系。对CAP的原则性方法必须将观察到的结果的真实因果驱动力与虚假相关性和环境随机性隔离开来。我们介绍了反事实沙普利信用分配,这是一种基于因果理论的新框架,通过反事实沙普利价值($\phi$-value)来归因功劳和责任。通过重新分配环境奖励,$\φ$价值在三个关键维度上增强了时间信用分配:稀疏因果关系、高随机性和延迟奖励,同时保持最优政策。我们推导出一个一致的估计量,高效计算 $\phi$ 值,从而实现了一类新的策略梯度方法——$\phi$-PPO,并结合了优先轨迹重放(PTR)。实证结果表明,$\phi$值与任务奖励的真实原因完全一致,在以往最先进方法无法融合的挑战环境中,样本效率更高。
Scalable Causal Imitation Learning
可扩展因果模仿学习
- Authors: Eylam Tagor, Mingxuan Li, Elias Bareinboim
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17003
- Pdf link: https://arxiv.org/pdf/2607.17003
- Abstract
Imitation learning enables learning a policy in an unknown environment with a latent reward signal using expert demonstrations, but it struggles when the imitator's and expert's observations are mismatched and unobserved confounders are present in expert demonstrations. By identifying appropriate adjustment sets via the sequential $\pi$-backdoor criterion, causal imitation learning (CIL) provides a framework for approximating the expert's policy from confounded data. However, existing CIL methods, Causal Behavioral Cloning (Causal BC) and Causal Generative Adversarial Imitation Learning (Causal GAIL), are designed for short-horizon, low-dimensional settings. When applied to continuous control tasks with long horizons and high-dimensional state-action spaces, these methods exhibit poor performance: Causal BC suffers from compounding errors, Causal GAIL is unstable and sample-inefficient, and sequential $\pi$-backdoor adjustment becomes impractical. We introduce Causal Soft Q Imitation Learning (SQIL) and Causal Inverse soft-Q Learning (IQ-Learn), two off-policy causal imitation learning algorithms that combine the causal adjustment framework with state-of-the-art inverse reinforcement learning objectives. Both algorithms operate on causally-adjusted state representations produced by an efficient approximation of the sequential $\pi$-backdoor criterion, exploiting the causal structure of continuous control environments to reduce the full-horizon adjustment to a fixed-size sliding window. We evaluate all methods in a suite of confounded environments and find that Causal SQIL and Causal IQ-Learn substantially outperform prior CIL algorithms on long-horizon tasks, sometimes surpassing the expert, whereas all causally unaware imitation methods fail to learn meaningful behavior.
- 中文摘要
模仿学习使得在未知环境中通过专家演示学习策略成为可能,但当模仿者与专家的观察不匹配且存在未被观察到的混杂因素时,模仿学习会变得困难。通过通过顺序的$\pi$后门准则识别合适的调整集,因果模仿学习(CIL)提供了一个框架,用于从混杂数据中近似专家的策略。然而,现有的CIL方法——因果行为克隆(因果BC)和因果生成对抗模仿学习(Causal GAIL)——是为短视野、低维环境设计的。当应用于具有长视野和高维状态-动作空间的连续控制任务时,这些方法表现较差:因果BC存在复利误差,因果GAIL不稳定且样本效率低下,顺序的$\pi$后门调整变得不切实际。我们引入了因果软Q模拟学习(SQIL)和因果逆软Q学习(IQ-Learn),这两种非策略因果模仿学习算法,将因果调整框架与最先进的逆强化学习目标结合起来。这两种算法都基于因果调整状态表示,这些表示通过对连续控制环境的因果近似产生,利用连续控制环境的因果结构,将全视距调整简化为固定大小的滑动窗口。我们在一组混淆环境中评估了所有方法,发现因果SQIL和因果IQ-Learn在长期任务中显著优于以往的CIL算法,有时甚至超过专家,而所有因果无知模仿方法都未能学习有意义的行为。
Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
以奖励为驱动的大型语言模型代理工作流:综合POMDP路由与自我纠正以实现自主决策
- Authors: Amez Amanj Ali, Kuo-Kun Tseng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17038
- Pdf link: https://arxiv.org/pdf/2607.17038
- Abstract
This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sparse reward attribution, and dynamic environmental interaction, by designing and optimizing an intelligent agent workflow. The proposed architecture is based on the synthesis of core AI paradigms: Visual, Language, Generative, Graph, Multimodal, Reinforcement, and Agent Intelligence. Unlike conventional baseline models that rely on static prompting and lack robust perception-action loops, our approach introduces a Partially Observable Markov Decision Process (POMDP) routing mechanism. This mechanism is augmented with an internal, self-correcting reward model that evaluates decision trajectories before execution. By integrating multimodal inputs and advanced reinforcement learning principles (such as proximal policy optimization and value function approximation), the agent maintains long-term structural memory and dynamically adapts its reasoning pathways to mitigate error accumulation. Empirical experiments on the ALFWorld embodied simulation environment and the WebShop online navigation benchmark demonstrate a 24.5% absolute improvement in task success rate and trajectory efficiency over mainstream baselines like the standard ReAct framework. Comprehensive ablation studies confirm the significant contribution of the reward-driven critique module in suppressing hallucination rates. This research bridges theoretical foundations of reinforcement learning and graph-based memory with autonomous agent workflows. Ultimately, the resulting architecture offers a practical, scalable reference framework for developing artificial intelligence technologies in complex, multi-step autonomous systems. Code is available at this https URL.
- 中文摘要
本文通过设计和优化智能体工作流,解决当前大型语言模型(LLM)代理应用中的关键技术挑战,包括长视野规划、稀疏奖励归因和动态环境交互。所提出的架构基于核心AI范式的综合:可视化、语言、生成、图、多模态、强化和代理智能。与依赖静态提示且缺乏稳健感知-动作循环的传统基线模型不同,我们的方法引入了部分可观察马尔可夫决策过程(POMDP)路由机制。该机制通过内部自我修正的奖励模型增强,在执行前评估决策轨迹。通过集成多模输入和高级强化学习原则(如近端策略优化和价值函数近似),智能体保持长期结构记忆,并动态调整推理路径以减少错误累积。在ALFWorld内嵌模拟环境和WebShop在线导航基准测试上的实证实验显示,任务成功率和轨迹效率相比主流基线如标准ReAct框架绝对提升了24.5%。全面的消融研究证实了奖励驱动批判模块在抑制幻觉率方面的重要贡献。本研究将强化学习和基于图的记忆的理论基础与自主智能体工作流相结合。最终,这一架构为在复杂、多步的自主系统中开发人工智能技术提供了一个实用且可扩展的参考框架。代码可在此 https URL 访问。
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
STBridge:在UMM中连接理解与生成的共享目标对齐
- Authors: Ye Wang, Hongjun Wang, Hao Fang, Tongyuan Bai, Zuwei Long, Peixian Chen, Wei Liu, Weibo Gu, Xing Sun, Rui Ma
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.17140
- Pdf link: https://arxiv.org/pdf/2607.17140
- Abstract
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic consistency. A model may describe the intended target correctly while generating an inconsistent edit. This exposes an understanding-generation alignment gap: linguistic and visual outputs live in different spaces, yet should be governed by the same target semantics. We study this gap in image editing, where an instruction defines a target state that can be both described and visually realized. Given a source image and an edit instruction, we compare a UMM's target caption with its edited image to test whether the two outputs converge on the same result. Our analysis shows that existing UMMs remain weakly aligned, especially for fine-grained entities, attributes, spatial relations, and local details, indicating that semantic unification is not achieved by architecture alone. To bridge this gap, we propose STBridge, a shared-target alignment framework that connects understanding and generation through a common target state. Here the target caption expresses the desired visual result, while the edited image realizes it visually, replacing separate task-specific paths with a shared information flow from target expression to target realization. STBridge follows an align-then-optimize strategy: supervised fine-tuning first establishes the shared-target channel, and sequential reinforcement learning further refines target-centered coordination. Across visual understanding, image generation, and image editing benchmarks, STBridge consistently improves over the initialization model. Alignment analysis confirms that STBridge narrows the gap between what the model describes and what it generates, demonstrating shared-target alignment as an effective post-training strategy for bridging understanding and generation in UMMs.
- 中文摘要
统一多模态模型(UMMs)旨在将视觉理解和生成整合到单一架构中,但仅靠架构统一并不能确保语义一致性。模型可能正确描述目标,但生成不一致的编辑。这暴露了理解与世代对齐的差距:语言输出和视觉输出存在于不同的空间,但应由相同的目标语义来支配。我们研究图像编辑中的这一空白,指令定义了一个既可描述又可直观实现的目标状态。给定源图像和编辑指令,我们比较UMM的目标说明与编辑后的图像,以测试两个输出是否收敛于相同结果。我们的分析显示,现有的UMMs仍然弱对齐,尤其是在细粒度实体、属性、空间关系和局部细节方面,表明语义统一并非仅靠架构实现。为弥合这一差距,我们提出了STBridge,一种通过共同目标状态连接理解与生成的共享目标对齐框架。这里,目标说明表达期望的视觉结果,而编辑后的图像则通过视觉实现实现,用共享的信息流替代任务特定路径。STBridge 遵循一种先对齐再优化的策略:先监督微调建立共享目标通道,顺序强化学习进一步完善以目标为中心的协调。在视觉理解、图像生成和图像编辑基准测试中,STBridge 在初始化模型基础上持续提升。比对分析证实STBridge缩小了模型描述与生成内容之间的差距,证明共享目标比对是连接UMM理解与生成的有效训练后策略。
EdgeCoInfer: Hierarchical Collaborative Inference for On-Device Multimodal Large Models
EdgeCoInfer:用于设备内多模态大型模型的层级协作推理
- Authors: Lin Tan, David K. Y. Yau, Songtao Guo
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2607.17143
- Pdf link: https://arxiv.org/pdf/2607.17143
- Abstract
Modern mobile applications predominantly execute concurrent Multimodal Large Language Models (MLLMs) to provide ubiquitous intelligence. However, satisfying this demand within edge environments faces significant challenges due to multi-task concurrency and strictly coupled hard constraints. To address these issues, we propose EdgeCoInfer, a framework enabling granularity-adaptive deployment by co-optimizing inter-model functional module sharing and \textbf{intra-model fine-grained partitioning}. We solve the underlying Mixed-Integer Non-Linear Programming (MINLP) problem via a Hybrid Evolutionary Hierarchical Reinforcement Learning (HE-HRL) paradigm, which synchronizes a Genetic Algorithm (GA) for discrete model placement with a Soft Actor-Critic (SAC) agent for continuous resource allocation. To navigate the sparse feasible region, we introduce a feasibility-guided constructive execution mechanism, integrating a constructive cut-step decoder with pre-act pruning and a two-phase curriculum strategy for stable adaptation. Experimental results demonstrate that EdgeCoInfer ensures a 100\% task completion rate in high-concurrency scenarios, achieving a 76\% reduction in system cost and 71.88\% memory savings compared to state-of-the-art baselines.
- 中文摘要
现代移动应用主要同时执行多模态大型语言模型(MLLMs),以提供无处不在的智能。然而,在边缘环境中满足这一需求面临重大挑战,因为多任务并发和严格耦合硬约束。为解决这些问题,我们提出了EdgeCoInfer框架,通过协同优化模型间功能模块共享和\textbf(模型内细粒度划分)实现粒度自适应部署。我们通过混合进化层级强化学习(HE-HRL)范式解决了混合整数非线性规划(MINLP)问题,该范式将用于离散模型放置的遗传算法(GA)与用于连续资源分配的软演员-批判者(SAC)代理同步。为应对稀疏可行区域,我们引入了可行性导向的构造性执行机制,将建设性切割步骤解码器与前动作剪枝及两阶段课程策略相结合,实现稳定适应。实验结果表明,EdgeCoInfer 在高并发场景下确保任务完成率达到 100%,相比最先进基线,实现了 76% 的系统成本降低和 71.88% 的内存节省。
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
医患对话的有力总结:TalTech为超越转录挑战而打造的系统
- Authors: Aivo Olev, Tanel Alumäe
- Subjects: Subjects:
Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- Arxiv link: https://arxiv.org/abs/2607.17230
- Pdf link: https://arxiv.org/pdf/2607.17230
- Abstract
This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.
- 中文摘要
本文介绍了TalTech提交的Beyond Transcription Challenge(BeTraC)项目,该挑战要求直接从长的医患对话录音生成SOAP笔记,无需中间转录。在筛选了开放权重语音LLMs的长音频稳健性后,我们调整了Voxtral Mini(轻量级轨道)和Voxtral Small(重型轨道),并由LoRA监督进行微调,随后采用DAPO强化学习,并以挑战指标Open Medical Concept F1为奖励。我们的系统在两个方向中均排名第一,且一项独立的LLM评审评估显示所有提交中幻觉率最低,表明基于概念匹配指标的强化学习不必影响事实可靠性。我们还发现,对文本转录进行微调能很好地转化为语音输入,并且似乎能提升域外真实录音的鲁棒性。
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
LenGuard-GPC:带引导提示一致性的长度防护空间推理强化学习
- Authors: Xingjian Tao, Yiwei Wang, Yujun Cai, Jing Tang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17243
- Pdf link: https://arxiv.org/pdf/2607.17243
- Abstract
Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate. Reinforcement learning with verifiable rewards is a natural fit for this task, but standard GRPO reward relies on sparse outcome-level feedback and gives no signal about where a reasoning trajectory goes wrong, nor any control over its length. We propose LenGuard-GPC, a dense reward framework that addresses both problems together. For each sampled trajectory, it compares the token-wise predictive distributions under a standard prompt and a guided prompt, and uses the resulting token-sum KL divergence as a dense reward signal. Since this KL penalty accumulates over tokens and would otherwise reward shorter responses regardless of their quality, we introduce a staged length bonus that keeps reasoning length within a controlled range without simply encouraging brevity. On six multi-view spatial reasoning benchmarks, LenGuard-GPC improves accuracy over vanilla GRPO while reducing average response length.
- 中文摘要
多视角空间推理需要视觉语言模型比较图像中的视觉证据,对齐对象对应关系,并在长时间的视觉环境中推断空间关系,在这种环境中,思维链推理往往冗长而不够准确。带有可验证奖励的强化学习非常适合这项任务,但标准的GRPO奖励依赖于稀疏的结果层级反馈,且不提供推理轨迹出错的信号,也无法控制其长度。我们提出了LenGuard-GPC,一个密集的奖励框架,同时解决这两个问题。对于每个采样轨迹,它比较标准提示和引导提示下的逐标记预测分布,并利用所得的标记和 KL 发散作为稠密的奖励信号。由于这种 KL 惩罚会累积于代币数量,否则无论回复质量如何都会奖励较短的回复,因此我们引入了分阶段长度奖励,将推理长度控制在可控范围内,而不仅仅是鼓励简洁。在六项多视角空间推理基准测试中,LenGuard-GPC在降低平均响应长度的同时,提升了比普通GRPO的准确性。
Distilled Reinforcement Learning for LLM Post-training
LLM后期培训的精炼强化学习
- Authors: Chen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng, Ge Lan, Yue Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17247
- Pdf link: https://arxiv.org/pdf/2607.17247
- Abstract
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at this https URL.
- 中文摘要
大型语言模型(LLM)训练后对提升推理能力、适应性和对齐至关重要。现有方法主要遵循两种范式:强化学习(RL)和策略提纯(OPD)。然而,强化学习依赖粗粒度的结果监督,导致学分分配困难且获取新知识的能力有限。而OPD则无条件通过KL分歧匹配教师日志,这造成了两难:相似的教师提供的新知识很少,而差异显著的教师往往提供无效指导,使OPD主要限于家族内部的提炼。我们提出了提炼强化学习(Distilled Reinforcement Learning,简称Distilled RL),将教师监督融入RL目标,提供细致指导,选择性地传递新知识,避免无条件模仿。蒸馏式强化学习包含三个组成部分:带裁波的反向重要性抽样、负样本重置和序列级几何归一化。通过简明且易于理解的案例研究,我们展示了蒸馏式强化学习能够有效将教师模型中此前无法获得的知识转移到学生模型中。在家族内和跨家族蒸馏环境中的大量实验表明,蒸馏RL在pass@1和节pass@k方面均显著优于标准强化学习和门诊治疗。我们的代码可在此 https URL 访问。
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning
WAR:同步代理强化学习的工作负载感知推广
- Authors: Ryan Xu, Atlas Zhao, David Bao, Frank Du
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Operating Systems (cs.OS)
- Arxiv link: https://arxiv.org/abs/2607.17299
- Pdf link: https://arxiv.org/pdf/2607.17299
- Abstract
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectories rapidly grow to tens of thousands of tokens, making synchronous RL training increasingly constrained by rollout. We propose WAR, a workload-aware rollout system that substantially accelerates synchronous agentic RL by jointly optimizing decoding and scheduling. WAR is built on a key observation: the optimal rollout optimization strategy depends on runtime load: (1) Under low load, WAR enables model-free speculative decoding with SuffixDecoding, which reuses suffix patterns from previously completed trajectories as speculative drafts for future rollouts. Unlike model-based drafters, SuffixDecoding introduces no additional draft model and avoids GPU contention with rollout generation. (2) Under high load, where saturated batched decoding leaves limited room for speculative speedup, WAR shifts the optimization focus to cache-aware scheduling. A global scheduler places requests across rollout replicas based on cache locality, trajectory progress and server load, reducing redundant KV-cache recomputation and mitigating load imbalance. By combining decoding-level suffix reuse with system-level rollout scheduling, WAR delivers robust throughput improvements across workload regimes without changing the underlying RL algorithm. WAR improves long-context agentic rollout throughput by 1.4x under low load and up to 1.6x under high load. These results show that WAR removes a major rollout bottleneck in synchronous agentic RL and provides a practical path toward scalable long-context agent training.
- 中文摘要
长期扩展生成已成为智能强化学习(RL)中主导的系统瓶颈。随着代理多次与环境交互,轨迹迅速增长到数万个令牌,使同步强化学习在部署时受到越来越多限制。我们提出了WAR,一种工作负载感知的展开系统,通过联合优化解码和调度,大幅加快同步代理强化学习的进程。WAR基于一个关键观察:最优的推展优化策略取决于运行时负载:(1) 在低负载下,WAR通过SuffixDecoding实现无模型的推测性解码,该后缀模式将先前已完成轨迹中的后缀模式作为未来推测草稿重用。与基于模型的制图器不同,SuffixDecoding 不引入额外的草稿模型,并且避免了与 Rollout 生成时的 GPU 争用。(2)在高负载下,饱和批处理解码为推测加速留有有限空间,WAR将优化重点转向缓存感知调度。全局调度器根据缓存局部性、轨迹进度和服务器负载在部署副本间提交请求,减少重复的KV缓存重算并缓解负载不平衡。通过将解码级别后缀重用与系统层面的展开调度结合,WAR在不改变底层RL算法的情况下,实现了跨工作负载的稳健吞吐量提升。WAR在低负载下将长上下文代理扩展吞吐量提升1.4倍,高负载时最高可提升1.6倍。这些结果表明,WAR消除了同步智能体RL中一个主要的部署瓶颈,并为实现可扩展的长上下文代理训练提供了切实可行的路径。
Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
玻尔兹曼理性理论的合理化:熵正则化政策的公理化刻画
- Authors: Silviu Pitis
- Subjects: Subjects:
Machine Learning (cs.LG); Theoretical Economics (econ.TH)
- Arxiv link: https://arxiv.org/abs/2607.17316
- Pdf link: https://arxiv.org/pdf/2607.17316
- Abstract
The softmax policy $\pi(a \mid s) \propto \exp(\beta Q(s,a))$ is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, exploration, and optimization have been offered in the RL literature, but none uniquely derives the softmax form from first principles. This leaves a basic tension unresolved: the entropy bonus in the soft Bellman equation violates the Independence axiom that underwrites the Markov decision process (MDP) reward structure. We dissolve this tension by distinguishing two kinds of randomness: chance and choice. By restricting von Neumann-Morgenstern (VNM) Independence to environmental lotteries over base prospects, we show that imposing independence of irrelevant alternatives (IIA) and monotonicity on the policy and value functions at choice nodes uniquely determines the Boltzmann policy, the entropy-regularized representation, and the soft Bellman equation. The choice between the soft and hard Bellman equations thus reduces to a design decision: whether the agent values its own ability to choose. We develop RL-specific consequences, including return monotonicity and convergence under generalized discounting, and synthesize the independent lines from economics and information theory that arrive at the same structure, offering a normative assessment of when IIA is appropriate for agent design.
- 中文摘要
softmax策略$\pi(a \mid s) \propto \exp(\beta Q(s,a))$是强化学习(RL)中随机选择的默认模型。基于鲁棒性、探索性和优化的各种论据在强化学习文献中被提出,但没有一种能独一无二地从第一原理推导出软极大形式。这留下了一个基本的张力未解:软贝尔曼方程中的熵加成违反了支撑马尔可夫决策过程(MDP)奖励结构的独立性公理。我们通过区分两种随机性:偶然性和选择,来化解这种紧张关系。通过将冯·诺依曼-摩根斯特恩(VNM)的独立性限制在环境彩票中对基前景,我们证明了在选择节点对政策函数和价值函数施加无关备选方案独立性(IIA)和单调性,唯一确定了玻尔兹曼策略、熵正则化表示和软贝尔曼方程。软贝尔曼方程与硬贝尔曼方程之间的选择因此归结为设计决策:智能体是否重视自身选择能力。我们发展了强化学习特有的后果,包括收益单调性和广义贴现下的收敛性,并综合了经济学和信息理论中得出相同结构的独立线索,提供何时IIA适合代理设计的规范性评估。
TAPAS: Throughput-adaptive Perception for Autonomous Systems
TAPAS:自主系统的吞吐量自适应感知
- Authors: Aman Vyas, Vasista Kodumagulla, Zain Taufique, Pasi Liljeberg, Anil Kanduri
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17317
- Pdf link: https://arxiv.org/pdf/2607.17317
- Abstract
Autonomous systems rely on a perception module to navigate through dynamic environments. In real-world scenarios, the perception module's throughput requirements vary at runtime due to changes in scene complexity. However, existing perception strategies assume a fixed FPS and static model-to-cluster mapping, resulting in either over/under provision of throughput requirements or unnecessary energy consumption across diverse scenes. Addressing this challenge requires tightly coupled \textit{scene complexity awareness} to estimate an appropriate FPS target and \textit{dynamic model-to-cluster mapping} to deliver the required throughput at minimum energy. We propose a throughput-adaptive perception strategy for mobile/edge platforms, enabling intelligent runtime resource allocation based on varying FPS targets. We use Reinforcement Learning (RL) with RRM (Reward Reasoning Model) and a GRU (Gated Recurrent Unit) agent to orchestrate perception tasks across heterogeneous mobile/edge platforms. We evaluate TAPAS on Jetson Orin NX across KITTI and unseen nuScenes. On the \textit{KITTI} dataset's test sequences, TAPAS achieves 93-100% throughput met rate while saving energy by 76%. On the unseen \textit{nuScenes} dataset, TAPAS maintains 97% throughput met rate with 64% lower energy compared to \textit{SOTA} approaches, proving its robustness.
- 中文摘要
自主系统依赖感知模块在动态环境中导航。在实际场景中,感知模块的吞吐量需求因场景复杂度的变化而在运行时变化。然而,现有的感知策略假设固定的帧率和静态模型到集群映射,导致吞吐量需求过剩或不足,或在不同场景中产生不必要的能耗。解决这一挑战需要紧耦合的\textit{场景复杂度意识}来估算合适的帧率目标,以及\textit{动态模型到集群映射}以最低能耗提供所需吞吐量。我们提出了一种针对移动/边缘平台的吞吐量自适应感知策略,实现基于不同FPS目标的智能运行时资源分配。我们使用强化学习(RL)结合RRM(奖励推理模型)和GRU(门控循环单元)代理,在异构移动/边缘平台上协调感知任务。我们评估Jetson Orin NX在KITTI及未公开nuScenes上的TAPAS。在\textit{KITTI}数据集的测试序列中,TAPAS实现了93%-100%的通量满足率,同时节能76%。在未见的\textit{nuScenes}数据集上,TAPAS保持了97%的吞吐量满足率,且能量比\textit{SOTA}方法低64%,证明了其鲁棒性。
Rethinking the Suitability of Reinforcement Learning Algorithms Under Practical Transfer Constraints
在实际转移约束下重新思考强化学习算法的适用性
- Authors: Hany Hamed, Abhishek Naik, Colin Bellinger, A. Rupam Mahmood
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.17326
- Pdf link: https://arxiv.org/pdf/2607.17326
- Abstract
Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which asks whether conclusions about algorithm suitability change under wall-clock rather than interaction-based budgets, and robustness under dynamics mismatch, which asks how different learning paradigms respond to variability in the training distribution induced by domain randomization. We provide two insights to reinforcement-learning practitioners. First, comparing the sample efficiency of different algorithms is often an insufficient criterion in transfer-oriented settings. The wall-clock time required to train a decent policy is an important consideration for practitioners, and we find that the sample-inefficient PPO algorithm can produce a performant policy faster than relatively more sample-efficient algorithms such as SAC and TD-MPC2, validating the common understanding of massively parallel training paradigms. Second, domain randomization can help different kinds of algorithms learn robust policies. In particular, although PPO, SAC, and TD-MPC2 represent different RL paradigms - on-policy, off-policy, and model-based learning and planning, respectively - we find that domain randomization affects all three algorithms in a similar way. To the best of our knowledge, this is the first controlled comparison of the effect of domain-randomization coverage on PPO, SAC, and TD-MPC2 under the same transfer protocol. Taken together, these two insights highlight the importance of evaluating RL algorithms not only by sample efficiency, but also by practical considerations such as training time and the algorithms' ability to produce usable policies.
- 中文摘要
迁移导向强化学习需要在超出标准样本效率的维度上评估算法。我们关注两个维度:实用效率,即在基于交互的预算下,算法适用性的结论是否发生变化;以及在动态不匹配下的鲁棒性,探讨不同学习范式如何应对由领域随机化引起的训练分布变异。我们为强化学习从业者提供两点见解。首先,在以转移为导向的环境中,比较不同算法的样本效率往往是一个不充分的标准。训练一个不错策略所需的墙时钟时间是实践者的重要考虑因素,我们发现样本效率较低的PPO算法能比SAC和TD-MPC2等相对更高效的算法更快地生成高效策略,验证了对大规模并行训练范式的普遍理解。其次,域随机化可以帮助不同类型的算法学习稳健策略。特别是,尽管PPO、SAC和TD-MPC2代表不同的强化学习范式——分别是开策略、关闭策略以及基于模型的学习与规划——但我们发现领域随机化对这三种算法的影响相似。据我们所知,这是首次在同一传输协议下,对域随机化覆盖率对PPO、SAC和TD-MPC2影响的受控比较。综合来看,这两点洞见凸显了评估强化学习算法的重要性,不仅要考虑样本效率,还要考虑训练时间和算法产生可用策略的能力等实际因素。
CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards
CORAL:学习淀粉样纤维配体结合并获得合作结合奖励
- Authors: Yasheng Sun, Bohan Li, Youqi Tao, Jürgen Schmidhuber
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.17412
- Pdf link: https://arxiv.org/pdf/2607.17412
- Abstract
A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics. Predicting how such ligands bind to fibril targets, however, presents two fundamental challenges. First, resolved co-crystal structures of amyloid-ligand complexes are exceptionally scarce; even with recent advances in cryo-EM only a handful have been structurally characterized, making supervised training of docking models impractical for this target class. Second, amyloid fibrils present a binding mode fundamentally different from globular proteins: ligands intercalate into longitudinal cross-$\beta$ grooves and stack cooperatively along the fibril axis, a geometry that existing docking models are not designed to capture. To address these challenges, we present CORAL (COopeRative Amyloid Ligand docking), a reinforcement learning framework that trains a generative docking model to produce ligand pose distributions tailored to the cross-$\beta$ groove geometry. Our reward explicitly incorporates cooperative ligand-ligand stacking energy alongside protein-ligand docking affinity, directly capturing the distinctive binding geometry of amyloid fibrils. We further introduce a curated evaluation set of amyloid-ligand complexes constructed from model-generated poses validated by domain experts. Experiments on both experimentally resolved structures and this evaluation set demonstrate improved pose quality and binding affinity correlation over existing docking baselines.
- 中文摘要
阿尔茨海默病和帕金森病等神经退行性疾病的一个标志是蛋白质异常聚集成淀粉样纤维,而选择性结合这些纤维的小分子有望成为诊断、影像探针和治疗手段。然而,预测这些配体如何结合纤维靶点,带来了两个根本性的挑战。首先,淀粉样-配体复合物的分辨共晶结构极为稀少;即使冷冻电磁技术有新进展,结构上只有少数模型被明确描述,使得对该目标类别的监督对接模型训练不切实际。其次,淀粉样纤维的结合方式与球状蛋白根本不同:配体嵌入纵向交叉 $/beta$ 沟槽,并沿纤维轴协同堆叠,这种几何形状是现有对接模型未设计的。为应对这些挑战,我们提出了CORAL(COopeRative 淀粉样样配体对接),这是一种强化学习框架,训练生成式对接模型,生成针对交叉$/beta$沟槽几何形状的配体态分布。我们的奖励明确包含了合作配体-配体堆积能量以及蛋白质-配体结合亲和力,直接捕捉淀粉样纤维独特的结合几何结构。我们还进一步介绍了一组由领域专家验证的模型生成姿态构建的精选类淀粉样蛋白-配体复合物评估集。实验中,实验通过实验解析结构和该评估集,显示其姿态质量和结合亲和力相关性优于现有对接基线。
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
TraversRL:可穿越的步行路径生成与强化学习
- Authors: Bin Han, Robert Wolfe, Bill Howe
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.17479
- Pdf link: https://arxiv.org/pdf/2607.17479
- Abstract
Automatically generating pedestrian pathways from aerial images requires producing a connected network suitable for routing, not just detecting where sidewalks appear. Sidewalks and crossings, in contrast to roads, may be partially occluded, implicitly defined, and exhibit complex connectivity patterns. Existing segmentation-based approaches focus on labeling pixels to infer segments, but often produce disconnected or fragmentary graphs that are unreliable for navigation. We introduce TraversRL, a vision-conditioned model that iteratively grows a pathway network from an aerial image, simulating a traveler navigating the built environment. TraversRL uses an action space of short and long direction-distance segments designed to adapt to complex patterns and span occlusions, and uses a combination of graph-level and step-wise rewards to balance overall connectivity with precise edge placement. Across three visual backbones and three intersection datasets, TraversRL substantially improves buffered IoU with the ground-truth graph relative to a state-of-the-art segmentation baseline, and more than doubles metrics of connectivity. Moreover, combining global and local rewards produces cleaner graphs with fewer spurious branches while further improving overall performance. These results demonstrate that modeling pathway extraction as a sequential decision process from the perspective of a traveler, while optimizing for final graph quality with reinforcement learning, produces significantly more reliable pedestrian networks.
- 中文摘要
从航拍图像自动生成行人路径需要建立一个适合路由的连接网络,而不仅仅是检测人行道出现的位置。与道路不同,人行道和斑马道可能被部分封闭,隐含定义,并表现出复杂的连通模式。现有基于分割的方法主要通过标记像素来推断段,但常常产生断裂或片段化的图表,这些图对导航来说不可靠。我们介绍了TraversRL,一种视觉条件模型,通过从航拍图像迭代构建路径网络,模拟旅行者在建成环境中导航。TraversRL 采用短距离和方向距离段的动作空间,设计以适应复杂模式和跨度遮挡,并结合图级和逐步奖励,平衡整体连通性与精准边缘布局。在三个视觉骨干和三个交叉数据集中,TraversRL相比最先进的切割基线大幅提升了与地面真实图的缓冲IoU,并将连接度指标提升了一倍多。此外,结合全局和局部奖励能产生更清晰的图,减少虚假分支,同时进一步提升整体表现。这些结果表明,从旅行者视角将路径提取建模为顺序决策过程,同时通过强化学习优化最终图质量,能显著提升行人网络的可靠性。
Reinforcement Learning: From Algorithms To Foundation Models
强化学习:从算法到基础模型
- Authors: Zihan Ding
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17560
- Pdf link: https://arxiv.org/pdf/2607.17560
- Abstract
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studies how an agent should act to maximise long-term reward in a dynamic environment. In richer settings, the problem extends beyond a single agent and fixed environment: intelligent behavior may require strategic interaction, adaptation to uncertainty, and reasoning over high-dimensional worlds. This thesis studies RL from two perspectives: algorithms in games and RL in the era of foundation models. The first part focuses on multi-agent RL in games. It examines how incentives, policies, and equilibrium concepts interact in competitive and general-sum environments, spanning two-player zero-sum games, large-scale video games, and multi-player settings with general structure. These works investigate learning in multi-agent systems and the behavior of RL methods in interactive environments. The second part studies RL with generative and foundation models, motivated by the idea that prior knowledge can enrich sequential decision making. Pretrained generative models and learned world models serve as representation tools and structured priors for planning, control, and policy optimization. The thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations. It also addresses long-horizon modeling through architectures with memory. Together, these contributions present a unified view of RL as objective-driven adaptation in complex sequential domains. From strategic games to generative world models, the thesis highlights how RL connects decision making, environment modeling, and emerging foundation-model capabilities, offering a broader perspective on the principles underlying intelligent behavior.
- 中文摘要
强化学习(RL)为在明确目标下的顺序决策提供了框架。在经典形式中,强化学习研究代理在动态环境中应如何行动以最大化长期回报。在更丰富的环境中,问题不仅限于单一主体和固定环境:智能行为可能需要战略性互动、适应不确定性以及对高维世界的推理。本论文从两个角度研究强化学习:游戏中的算法和基础模型时代的强化学习。第一部分聚焦于游戏中的多智能体强化学习。它考察了激励、政策和均衡概念在竞争性和一般和环境中的相互作用,涵盖了双人零和游戏、大规模视频游戏以及具有一般结构的多人游戏环境。这些研究探讨了多智能体系统中的学习以及强化学习方法在交互环境中的行为。第二部分研究生成式和基础模型的强化学习,基于先验知识能丰富连续决策的理念。预训练生成模型和学习过的世界模型作为表示工具和结构化先验,用于规划、控制和政策优化。论文开发基于扩散的世界模型,研究强化学习以实现高效的视频生成,探索生成模型作为策略类,并研究交互式视频世界模型,其中行动塑造未来观测。它还通过带内存的架构实现长视野建模。这些贡献共同呈现了强化学习作为复杂序列领域中目标驱动适应的统一观点。从战略博弈到生成世界模型,论文强调强化学习如何连接决策、环境建模和新兴的基础模型能力,提供了对智能行为背后原理的更广泛视角。
AGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models
AGG:雅可比聚合群梯度用于扩散模型高效GRPO训练
- Authors: Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan chen
- Subjects: Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2607.17572
- Pdf link: https://arxiv.org/pdf/2607.17572
- Abstract
Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matching models introduces a severe computational bottleneck: gradients must be back-propagated through the high-capacity DiT backbone at \emph{every} timestep of the sampling trajectory, making high-resolution text-to-image (T2I) training prohibitively expensive. Training-free DiT inference acceleration methods (e.g., $\Delta$-DiT, ScalingCache) exploit the fact that DiT hidden states and velocity predictions vary \emph{smoothly and nearly linearly} along the trajectory. We ask whether the same linearity can reduce the backward-pass cost of DiT RL training, and answer affirmatively with \textbf{JAGG} (\textbf{J}acobian-\textbf{A}ggregated \textbf{G}roup \textbf{G}radient), which reduces full transformer backward passes from $W$ to $2$ per group of $W$ consecutive steps. JAGG approximates intermediate-step Jacobians via $t$-weighted interpolation of the endpoint Jacobians, then aggregates per-step upstream signals into two composite gradients applied through a single joint backward pass. We prove this interpolation is \emph{exact} when the velocity is linear in $(z,t)$, and a cosine-similarity routing rule (\texttt{jagg_frac}) deploys JAGG only where the assumption holds. Experiments on T2I benchmarks show JAGG delivers $\sim$2$\times$ backward speedup with negligible quality degradation.
- 中文摘要
群相对策略优化(Group Relative Policy Optimization,简称GRPO)是一种强大的强化学习算法,用于将生成模型与人类偏好对齐。虽然在大型语言模型中取得成功~\cite{shao2024deepseekmathingpushinglimitsmamatical},但其扩展到扩散和流匹配模型时引入了严重的计算瓶颈:梯度必须在采样轨迹的每个时间步反向传播通过高容量DiT骨干,使得高分辨率文本到图像(T2I)训练成本高昂。无训练的DiT推断加速方法(例如$\Delta$-DiT、ScalingCache)利用了DiT隐藏状态和速度预测沿轨迹平滑且近乎线性地变化的事实。我们询问相同的线性是否能降低DiT强化学习的后向传递成本,并用\textbf{JAGG}(\textbf{acobian-\textbf{A}aggregated \textbf{G}roup \textbf{G}radient)回答肯定,该方法将每组连续$W步的全变换器后向传递从$W$减少到$2$。JAGG通过对端点雅可比矩阵进行$t$加权插值,近似中间步的雅可比矩阵,然后将每步上游信号聚合为两个复合梯度,通过单一关节的后向传递施加。当速度在 $(z,t)$ 为线性时,我们证明该插值是 \emph{exact},且余弦相似路由规则 (\texttt{jagg_frac})仅在假设成立时部署 JAGG。T2I基准测试的实验显示,JAGG能以极小的质量劣化实现倒退加速。
Concentration and Mean-Square Bounds for Contractive Stochastic Approximation: A Unified Elementary Approach
收缩随机近似的集中度与均方界限:统一初等方法
- Authors: Siddharth Chandak
- Subjects: Subjects:
Machine Learning (cs.LG); Systems and Control (eess.SY); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2607.17595
- Pdf link: https://arxiv.org/pdf/2607.17595
- Abstract
We establish mean-square and concentration bounds for stochastic approximation (SA) with arbitrary norm contractive mappings, under a multiplicative noise model where the noise may scale affinely with the norm of the iterates, and the iterates are potentially unbounded. These settings arise in reinforcement learning, where operators are often contractive in the $\ell_\infty$ norm and the noise scales with the iterates. To address the arbitrary norm, earlier works replace the non-smooth squared norm with a smooth Lyapunov function constructed via the generalized Moreau envelope. For concentration analysis, these works handle multiplicative noise and unbounded iterates through a multi-stage bootstrapping argument that starts from a time-varying worst-case bound and iteratively refines it. We instead present a unified and elementary analysis that yields both bounds. Using an averaged noise sequence and corresponding auxiliary iterates, we obtain a one-step Lyapunov drift inequality for the normed error directly, without smoothing the norm or constructing an envelope. For the mean-square bound, we combine this drift inequality with an induction argument showing that the iterates remain bounded in expectation. For the concentration bound, we develop a probabilistic induction over a sequence of "good" events on which the iterates are controlled, allowing the standard Azuma-Hoeffding bound to be applied. Our approach yields the first sub-Gaussian tailed maximal (all-time) concentration bound for SA under multiplicative noise, by allowing the stepsize to depend logarithmically on the confidence level. Beyond the specific setting considered here, we discuss the generalizability of these proof techniques to other noise models and iterative algorithms.
- 中文摘要
我们在乘法噪声模型下,利用任意范数收缩映射建立随机近似(SA)的均方和集中界限,噪声可能与迭代范数相邻近,且潜在无界。这些设置出现在强化学习中,算符通常在$\ell_\infty$范数下是收缩的,噪声随迭代次数增长。为了解决任意范数,早期研究用通过广义莫罗包络构造的光滑李雅普诺夫函数替代了非光滑平方范数。对于集中分析,这些工作处理乘法噪声和无界迭代,通过多阶段自助推断论证,从时间变化的最坏情况上界出发,并迭代细化。我们反而提出了一个统一且初步的分析,得出了两个界限。利用平均噪声序列和相应的辅助迭代,我们直接得到赋范误差的一步李雅普诺夫漂移不等式,无需平滑范数或构造包络。对于均方界限,我们将漂移不等式与归纳法结合,证明迭代次数在期望上保持有界。对于集中界限,我们对一系列“良好”事件发展概率归纳,迭代受控,从而应用标准的东-霍夫丁界限。我们的方法通过允许步长与置信水平对数依赖,实现了乘法噪声下SA的首个亚高斯尾最大(历时)浓度约束。除了这里讨论的具体环境外,我们还讨论了这些证明技术在其他噪声模型和迭代算法中的推广性。
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
ConsiSpace:学习几何一致性对视频空间推理至关重要
- Authors: Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.17599
- Pdf link: https://arxiv.org/pdf/2607.17599
- Abstract
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.
- 中文摘要
视频空间推理对于导航导向感知和长视频问答至关重要,模型必须在变化视角下推断跨长视野的空间关系。然而,现有的多模态大型语言模型(MLLM)仍然主要以语义为中心,且常常无法可靠地聚合冗余视频观察中的一致空间证据,导致推理效率低下或不稳定。为解决这些问题,我们提出了ConsiSpace,这是一个几何一致性感知框架,用于几何敏感视频空间推理,将空间一致性转化为证据组织原则和显式后SFT学习信号。我们构建了一个几何一致性记忆(GCM),包括隐式证据标记和显式几何线索,并利用高效的组织策略紧凑地保存任务相关的空间证据。此外,我们利用统一一致性自监督强化学习(UC-SSRL),在监督微调后提升交叉视图稳定性,并获得答案、度量和拓扑一致性的奖励。在三个空间推理基准测试VSI-Bench、OSI-Bench和MMSI-Video-Bench上的广泛实验显示,平均得分比最强基线提升了12.6分。
On Optimal Event-Triggered Distributed Control for Stochastic Multi-Agent Systems via Reinforcement Learning
关于通过强化学习实现随机多智能体系统最优事件触发分布式控制
- Authors: Ziming Wang, Bingbing Li, Karl H. Johansson, Apostolos I. Rikos
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2607.17635
- Pdf link: https://arxiv.org/pdf/2607.17635
- Abstract
We propose a reinforcement learning (RL) based optimal distributed control algorithm for the multi-agent systems (MASs) with stochastic uncertainties. Unlike existing methods, during the optimized backstepping design process, we use the actor-critic-identifier structure. The actor neural network is used to reflect control behavior, the critic neural network works to evaluate control performance and the unknown stochastic uncertainties are handled by identifier neural network. Furthermore, a low-pass filter effectively suppresses problems stemming from non-affine nonlinear faults and a hybrid event-triggered control (ETC) strategy is proposed to reduce control frequency. We analyze our algorithm's operation, and we provide a Lyapunov-based stability proof that guarantees all errors are bounded, ensuring precise tracking between the leader and followers. We validate its correctness in a single-axis robotic manipulator simulation and finally, we compare against the non-optimal control algorithm highlighting our optimal control algorithm's operational advantages.
- 中文摘要
我们提出了一种基于强化学习(RL)的最优分布式控制算法,适用于具有随机不确定性的多智能体系统(MASs)。与现有方法不同,在优化回溯设计过程中,我们采用了actor-critic-identifier结构。演员神经网络用于反映控制行为,批判神经网络用于评估控制性能,未知的随机不确定性则由标识性神经网络处理。此外,低通滤波器有效抑制非仿射非线性故障引发的问题,并提出了混合事件触发控制(ETC)策略以降低控制频率。我们分析算法的运行,并提供了基于李雅普诺夫的稳定性证明,保证所有误差均有界,确保领导者与跟随者之间的精确追踪。我们在单轴机器人机械臂模拟中验证其正确性,最后与非最优控制算法进行比较,突出该算法的操作优势。
Mobile Network Control with a World Model
用世界模型进行移动网络控制
- Authors: Maxime Bouton, Ioanna Mitsioni, Simon Lindståhl, Jaeseong Jeong
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17747
- Pdf link: https://arxiv.org/pdf/2607.17747
- Abstract
The increasing complexity of mobile networks necessitates intelligent and dynamic control strategies for efficient, energy-conserving management. We propose a world model-based approach for network control that enables adaptive configuration of crucial parameters. The world model is trained from historical data and predicts the impact of its actions on future network states. Our controller leverages the model's uncertainty estimate to robustly find optimal network configuration changes. Furthermore, the optimization objective can be changed dynamically without model retraining. We demonstrate the effectiveness of the approach in simulated closed-loop control of a mobile network energy-saving feature. Our results show improved performance in balancing energy savings with quality of service, compared to traditional methods and reinforcement learning approaches. Finally, we show the world model performance on real network data from, and evaluate counterfactual actions proposed by the controller under various throughput constraints.
- 中文摘要
移动网络日益复杂,要求智能且动态的控制策略以实现高效、节能的管理。我们提出了一种基于世界模型的网络控制方法,能够自适应配置关键参数。世界模型基于历史数据训练,预测其行为对未来网络状态的影响。我们的控制器利用模型的不确定性估计,稳健地找到最优的网络配置变更。此外,优化目标可以动态更改,无需模型重新训练。我们展示了该方法在移动网络节能功能的闭环模拟控制中的有效性。我们的结果显示,与传统方法和强化学习方法相比,在节能与服务质量之间取得更佳的平衡表现。最后,我们展示了基于真实网络数据的世界模型性能,并评估控制器在各种吞吐量约束下提出的反事实行为。
Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
推广与指导:分解奖励以实现少数样本逆强化学习
- Authors: Ziyi Liu, Grace Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.17760
- Pdf link: https://arxiv.org/pdf/2607.17760
- Abstract
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.g., picking up mugs with varying shapes), making it impractical to collect demonstrations that fully specify a new task under every possible scenario. In practice, while demonstrations for the target task are limited, it is often easier to obtain datasets of heterogeneous but related behaviors. This motivates the problem of few-shot IRL with multi-task demonstrations (FM-IRL), where an agent must learn a new task with substantial variations from only a limited number of target-task demonstrations, together with sufficient demonstrations of related tasks and online agent experience. To do so, we must both recover the expert distribution of the new task and provide guidance when the agent deviates from it. We introduce Multitask discriminator Proximity-Guided IRL (MPG), which learns two complementary reward components: (1) a generalizable discriminator that transfers shared structure across related tasks to identify expert behavior in a new task, and (2) a proximity function that measures how far a state deviates from expert behavior and provides corrective guidance during exploration. We demonstrate the effectiveness of our method on multiple challenging navigation and manipulation tasks under significant variations (e.g., object configurations, table layouts, and initial robot poses), achieving an average success rate of 81.2%, outperforming the strongest per-task baseline by an average of 24.7 percentage points.
- 中文摘要
逆向强化学习(IRL)提供了一个强大的演示学习框架。然而,现实任务常常表现出显著的自然差异(例如,拿起形状各异的杯子),这使得收集能够在每种可能场景下完全指定新任务的演示变得不切实际。实际上,虽然目标任务的演示有限,但获取异构但相关行为的数据集通常更容易。这引发了多任务演示(FM-IRL)中少射真实任务演示(FM-IRL)的问题,即代理必须通过有限数量的目标任务演示、足够多的相关任务演示和在线代理经验,学习一个具有显著变化的新任务。为此,我们必须既恢复新任务的专家分布,又在主体偏离时提供指导。我们引入了多任务判别器近似引导IRL(MPG),它学习两个互补的奖励组成部分:(1)一个可推广的判别器,将共享结构传递到相关任务之间以识别新任务中的专家行为;(2)一个接近函数,衡量状态偏离专家行为的程度,并在探索过程中提供纠正指导。我们展示了该方法在多种具有挑战性的导航和操作任务中(如物体配置、桌面布局和初始机器人姿势)下的有效性,平均成功率为81.2%,平均比最强的单任务基线高出24.7个百分点。
Theoretical Foundations of $\max$@$k$ Reinforcement Learning
$\max$@$k 强化学习的理论基础
- Authors: Riccardo Poiani, Martino Bernasconi, Andrea Celli
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.17823
- Pdf link: https://arxiv.org/pdf/2607.17823
- Abstract
Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ responses rather than sampling a single response, and performance is then measured using a retry-aware metric such as $\max$@$k$. Despite their practical importance, the theoretical foundations of learning under such criteria remain limited. In this work, we provide a theoretical study of the $\max$@$k$ learning problem in finite-horizon reinforcement learning. We show that optimizing the $\max$@$k$ objectives is fundamentally different from standard expected-return maximization. In particular, we prove that Markovian policies are in general insufficient, identify a compact state augmentation that restores optimality, and explicitly characterize the performance gap that can arise between history-dependent and non-history-dependent policies. Moreover, we show that learning $\max$@$k$-optimal policies is statistically harder than standard reinforcement learning and provide an efficient algorithm that achieves the optimal sample complexity rate.
- 中文摘要
强化学习是现代大型推理模型的基石技术。通常,对于代码生成和定理证明等复杂任务,智能体通过生成$K$的响应来评估,而不是抽样单个响应,然后用如$\max$@$k$这样的重试感知指标来衡量性能。尽管这些标准具有实际意义,但基于这些标准的学习理论基础仍然有限。本研究旨在理论研究有限视界强化学习中的$\max$@$k$学习问题。我们表明,优化$\max$@$k$目标与标准期望回报最大化有根本不同。特别是,我们证明了马尔可夫策略普遍不足,识别出恢复最优性的紧致状态增强,并明确描述了历史依赖策略与非历史依赖策略之间可能出现的性能差距。此外,我们证明学习$\max$@$k$最优策略在统计上比标准强化学习更难,并提供了一种高效算法,实现最优样本复杂度率。
Distributional Soft Bellman Operator under the Cramér Geometry
Cramér 几何下的分布软贝尔曼算子
- Authors: Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.17897
- Pdf link: https://arxiv.org/pdf/2607.17897
- Abstract
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrt{\gamma}$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.
- 中文摘要
分布软策略迭代(DSPI)为将分布强化学习(DRL)与最大熵控制结合提供了重要框架,其中策略评估步骤由分布软贝尔曼算符控制,作用于熵正则化的收益。对此类评估步骤的理论分析需要一个概率度量,通过该概率指标可以控制贝尔曼更新,通常通过证明操作员收缩任意两个候选回报分布估计之间的距离来实现。本文重点关注克拉梅尔几何,这是一种基于累积分布函数(CDF)的度量,结构为$L^2$,并研究固定策略分布软贝尔曼算符是否具有该收缩性质,从而在该度量下存在唯一的不动点。直接在可接受的CDF域域上,我们构造了CDF级分布软贝尔曼算子,证明它是$\sqrt{\gamma}$-收缩,并结合收敛迭代策略评估获得相应的唯一不动点。CDF表述还表明,这一有限克拉梅尔域性质源于对一步奖励熵变化的均匀第一矩条件,而非奖励项和熵项的独立一致有界假设。然后我们将同一评估问题通过共轭传输到谱域,得到同一判定过程的等价希尔伯特空间表示。综合来看,这些结果确定了与 DSPI 策略评估步骤相关的 Cramér 几何 Bellman 不动点,为研究近似批评者、评估误差和 DSPI 风格算法中的批评者损失设计提供了参考点。
Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss
在通信丢失下实现多智能体协调的价值感知预测
- Authors: Kemal Devrim Kafadar, Eren Özaltun, Mahmud Efnan Şanlı, Feyza Orak, Emirhan Gazi, Kubilay Kağan Kömürcü, Nazım Kemal Üre
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.17914
- Pdf link: https://arxiv.org/pdf/2607.17914
- Abstract
Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain operation during these intermittent communication failures, agents can employ internal prediction models to estimate missing shared state information. However, predictors trained with standard reconstruction objectives treat all transitions equally. In a Reinforcement Learning context, this forces the model to waste capacity learning stochastic exploration noise and the outdated dynamics of suboptimal policies. In this paper, we propose a value-aware extension of Multi-Agent Observation Sharing under Communication Dropout (MARO) to patch communication gaps; we refer to this method as Value-Aware MARO. By dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture, our objective explicitly couples the predictor's learning process to the policy's evolution. This formulation focuses the model's capacity on the intentional, high-return dynamics actively reinforced by the agents. We evaluate our framework on several tasks within the Multi-Agent Particle Environment under varying communication reliability levels. Experimental results demonstrate that our approach maintains performance under declining communication reliability, particularly below 40%. While our method performs comparably in tasks where the baseline already maintains high coordination, our value-aware weighting effectively prevents the performance collapse observed in the standard predictor during high-attrition scenarios. In these environments, our method achieves an average improvement in mean returns of more than 20% and reduces performance variance by a mean of 64.7% compared to the standard unweighted baseline.
- 中文摘要
稳健的多智能体协调高度依赖智能体间通信,而这种通信在现实部署中常常被物理和环境限制所中断。为了在这些间歇性通信失败期间保持运行,代理可以利用内部预测模型来估算缺失的共享状态信息。然而,用标准重建目标训练的预测器对所有转变一视同仁。在强化学习的背景下,这迫使模型浪费容量学习的随机探索噪声和过时的次优策略动态。本文提出在通信中途(MARO)下对多代理观察共享的价值意识扩展,以补丁通信缺口;我们称这种方法为价值感知MARO。通过动态加权预测变量损失函数,利用基于底层actor-critic架构的优势估计,我们的目标明确将预测变量的学习过程与策略演变耦合。该表述将模型容量集中于由代理主动强化的有意高回报动态。我们在多智能体粒子环境中,在不同通信可靠性水平下评估了该框架的多个任务。实验结果表明,我们的方法在通信可靠性下降(尤其是低于40%)的情况下仍能保持性能。虽然我们的方法在基线已保持高度协调的任务中表现相当,但我们的价值感知权重有效防止了标准预测变量在高流失场景中观察到的性能崩溃。在这些环境下,我们的方法平均平均回报提升超过20%,绩效方差平均降低64.7%,相比标准无加权基线。
PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks
PRIME:无人机辅助紧急通信网络的多智能体环境中可塑性恢复
- Authors: Wen Qiu, Zhiqiang He, Wei Zhao, Hiroshi Masui
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2607.17922
- Pdf link: https://arxiv.org/pdf/2607.17922
- Abstract
Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn. The obvious remedy, resetting dormant neurons, is unsafe under shared-parameter multi-agent training: many neurons that appear inactive are still receiving strong training gradients, and whether a neuron appears dormant depends on which agent's observations it processes. PRIME (Plasticity Recovery In Multi-agent Environments) therefore verifies both directions before intervening. Extending the bidirectional Silent Neuron framework to cooperative multi-agent reinforcement learning, it aggregates activation and gradient statistics over the full team batch, reads the backward signal from the gradient the training loss has already deposited , not from a hand-crafted proxy, and reinitializes only neurons that are simultaneously activation-dormant and gradient-silent. Useful representations are preserved while learning capacity is restored. On a phase-switching UAV emergency communication simulator, PRIME improves interquartile mean return by 24.9\% over MAPPO and holds dormant neuron fractions at 10--20\% versus 40--45\%; ablations attribute the gains to the gradient signal and team-level aggregation rather than to the specific reset operator. A dynamic regret bound shows that the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.
- 中文摘要
这些网络的大多数强化学习控制器假设是静止状态,少数处理变化的控制器对外部环境做出反应,而不检查网络的内部状态。我们表明,持续的非平稳性直接损害了这种内部状态:随着目标转移,神经元逐渐进入休眠状态,共享策略失去学习能力。显而易见的解决办法是重置休眠神经元,但在共享参数多智能体训练下是不安全的:许多看似不活跃的神经元仍接收到强烈的训练梯度,神经元是否处于休眠状态取决于它处理的是哪个智能体的观察。因此,PRIME(多代理环境中的可塑性恢复)在介入前会对双向进行验证。它将双向无声神经元框架扩展到合作多智能体强化学习,汇总整个团队批次的激活和梯度统计数据,读取训练丢失已积累的梯度的反向信号,而非手工制作的代理,并仅初始化同时处于激活休眠和梯度静默状态的神经元。有用的表示得以保留,同时学习能力得以恢复。在相位切换无人机紧急通信模拟器上,PRIME比MAPPO提高了24.9%的四分位平均回波,并将休眠神经元分数保持在10--20\%对40--45%;消融分析将增益归因于梯度信号和团队层级聚合,而非特定复位操作员。动态遗憾界限表明,微扰成本随小静默子空间维数而成比例,而非全参数数。
Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization
优势中汇总,而非比率:合作多智能体策略优化的典型形式分析
- Authors: Zijian Zhao, Sen Li
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.17924
- Pdf link: https://arxiv.org/pdf/2607.17924
- Abstract
Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents\footnote{In this paper, "neighbors" refer not only to physical proximity but also to agents whose actions influence one another.} to aggregate in order to effectively utilize global information for cooperation. This decision must be made along two dimensions: in the advantage (which agents' rewards contribute to the credit signal) and in the ratio (which agents' likelihood ratios form the clipped importance weight). Existing methods occupy scattered, underexplored points on these two axes: IPPO treats both separately; MAPPO pairs a team-level advantage with per-agent ratios; HAPPO employs sequential ratios with per-agent advantages; and single-agent reductions operating on factorized joint policies aggregate both into fully joint products. We formalize these two design choices as support matrices $\SA$ and $\SR$, and prove a canonical structure: the expected multi-agent policy optimization objective depends on the pair $(\SA,\SR)$ only through their matrix product $\tS=\SR\SA$. This yields two key consequences: (i) Redundancy: the two support matrices are interchangeable with respect to the signal, meaning neither aggregation pattern is inherently superior.(ii) Variance Ordering: the advantage aggregates rewards as a sum (additive variance with an interior bias-variance optimum at the coupling neighborhood), whereas the ratio aggregates likelihood ratios as a product (multiplicative variance that grows exponentially with support size, with no accompanying bias reduction). The resulting design principle is unambiguous: aggregate neighbors in the advantage, sized to the coupling neighborhood, and keep the ratio per-agent.
- 中文摘要
多智能体策略优化,以基于PPO的方法为代表,是合作式多智能体强化学习(MARL)的一个关键分支。一个核心设计问题是有多少邻居代理\脚注{本文中,“邻居”不仅指物理接近,还指行动相互影响的代理。}聚合以有效利用全球信息进行合作。这一决定必须从两个维度进行:优势(哪些代理人的奖励对信用信号有贡献)和比率(哪些代理人的似然比率构成截断重要权重)。现有方法在这两个轴线上分散且未被充分探索:IPPO分别处理两者;MAPPO将团队层面优势与每位代理人比例相结合;HAPPO采用顺序比率,具有每位代理人的优势;以及基于因果化联合保单的单一代理人减免,两者均汇总成完全联合产品。我们将这两个设计选择形式化为支持矩阵 $\SA$ 和 $\SR$,并证明了一个典型结构:期望的多智能体策略优化目标仅依赖于 $(\SA,\SR)$ 对,仅通过其矩阵积 $\tS=\SR\SA$。这导致两个关键后果:(i) 冗余性:两个支持矩阵相对于信号可互换,意味着任何聚合模式本质上都不优越。(ii) 方差排序:优势将奖励汇总为和(加法方差,且在耦合邻域处具有内部偏差-方差最优),而比率则将似然比汇总为乘积(乘法方差随支撑规模指数增长,且无伴随偏差减少)。由此产生的设计原则明确无歧义:将优势中的邻居聚合,大小为耦合邻域,并保持每个代理的比率。
A Geometric Perspective on Stabilizing Value Conflict Resolution
稳定价值冲突解决的几何视角
- Authors: Saket Reddy, Andy Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17946
- Pdf link: https://arxiv.org/pdf/2607.17946
- Abstract
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Geometrically, we show that CoT correlates with further smoothing the model's loss landscape in its sharpest direction, helping resolve the optimization instability of traditional scalar rewards. We also demonstrate via relevant downstream benchmarks that value conflict-focused CoT may generalize to different kinds of moral reasoning, demonstrating that this CoT has the potential to be an effective mechanism for better moral reasoning. To capitalize on this potential, we create a new value conflict-focused CoT design that further smooths the sharpest direction of the loss landscape and increases moral reasoning performance. This finding shows that explicitly modifying and improving the design of reasoning dynamics offers a promising avenue for improving model performance on user requests with complex value conflicts, advancing pluralistic alignment in LLMs.
- 中文摘要
大型语言模型(LLMs)在使用人类反馈强化学习(RLHF)的压缩标量奖励训练时,常常难以应对价值冲突。为应对这一挑战,我们探讨了思维链(CoT)推理如何帮助提升该领域的表现。几何上,我们表明CoT与模型损失景观在最锐利方向上进一步平滑相关,有助于解决传统标量奖励的优化不稳定性。我们还通过相关的下游基准展示了价值冲突导向的CoT可以推广到不同类型的道德推理,表明该CoT有潜力成为更好道德推理的有效机制。为了发挥这一潜力,我们创建了一种以价值冲突为核心的新CoT设计,进一步平滑了损失格局的最锐利方向,并提升了道德推理能力。这一发现表明,明确修改和改进推理动力学设计,为提升复杂值冲突用户请求的模型性能提供了有前景的途径,推动了大型语言模型中的多元对齐。
Information-Based Exploration via Random Features for Reinforcement Learning
基于信息的随机特征探索强化学习
- Authors: Waris Radji, Odalric-Ambrym Maillard
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.17981
- Pdf link: https://arxiv.org/pdf/2607.17981
- Abstract
Representation learning has enabled classical exploration strategies to be extended to deep Reinforcement Learning (RL), but often makes algorithms more complex and theoretical guarantees harder to establish. We introduce Random Feature Information Gain (RFIG), grounded in Bayesian kernel methods theory, which uses random Fourier features to approximate information gain and compute exploration bonuses in non-countable spaces. We provide error bounds on information gain approximation and avoid the black-box aspects of neural network-based uncertainty estimation, for optimism-based exploration. We present practical details that make RFIG scalable to deep RL scenarios, enabling smooth integration into standard deep RL algorithms. Experimental evaluation across diverse control and navigation tasks demonstrates that RFIG achieves competitive performance with well-established deep exploration methods while offering superior theoretical interpretation.
- 中文摘要
表征学习使经典探索策略得以扩展到深度强化学习(RL),但往往使算法更复杂,理论保证更难建立。我们介绍了随机特征信息增益(RFIG),基于贝叶斯核方法理论,利用随机傅里叶特征近似信息增益并计算不可数空间中的探索加成。我们提供了信息增益近似的误差界限,避免了基于神经网络的不确定性估计中的黑箱特性,以便基于乐观的探索。我们提供了使RFIG能够扩展到深度强化学习场景的实用细节,从而实现与标准深度强化学习算法的平滑集成。在多样化控制和导航任务中的实验评估表明,RFIG凭借成熟的深度探测方法实现了竞争性能,同时提供了更优越的理论解释。
PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning
PAMD:视觉强化学习中双模拟表示的结构化自适应距离
- Authors: Daegyeong Roh, Juho Bae, Han-Lim Choi
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.18004
- Pdf link: https://arxiv.org/pdf/2607.18004
- Abstract
Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by reward and transition similarity. In practice, the choice of the latent distance can strongly affect performance: using a fixed, pre-specified global norms (e.g., $\ell_p$ norms or other hand-designed metrics) may be overly restrictive to capture the behavioral distance. In contrast, unconstrained pairwise distances may admit degenerate solutions that drive the metric loss down without improving the representation. To address this gap, we introduce PAMD: Pairwise Adaptive Mahalanobis Distance, which parameterizes a positive-definite, pair-conditioned metric for measuring latent state similarity. PAMD is a simple plug-in for existing bisimulation-based methods, offering a more expressive yet structured alternative to fixed, pre-specified latent distances. We empirically validate our method on visual MuJoCo continuous-control tasks, where final performance of several recent bisimulation-based RL algorithms is substantially improved when equipped with the distance we propose.
- 中文摘要
许多视觉强化学习(RL)算法通过将潜在距离与由奖励和过渡相似性诱导的行为距离匹配来学习表征。实际上,潜在距离的选择会强烈影响性能:使用固定的预先指定的全局规范(例如,$\ell_p$范数或其他手工设计的指标)可能过于限制,难以捕捉行为距离。相比之下,无约束的两两距离可能会存在退化解,从而降低度规损失而不改善表示。为弥补这一空白,我们引入了PAMD:成对自适应马哈拉诺比斯距离,它为衡量潜态相似度提供了正定、配对条件的度量。PAMD是一个简单的插件,适用于现有基于双模拟的方法,提供了比固定预设潜距更具表现力且结构化的替代方案。我们在视觉MuJoCo连续控制任务中实证验证了该方法,在配备我们所提出的距离后,多个近期基于双模拟的强化学习算法的最终性能显著提升。
Generalised Bellman recurrence and three dualities in sequential decision-making
广义贝尔曼递归与顺序决策中的三对偶性
- Authors: Fernando E. Rosas, David Hyland, Daniel Polani
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18077
- Pdf link: https://arxiv.org/pdf/2607.18077
- Abstract
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.
- 中文摘要
是什么赋予了贝尔曼方程的其形式?我们证明最优值函数的递归性质由三个条件推导出:动力学通过足够的统计分解,返回可递归分解,以及不确定性聚合与两者兼容。当这三种条件都成立于共同态时,贝尔曼方程由它们的相互一致性产生;当一个失败时,通常可以通过增强状态或变形回报或动态来恢复可处理性。同样的条件也被证明会产生三种对偶性:概率与回报之间,回报与聚合之间,以及聚合与概率之间的对偶。我们的框架揭示了这些二元性源自单一构造,统一了强化学习、控制和决策理论中各自发展的方法。
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
稀疏证据可以满足:代理证据寻求多模态视频错误信息检测
- Authors: Haochen Zhao, Yongxiu Xu, Xinkui Lin, Dong Xie, Jiarui Lu, Yuqi Qian, Yubin Wang, Hongbo Xu, Gaopeng Gou
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18080
- Pdf link: https://arxiv.org/pdf/2607.18080
- Abstract
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from verification: first identifying sparse, decision-relevant clues and then judging veracity based on the acquired evidence. Accordingly, we propose SIEVE, a framework for Sparse Interactive Evidence Verification via Extraction in multimodal video misinformation detection. An evidence-seeking agent actively explores the available multimodal evidence and constructs a compact evidence package, which is then used by a verifier to determine veracity. The agent is trained with supervised evidence-seeking trajectories and an evidence-aware reinforcement learning objective that promotes informative evidence acquisition while discouraging unnecessary or invalid interactions. Experiments on multiple video misinformation benchmarks show that SIEVE consistently outperforms the evaluated baselines and supports reliable verification using compact evidence packages. Moreover, the resulting acquisition process provides an explicit and inspectable evidence trail, improving the transparency and groundedness of multimodal misinformation detection.
- 中文摘要
多模态视频虚假信息检测通常被表述为一种整体视频理解任务,即一次性处理和判断整个视频及其相关内容。然而,现实中的错误信息往往表现出稀疏且构成性的证据结构:可靠的判断可能仅依赖少数耦合线索,而大多数视频内容提供的额外信息有限。穷尽多模态推理因此可能引入大量冗余和模糊的决定性证据。这促使证据获取与验证脱钩:先识别稀疏且与决策相关的线索,然后根据获得的证据判断真实性。因此,我们提出了SIEVE,这是一个用于多模态视频错误信息检测中稀疏交互式证据核查的框架。证据寻求代理主动探索可用的多模态证据,构建一个紧凑的证据包,验证者随后用它来判断真实性。代理接受监督式证据寻求轨迹和证据感知强化学习目标的训练,促进信息性证据获取,同时阻止不必要或无效的互动。对多个视频错误信息基准测试的实验表明,SIEVE始终优于评估基线,并支持使用紧凑证据包进行可靠的验证。此外,由此产生的采集过程提供了明确且可检验的证据痕迹,提升了多模态错误信息检测的透明度和稳固性。
Importance Sampling and PCA for Finding Failures in Commercial Autonomous Vehicles
重要性抽样和PCA用于发现商用自动驾驶车辆故障
- Authors: Hailey Warner, Duncan Eddy, Shreya Parjan, Caroline Cahilly, Harrison Delecki, Matthias Kleinstauber, Chaitanya Shinde, Jerry Lopez, Mykel J. Kochenderfer
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.18106
- Pdf link: https://arxiv.org/pdf/2607.18106
- Abstract
Methods for discovering rare failures in autonomous systems have so far been demonstrated almost exclusively in simulations with simple, academic driving stacks, leaving open whether they generalize to the more robust planners used in commercial systems. We address this gap by applying two rare-event discovery algorithms to a commercial autonomous trucking stack. Adaptive stress testing (AST) uses reinforcement learning to search for the most likely noise trajectories leading to a simulated collision, while diffusion-based failure sampling (DiFS) trains a denoising diffusion model to sample a diverse set of failures. We show that both algorithms find simulated collisions during merge and cut-in maneuvers where traditional Monte Carlo simulation does not. To make these failures actionable, we introduce a statistical analysis based on principal component analysis (PCA) that classifies failures into common modes and identifies the timesteps that most influence the outcome. We cluster the principal components and invert the PCA transform to recover generalized noise trajectories, and show that these trajectories reproduce failures in identical and similar scenarios. This provides a path from failure discovery to systematic diagnosis of perception-level flaws.
- 中文摘要
迄今为止,发现自动系统罕见故障的方法几乎完全通过简单学术驱动栈的仿真展示,是否推广到商业系统中使用的更稳健的规划器尚不确定。我们通过将两种罕见事件发现算法应用于商业自动驾驶卡车运输堆栈来弥补这一空白。自适应应力测试(AST)利用强化学习寻找导致模拟碰撞的最可能噪声轨迹,而基于扩散的失效采样(DiFS)则训练去噪扩散模型,以采样多样化的失败。我们证明,这两种算法都能在合并和插入机动中发现模拟碰撞,而传统蒙特卡洛模拟则无法做到。为了使这些失效可操作,我们引入了基于主成分分析(PCA)的统计分析,将失效分类为常见模式,并识别对结果影响最大的时间步长。我们聚类主成分并反演PCA变换以恢复广义噪声轨迹,并证明这些轨迹在相同且相似的场景下重现失效。这为从发现失误到系统诊断感知层面缺陷提供了一条路径。
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
LLM作为教练:针对不可验证任务的体验式学习
- Authors: Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.18110
- Pdf link: https://arxiv.org/pdf/2607.18110
- Abstract
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.
- 中文摘要
开放式任务的强化学习(RL)将基于评分标准的评估压缩为标量奖励,丢弃丰富的文本反馈,并将回答与不同质量特征混为一谈。我们提出体验式学习(EL),将将LLM作为评委的反馈模式转变为LLM作为教练。教练将对每一项政策回应的评估提炼为可转移的体验知识,这为教师模型提供了条件,并通过政策上下文的提炼内化。与标量奖励相比,这种更高带宽的反馈通道提供了密集的监督,并保持了高质量响应中细致的偏好。在两大政策家族中,无论是政策本身的反馈还是专有模型,EL在未完成和看不见的开放式任务上,始终优于基于评分标准的强化学习。值得注意的是,EL在训练分布之外的推广性更好,并且减少了奖励黑客行为。这些发现确立了体验式知识作为一种更丰富且更具推广性的学习信号,用于非验证任务的培训后。
Isaac Sim-to-Real: Reinforcement Learning based Locomotion for Quadrupeds
Isaac模拟到现实:基于强化学习的四足动物运动
- Authors: Jordan Dowdy, Jean Chagas Vaz
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.18135
- Pdf link: https://arxiv.org/pdf/2607.18135
- Abstract
Learning-based approaches to locomotion have risen in popularity in recent years, showing the capability for complex legged locomotion and whole-body control. Reinforcement learning (RL), the primary learning-based approach for locomotion, often utilizes a high-performance simulation tool, providing a controlled and efficient training and development environment. However, policies that perform well in simulation frequently encounter unexpected challenges when deployed on a physical system, known as the sim-to-real gap. This work presents a robust RL locomotion framework capable of whole-body control. The proposed RL framework utilizes Nvidia's new set of simulation tools, Isaac Sim, and its companion RL framework, Isaac Lab, for training, achieving a zero-shot sim-to-real policy. The performance of our policy is validated on physical hardware using the Unitree Go1, with experimental results showing similar velocity tracking performance to the quadruped's integrated controller, with a greater ability to recover from large disturbances, and achieve linear velocities of 2.0 m/s and angular velocities of 1.8 rad/s.
- 中文摘要
近年来,基于学习的运动方法日益流行,展示了复杂腿部运动和全身控制的能力。强化学习(RL)是主要基于学习的移动方法,通常使用高性能模拟工具,提供受控且高效的训练与发展环境。然而,在仿真中表现良好的策略在部署于物理系统时常常会遇到意想不到的挑战,即所谓的模拟到现实差距。这项工作提出了一个能够全身控制的强化学习运动框架。所提的强化学习框架利用了英伟达的新模拟工具Isaac Sim及其配套的强化学习框架Isaac Lab进行训练,实现了零机会模拟到真实的策略。我们的策略性能在使用Unitree Go1的物理硬件上验证,实验结果显示其速度追踪性能与四足机的集成控制器相似,且更能从大扰动中恢复,实现2.0 m/s的线速度和1.8 rad/s的角速度。
Keyword: diffusion policy
HyperDCM: Dynamic Cluster Memory Replay in Hyperbolic Space for Continual Robotic Navigation Across Scenes
HyperDCM:双曲空间中的动态集群记忆回放,实现场景间持续机器人导航
- Authors: Zhengfei Lu, Jian Yang, Muyu Wang, Shaowen Chen, Jinpeng Mi, Ke Li, Xiong You, Qi Wu, Xuan Tang, Xian Wei
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.16267
- Pdf link: https://arxiv.org/pdf/2607.16267
- Abstract
Continual learning in visual navigation remains challenging due to catastrophic forgetting and the difficulties associated with adapting to diverse and evolving environments. To address these issues, we propose Hyperbolic Dynamic Cluster Memory (HyperDCM), a structure-aware memory mechanism that enhances diffusion policy-based navigation through scene graph modeling and principled memory replay. HyperDCM extracts semantic scene triples from RGB observations using large vision-language models, encodes them into scene graph embeddings via a Relational Graph Convolutional Network (R-GCN), and projects the embeddings into hyperbolic space to enhance structural separability and retention in continual navigation. A dynamic clustering and structure-sensitive update strategy selects representative samples for memory replay, thereby preserving knowledge diversity and mitigating catastrophic forgetting. Experiments on multi-scene indoor and outdoor datasets demonstrate that HyperDCM achieves superior retention of past navigation capabilities and improved generalization compared to representative continual learning baselines adapted to diffusion policy navigation.
- 中文摘要
由于灾难性遗忘和适应多样且不断变化的环境带来的困难,视觉导航的持续学习依然充满挑战。为解决这些问题,我们提出了双曲动态集群存储器(HyperDCM),这是一种结构感知的记忆机制,通过场景图建模和原则性内存回放增强基于扩散策略的导航。HyperDCM利用大型视觉语言模型从RGB观测中提取语义场景三元组,通过关系图卷积网络(R-GCN)将其编码为场景图嵌入,并将嵌入投影到双曲空间,以增强结构可分离性和持续导航的保留性。动态聚类和结构敏感更新策略选择具有代表性的样本进行记忆回放,从而保持知识多样性并减少灾难性遗忘。多场景室内外数据集的实验表明,HyperDCM相比适应扩散策略导航的代表性持续学习基线,在保留过去导航能力方面更为出色,泛化能力也更佳。
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
通过延迟感知引导融合实现异步多模态扩散策略组合
- Authors: Zihao He, Hongjie Fang, Shirun Tang, Cewu Lu, Haoshu Fang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.17257
- Pdf link: https://arxiv.org/pdf/2607.17257
- Abstract
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to improve manipulation performance. However, these modalities often differ not only in information content but also in sensing rates and inference latencies. Existing multimodal diffusion policies typically rely on synchronous fusion or manually designed multi-frequency architectures, which either slow down high-frequency feedback or limit extensibility to new modality combinations. We propose LAG-Fusion, a latency-aware guidance fusion framework for asynchronous multimodal diffusion policy composition. LAG-Fusion allows modality-specific policies to operate at their native inference rates and contribute denoising guidance whenever available. To make asynchronous composition consistent, we derive a reference-frame rebasing rule for diffusion variables under relative action representations, enabling delayed guidance to be aligned before fusion. We instantiate LAG-Fusion in contact-rich manipulation by composing a low-frequency vision policy with a high-frequency force policy. Experiments under heterogeneous modality latencies show that LAG-Fusion improves policy responsiveness and task performance over synchronous fusion and specially designed force-aware baselines.
- 中文摘要
扩散政策显示出机器人模仿学习的强大潜力,近期扩展还加入了更多方法以提升操作性能。然而,这些模态不仅在信息内容上不同,还在感知速率和推理延迟上也存在差异。现有的多模态扩散策略通常依赖同步聚变或手动设计的多频架构,这些架构要么减慢高频反馈速度,要么限制对新模态组合的扩展性。我们提出了LAG-Fusion,一种用于异步多模态扩散策略组合的延迟感知引导融合框架。LAG-Fusion允许特定模态的策略以其原生推断率运行,并在可用时提供去噪指导。为了使异步合成一致,我们推导出一个参考系重基规则,适用于相对作用表示下的扩散变量,从而在融合前对准延迟导引。我们通过将低频视觉策略与高频力策略组合,实现了接触丰富操控中的LAG-Fusion。在异构模态延迟下的实验表明,LAG-Fusion相比同步融合和专门设计的力感知基线,在策略响应性和任务性能上有所提升。