生成时间: 2026-10-06 00:55:27 (UTC+8); Arxiv 发布时间: 2026-10-05 20:00 EDT (2026-10-06 08:00 UTC+8)
今天共有 45 篇相关文章
Keyword: reinforcement learning
Lexicographic Multi-Objective On-Policy Distillation
词典序多目标政策提炼
- Authors: Doseok Jang, Jon Ander Campos, Youran Qi
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02359
- Pdf link: https://arxiv.org/pdf/2610.02359
- Abstract
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
- 中文摘要
可验证奖励强化学习(RLVR)通常能优化答案的正确性,但有用的语言模型行为同样需要高质量的推理和简洁的回答。现有的多奖励后训练方法通常会对奖励进行等级化或合并专家,但未明确保护奖励优先顺序。当权衡不对称时,这就存在问题:例如,简洁性不应以牺牲正确性为代价提升。我们引入了词典序多目标策略提炼(LMOPD),这是一种多教师方法,用于在明确优先级下整合奖励专用策略。每次学生推广,LMOPD会选择第一个目标的专家,其门检测到缺陷,然后局部投影其中心对数策略修正,以去除与高优先级专家相反的组件。我们在三项数学基准测试下,评估30B-A3B专家混合变压器模型,在两位和四位专家的设置下,衡量保留的专家增益。在两位专家的情况下,LMOPD的点估计完全保留了准确性和推理质量的提升,同时获得了46.9%美元的简洁性提升。四位专家时,LMOPD保留了约90%%的准确性和推理正确性提升,而下一个最佳评估基线仅约57%%。匹配的四专家分解显示,词典序路由优于随机路由,且预测进一步强化了两种最高优先级能力。在这两个量表中,LMOPD比我们评估的现有基线更有效地保留了最高优先级的能力,展示了明确优先级对专业整合的价值。
Energy Saving in 5G and Beyond Networks: A Quantum Reinforcement Learning Approach
5G及更远网络中的节能:一种量子强化学习方法
- Authors: Muhammad Usman, Nguyen Van Huynh, Marianna Lezzi, Mariangela Lazoi
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02403
- Pdf link: https://arxiv.org/pdf/2610.02403
- Abstract
Energy saving has become a critical challenge in 5G and beyond networks. The rapid growth of connected devices has increased the overall network energy demand, driving operational expenditure to unsustainable heights. The Base Station (BS) accounts for the largest share of energy usage, typically consuming around 60-70\% of the Radio Access Network (RAN)'s total energy. Therefore, to address this issue, this article optimizes the BS's energy usage while accounting for the dynamic behavior of User Equipment (UE). Deep Reinforcement Learning (DRL) is a natural candidate for determining effective energy saving policies, such as automatically switching BSs on or off when user density is low or adjusting transmission power to balance energy efficiency and Quality of Service (QoS). However, its heavy training burden and the exponential growth of state and action spaces in dense 5G environments make exploration increasingly difficult. To overcome these limitations, we introduce a novel Quantum Reinforcement Learning (QRL) algorithm that leverages quantum principles, including superposition and entanglement, through parameterized quantum circuits, enabling significantly faster convergence than DRL, which relies on conventional deep neural networks. Extensive simulations demonstrate that the proposed QRL can substantially reduce energy consumption while maintaining QoS, even when UEs are highly dynamic and frequently switch their association with BS antennas. Additionally, QRL consistently outperforms DRL and Q-Learning in both convergence speed and learning complexity.
- 中文摘要
节能已成为5G及其他网络领域的关键挑战。连接设备的快速增长增加了整体网络能源需求,使运营支出达到不可持续的高度。基站(BS)占能耗最大份额,通常占无线接入网(RAN)总能耗的60%-70%。因此,为解决这一问题,本文在考虑用户设备(UE)动态行为的前提下优化BS的能耗使用。深度强化学习(DRL)是确定有效节能策略的自然候选方法,例如在用户密度低时自动开关BS,或调整传输功率以平衡能效和服务质量(QoS)。然而,其沉重的训练负担以及在密集5G环境中状态和动作空间的指数增长,使探索变得越来越困难。为克服这些限制,我们引入了一种新型量子强化学习(QRL)算法,利用包括叠加和纠缠在内的量子原理,通过参数化的量子电路实现了显著快于依赖传统深度神经网络的DRL的收敛速度。大量模拟表明,即使UE高度动态且频繁切换与BS天线关联,QRL也能在保持QoS的同时大幅降低能耗。此外,QRL在收敛速度和学习复杂度上始终优于DRL和Q-Learning。
Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments
核物理散射实验中靶极化优化的强化学习技术
- Authors: Armen Kasparian, Torri Jeske, Monibor Rahman, Chris Keith, James Maxwell, Thomas Britton, Malachi Schram, David Lawrence
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02452
- Pdf link: https://arxiv.org/pdf/2610.02452
- Abstract
The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and evolving material properties, a task that is traditionally performed through manual trial-and-error by expert operators. This work presents a data-driven control framework that combines surrogate modeling with reinforcement learning to optimize the target polarization. Using operational data from the APOLLO cryogenic target system, we train and evaluate multilayer perceptron and Gaussian process regression models to predict polarization as a function of microwave frequency, beam current, and accumulated radiation dose. We show that Gaussian process-based models provide calibrated uncertainty estimates and reliably identify regions outside the training distribution, while MLPs exhibit limited sensitivity to distributional shift. To enable learning and control across multiple target samples, we introduce a Gaussian process approximation and embed the surrogate model within a standardized simulation environment. A reinforcement learning agent is trained using a lower-confidence-bound reward formulation that balances performance maximization against uncertainty. We are able to show an almost 2x improvement on the operators actions utilizing our RL agent.
- 中文摘要
核物理实验中动态偏振靶的运行依赖于微波频率的持续调谐,以补偿辐射损伤和材料性质演化,这一任务传统上由专家操作员通过手工试错来完成。本研究提出了一个数据驱动的控制框架,结合了替代建模与强化学习,以优化靶极化。利用APOLLO低温靶系统的操作数据,我们训练并评估多层感知器和高斯过程回归模型,预测偏振随微波频率、束流和累积辐射剂量的变化。我们表明,基于高斯过程的模型提供了校准的不确定性估计,并可靠地识别训练分布外的区域,而MLP对分布偏移的敏感性有限。为实现多目标样本的学习和控制,我们引入了高斯过程近似,并将替代模型嵌入标准化模拟环境中。强化学习代理采用较低置信度界限的奖励公式进行训练,平衡性能最大化与不确定性。我们能够证明操作员使用强化学习代理的操作几乎提升了2倍。
Tropical Reinforcement Learning
热带强化学习
- Authors: Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02478
- Pdf link: https://arxiv.org/pdf/2610.02478
- Abstract
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
- 中文摘要
大型语言模型的强化学习通常最大化期望回报,将所有成功轨迹的概率加总。然而,经典求和表述只能报告模型策略成功的频率,而不能报告具体哪种解法实际有效,且由于概率和为1,强化一个解可能会让模型忘记另一个从未被证明错误的解。这使得期望回馈不适合组合推理,因为解必须由模型在多次尝试中产生但很少一起产生的推理步骤组装而成。为此,我们提出了热带强化学习,它基于简单的代数变更:不加备选解的概率,而是取其最大值,从而得到热带半环。状态的值随后成为其最可能验证解的对数概率,并附带一条可重放和重复使用的显式路径。这使得真正的组合成为可能,因为即使来自不同的部署,最佳前缀和后缀相遇也可以在共享状态上合并。为了实现这一点,我们引入了TROPIC,一种针对确定性、可重置环境且结果可验证的训练算法。在四个代理任务(Sokoban、Countdown、FrozenLake、WebShop)上,TROPIC的表现比最强的策略基线高出多达16个百分点。改变强化学习的代数,而不仅仅是其估计量,可以显著提升语言模型中的组合推理能力
Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
多保真度策略梯度稳定数据稀缺的强化学习
- Authors: Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.02505
- Pdf link: https://arxiv.org/pdf/2610.02505
- Abstract
Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develop MFPG for modern actor-critic learning in GPU-parallel simulation and on a physical robot. Our analysis and experiments show that naive extensions to proximal policy optimization (PPO) can lose cross-fidelity correlation or inflate variance. Our MFPG-PPO addresses these failures by redesigning the sampling, advantage estimation, and control variate construction to preserve cross-fidelity correlation, and by monitoring estimator uncertainty to prevent variance inflation. We also introduce a budget-aware MFPG-PPO to divide a fixed sampling budget among high- and low-fidelity data sources. Across simulated robot locomotion tasks of varying LF-to-HF transfer difficulty and HF data budgets, MFPG-PPO improves upon PPO trained on HF data alone in nearly all settings, and consistently matches the performance of PPO trained with 16x more HF data on the hardest task at the smallest HF budgets. In contrast, most baselines that use LF data perform well only where direct LF-to-HF transfer succeeds. MFPG-PPO enables stable learning on a physical Franka arm using only 4 real-robot episodes per update and no human demonstrations.
- 中文摘要
当昂贵且稀缺的目标域数据产生噪声梯度估计时,策略梯度方法可能变得不稳定。我们通过用丰富、廉价但有偏的低保真度(LF)数据补充有限的高保真度(HF)目标域数据,例如简化模拟器,来应对这一挑战。大多数现有方法直接基于LF数据优化偏置目标。相比之下,新近推出的多保真度策略梯度(MFPG)框架仅利用LF数据构建控制变量,降低方差并提高HF数据效率,同时不影响策略梯度估计量。然而,MFPG已发表的研究仅限于REINFORCE在小规模模拟任务中。我们开发MFPG用于GPU并行模拟和物理机器人中的现代actor-critic学习。我们的分析和实验表明,对近保真度优化(PPO)的朴素扩展可能会失去交叉保真度相关性或膨胀方差。我们的MFPG-PPO通过重新设计抽样、优势估计和控制变量构建来维护交叉保真度相关性,并监控估计器不确定性以防止方差膨胀,从而解决了这些失败。我们还引入了预算感知型MFPG-PPO,用于将固定采样预算分配到高保真度和低保真度数据源之间。在模拟机器人移动任务中,LF到HF传输难度和HF数据预算各异,MFPG-PPO在几乎所有设置下均优于仅用HF数据训练的PPO,并在最困难任务和最小HF预算下,性能始终与用16倍多HF数据训练的PPO相当。相比之下,大多数使用LF数据的基线只有在直接LF到HF传输成功的情况下表现良好。MFPG-PPO支持在Franka机械臂上稳定学习,每次更新仅需4集真实机器人视频,无需人工演示。
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
学习下一步调查内容:长远研究代理的元推理
- Authors: Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02525
- Pdf link: https://arxiv.org/pdf/2610.02525
- Abstract
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.
- 中文摘要
随着证据积累,长视野研究代理必须决定如何调查以及下一步调查什么。这很难学习,因为在长执行轨迹中此类决策稀少,其后果可能在多个调查后显现。我们引入迭代研究代理的元推理(MIRA),这是一种将研究分配与执行分开的层级架构。外环元推理器从持久的研究记录中策划上下文,然后为下一次调查编写工作单或结束该集。新的内环执行者执行每个工作订单,使执行成为元推理动作之间过渡的一部分。在没有策略培训的情况下,MIRA提升了长视野推断能力,并在定理证明和开放式神经架构研究中更有效地分配额外计算量。其决策边界还为功劳分配提供了自然单位。在每个边界处,我们训练生成批评者预测部分状态的预期剩余收益,表现优于代币级替代方案。跨环境预训练提升预测和适应能力,产生可转移的先验以评估部分进展。我们利用此先验初始化MIRA-AC,一个生成行为者-批评者,联合训练以预测剩余收益并选择下一个调查,无需独立批评模型。MIRA-AC将策略优化集中于元推理决策,实现高效的长视野强化学习,而无需直接优化其发起的较长执行轨迹。基于模型自身代理爬山信号训练MIRA-AC,可提升四个自动研究环境中的黄金性能;该行为者通过跨环境值初始化进行转移。综合来看,这些结果表明元推理可以作为明确策略学习,用于指导长视野自治研究。
Reward Inflation: A Healthy Stimulus for Reinforcement Learning
奖励膨胀:强化学习的健康刺激
- Authors: Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02545
- Pdf link: https://arxiv.org/pdf/2610.02545
- Abstract
Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
- 中文摘要
奖励是强化学习(RL)中的主要学习信号。然而,虽然奖励强度通常在整个训练过程中保持固定,但其时间调制机制仍未被充分探索。本文提出了奖励膨胀,即在训练过程中奖励逐渐递增的现象,并证明它可以作为强化学习的健康刺激。理论上,奖励膨胀会诱导隐含的近期权重,在政策更新期间加重近期的过渡,从而加快适应速度。我们还进一步证明,通过在政策饱和时保持梯度信号,奖励膨胀抑制了休眠神经元的出现,有助于保持可塑性。关于ALE博弈和MuJoCo任务的实证结果证实了这些发现,表明适当的奖励膨胀水平能惠及广泛的任务。最后,我们引入了美联储,这是一种适应性变体,可以随时调整通胀水平,并发现它常常优于固定通胀。
Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
通过分解值梯度流进行测试时间多代理协调
- Authors: Dongsu Lee, Haoran Xu, Amy Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.02554
- Pdf link: https://arxiv.org/pdf/2610.02554
- Abstract
Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.
- 中文摘要
离线多智能体强化学习(MARL)面临持续性的权衡。表达式生成策略可以在数据中表示多模态协调,但无法区分高价值区域;而价值优化策略利用已学习到的Q函数,但将多模态压缩为单一主导模式。单代理的模式崩溃可能破坏联合协调,同时跨代理漂移则可能将联合策略推入动作空间中未见的区域。我们提出通过最优统一传输(SCOUT)实现可扩展协调,这是首个将生成基础模型与测试时间动作细化学习得值函数结合的离线MARL框架。SCOUT训练两个解耦组件:一个流匹配行为先验和一个分解值函数。测试时,它通过Stein变分梯度下降将行为样本传输到高价值区域。传输步数控制自适应测试时间尺度,取代固定正则化系数。根据个人-全局最大值(IGM)原理,我们证明了联合软值缺口上的单项KL界限,随着传输收敛而消失,且存在与IGM违背成正比的不可约加残。实证上,SCOUT在离散和连续离线MARL基准测试中实现最佳平均性能,并在所有离线到在线配置中均有性能提升。
Test-time Calibration Learning for Large Language Model Reasoning
大型语言模型推理的测试时校准学习
- Authors: Zizhuo Zhang, Xiong Peng, Jingwei Sun, Rong Yao, Shixiong Kai, Mingxuan Yuan, Bo Han
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02695
- Pdf link: https://arxiv.org/pdf/2610.02695
- Abstract
Reliable large language models (LLMs) must not only produce accurate answers but also express confidence that faithfully reflects their probability of being correct. Such calibration is essential for identifying uncertain predictions and supporting reliable decision-making in real-world deployment. Recent studies incorporate calibration learning into reinforcement learning (RL), jointly optimizing answer correctness and verbalized confidence using ground-truth correctness supervision. However, their reliance on labeled data limits their applicability in practical test-time settings, where ground-truth labels are unavailable and calibration may need to adapt to newly encountered target tasks. To address this challenge, we propose Test-Time Calibration Learning (TTCL), a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data. Specifically, TTCL derives self-supervision signals for both correctness and calibration from multiple model-generated responses, enabling calibration learning at test time without ground-truth labels. Theoretical analysis further establishes TTCL as a bounded surrogate for the ideal calibration objective. Extensive experiments on mathematical reasoning and factual question answering demonstrate that TTCL consistently improves both accuracy and calibration across diverse models and tasks. On base models, TTCL achieves an average relative accuracy improvement of +40.13% and an ECE reduction of +70.80% across eight benchmarks. Moreover, TTCL can further improve both accuracy and calibration for already calibrated models under domain shift, particularly when source-domain calibration transfers poorly to target tasks. In the math-to-factQA setting, TTCL achieves an average relative accuracy gain of +20.35% and reduces ECE by +53.83%. The source code is released at this https URL.
- 中文摘要
可靠的大型语言模型(LLM)不仅必须产生准确答案,还必须表达忠实反映其正确概率的信心。这种校准对于识别不确定的预测和支持现实应用中的可靠决策至关重要。近期研究将校准学习纳入强化学习(RL),通过基础真实性监督共同优化答案正确性和口头置信度。然而,它们对标记数据的依赖限制了其在实际测试时间环境中的适用性,因为缺乏地面真实标签,校准可能需要适应新遇到的目标任务。为应对这一挑战,我们提出了测试时间校准学习(TTCL)框架,该框架结合推理准确性和口头置信度直接调整未标记目标任务数据。具体来说,TTCL能够从多个模型生成的响应中推导出正确性和校准的自监督信号,使测试时无需地面真实标签即可校准学习。理论分析进一步确立了TTCL作为理想校准目标的有界替代品。大量数学推理和事实性问题解答实验表明,TTCL在不同模型和任务中持续提升准确性和校准。在基模型上,TTCL在八个基准测试中平均相对准确率提升+40.13%,ECE降低+70.80%。此外,TTCL还能进一步提升已校准模型在域位移下的准确性和校准,尤其是在源域校准对目标任务传递不佳时。在数学到事实质询环境中,TTCL的平均相对准确率提升为+20.35%,并将ECE降低+53.83%。源代码发布于此 https URL。
Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
从演进错误中学习:策略上蒸馏的自适应迭代修复
- Authors: Rui Li, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02700
- Pdf link: https://arxiv.org/pdf/2610.02700
- Abstract
On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
- 中文摘要
策略上自我蒸馏(OPSD)提供密集的令牌级反馈,针对学生自身策略中抽样的轨迹,比强化学习的结果级奖励更为丰富的训练信号。这种反馈来自教师,教师基于学生无法获得的完整参考解。参考解指定目标,但不说明如何从学生当前错误向目标移动,形成解条件的捷径风险。我们引入AIR-OPD,一种自适应迭代修复框架,用于策略中提炼,提供错误到修复的监督。给定失败响应,指导生成器合成当前错误的修复指导。学生用该指导采样策略重试。如果重试仍不正确,生成器为新观察到的错误生成新的修复指导。每轮中,固定教师作为特权上下文获得指导,并在学生最新失败回答中错误对齐的区域进行监督。结果感知阶段权重偏向早期修复阶段和即时重试通过验证的学分阶段。我们基于DAPO-Math-17K数据集训练AIR-OPD,并评估AIME24、AIME25和HMMT25,同时在MMLU-Pro和GPQA上进行非发行测试。我们考察了两个指导来源:来自当前学生政策的自我指导和来自更大模型的外部指导。对于Qwen3-4B和Qwen3-8B,AIR-OPD均获得最佳数学推理平均值,较最强基线提升最多3.6分,同时保持非发行基准测试的基础模型表现。
Bellman Error Minimization Via Linear Programming Normalization
通过线性规划归一化实现贝尔曼错误最小化
- Authors: Haining Yu
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02730
- Pdf link: https://arxiv.org/pdf/2610.02730
- Abstract
This paper proposes a new functional approximation approach to reduce Bellman error in high-dimensional dynamic programming and Reinforcement Learning problems. Using a classic dynamic programming problem (network capacity control in revenue management) as the motivational example, the paper illustrates that deep neural networks and linear programming approximation algorithms can be combined to derive approximate solutions to dynamic programming problems. Simulation results show the proposed approximation algorithms achieves competitive performance when compared with benchmark.
- 中文摘要
本文提出了一种新的函数近似方法,以减少高维动态规划和强化学习问题中的贝尔曼误差。以经典动态规划问题(收入管理中的网络容量控制)为动机示例,论文展示了深度神经网络与线性规划近似算法的结合,可以推导动态规划问题的近似解。模拟结果显示,所提出的近似算法在与基准测试相比具有竞争力。
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
前瞻性回顾:通过预测与现实差距进行自我校准的强化学习
- Authors: Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02740
- Pdf link: https://arxiv.org/pdf/2610.02740
- Abstract
Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.
- 中文摘要
长视野智能体的强化学习完全依赖回溯性训练信号:仅在观察环境后果后才给予功劳,智能体在动作时的信念对梯度是隐形的。我们引入了前瞻性回顾(PH),这是一种自我校准训练原理,通过基于智能体未来预测(反馈前)与回顾性评估(反馈后)之间间隙的信号,补充任何回顾性基础方法。这种每次推出的惊喜识别了智能体自模型最不准确的样本,并通过停止梯度惊喜加权优势放大了其梯度贡献。由于前瞻性预测变量与策略共享参数,两者协同演化,逐渐聚焦于智能体剩余的盲点。我们将该原则与特权信息差距联系起来,并展示了最小化惊讶残差为代理误校准率提供了下降路径;因此,校准是优化的副产品,而非额外目标。在单回合可验证任务和多回合个人代理任务(在GRPO、策略上提炼及其组合下),PH在模型尺度上均能提升任务性能和校准,实现一致的提升。值得注意的是,主导的误校准模式在结构性间转换,单回合中过度自信失败,多回合中信心不足成功,但同一训练原则成功应对两者。
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
强化学习前的门诊:基于评分标准的暖启动强化学习,配合政策提炼
- Authors: Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak, Yun He, Richard Yuanzhe Pang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02781
- Pdf link: https://arxiv.org/pdf/2610.02781
- Abstract
Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.
- 中文摘要
许多有用的语言模型任务无法通过精确的结果验证来评估。基于评分标准的评分标准强化学习(RL)解决了这一问题,使开放式回答与明确标准进行评分。然而,由于奖励是在完整回答之后分配的,训练信号并未直接指出哪些个人决策对最终得分产生了贡献。我们提出了一个两阶段的培训框架,首先将评分标准作为教师的特权上下文进行密集的代币级监督,然后作为进一步强化学习的奖励。第一阶段是评分标准特权的政策提炼(RP-OPD),没有评分标准的学生会匹配一位有评分标准的教师在学生生成前缀下的下一枚代币分布。第二阶段,强化学习直接优化评分标准奖励,并在观察到的提炼瓶颈中有所提升。我们利用开放权重模型评估该框架在健康和科学任务中的表现。在HealthBench、ResearchQA和RubricHub Science中,我们比较了培训后方法,并在强化学习前调整SFT或RP-OPD培训的量,发现我们的两阶段框架在评估方法中得分最高。RP-OPD + RL在RubricHub Science上显示出有限的奖励黑客迹象,而SFT + RL基线则因声称符合评分标准但未提供必要内容而获得高额奖励。这些发现支持使用评分标准指导政策提炼,然后再应用基于评分标准的强化学习。
VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
VIGOR:基于模型的强化学习中通过潜在空间一致性实现零样样视觉推广
- Authors: Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.02801
- Pdf link: https://arxiv.org/pdf/2610.02801
- Abstract
Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.
- 中文摘要
基于模型的强化学习(MBRL)通过在学习的潜在动态内进行规划,实现了强的样本效率,但在背景变化、光照变化或摄像机移动等未见视觉干扰下,其性能会显著下降。与无模型强化学习(编码器扰动仅影响单步预测)不同,MBRL存在两层弱点:视觉干扰首先将编码器输出推出分布范围,随后这些误差通过递归潜在的扩展在规划视野中叠加。我们提出通过潜在空间一致性实现模型驱动RL(VIGOR)的视觉泛化,该框架实现零样本推广至看不见的视觉干扰,同时保持MBRL骨干的样本效率。VIGOR集成了三个相互依赖的组件:(i)非对称弱强增强,将仅弱和弱至强潜在视图配对于单一批;(ii)动态层级一致性,通过直接潜在回归强制增强不变转移预测;(iii)编码器层级稳定,防止在动态层一致性施加的交叉增强监督下编码器漂移。对DeepMind控制套件(DMC)和Robosuite的评估显示,VIGOR优于最先进的无模型和基于模型的基线,在DMC上领先第二优基线3.4%,在Robosuite中超过43.6%。消融进一步表明,VIGOR的鲁棒性与增强无关:用不同扰动家族的替代增强能保持强强的泛化性,确认驱动稳健性的是由潜在空间一致性而非增强选择。
Text-Centric Post-Training for Omni-Modal Reasoning
全模态推理的文本中心后期培训
- Authors: Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Qimin Wu, Jingru Fan, Chen Qian, Yanfeng Wang, Yu Wang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02819
- Pdf link: https://arxiv.org/pdf/2610.02819
- Abstract
Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.
- 中文摘要
提升全域大型语言模型中的联合视听推理通常会产生大量数据构建和训练成本。我们的诊断显示,尽管所有对应的单跳问题均正确,但多跳推理仍存在困难,并暗示感知和推理目标局部优化存在部分解耦。这激励了训练后对这些能力的不同重视。纯文本推理训练在数据源、模型尺度和族群中均有显著提升。在表现最佳的纯文本配置中,监督微调及强化学习(RL)使Qwen2.5-Omni-7B的九个推理分数几何平均值较基础模型提升25.83%,以减少56.6%的GPU小时优于完整的原生视听路径。在完全由文本生成的LLM上训练时,在构建或训练中没有视听数据的情况下,这一几何平均值提高了21.01%。然而,仅文本训练会降低感知能力。因此,我们提出了以文本为中心的后训练范式:纯文本训练提供主要的推理优化,减少数据数据的原生视听RL则进一步优化感知。精细化使用约90%的输入令牌数量比全数据视听RL少,恢复了基础层级以上的感知,并保留了93.5%的表现最佳纯文本流水线的推理收益。
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
元评分标准:学习基于评分标准的强化学习奖励
- Authors: Yuxuan Fan, Jaehong Yoon
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02824
- Pdf link: https://arxiv.org/pdf/2610.02824
- Abstract
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00--20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
- 中文摘要
基于评分标准的强化学习将奖励驱动优化扩展到开放式任务,通过给个别响应要求部分积分。然而,评分标准评委即使所需信息或行动在回答中缺失,也能给予高评分,这种失败模式我们称之为空性信用。此类奖励在所需信息被移除后依然存在,并可能逆转响应的GRPO优势符号。为解决这个问题,我们引入了MetaRubric,它交替将证据感知的策略优化与响应引导的评分调整交替进行。我们通过在每个提示中更改一个任务相关事实来构建反事实对应。在策略优化过程中,只有当响应包含足够证据满足所需评分标准时,才会给予认可。在每个策略优化阶段结束后,当前的策略响应引导修订至原始和反事实标准,同时保留原始提示初始评分标准在每个提示事实下解释的意义。我们还调整阶段边界的标准权重,以更好地处理观察到的策略错误。在多个骨干链上,MetaRubric 比静态评判 GRPO 提高了 PubMedQA 准确性 6.00 至 20.40 个百分点,并在 HealthBench-Hard 和两个多模态医疗基准测试中进一步提升。
All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR
只工作不玩耍让杰克变得无趣:理解并防止RLVR中灾难性策略崩溃
- Authors: Qiyuan Huang, Tianshi Xu, Meng Li
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02835
- Pdf link: https://arxiv.org/pdf/2610.02835
- Abstract
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at this https URL.
- 中文摘要
在采用可验证奖励强化学习(RLVR)对大型语言模型(LLMs)进行后训练时,GRPO式算法可能在后期阶段出现严重崩溃。基于提示的探测表明,这并非良性的战略剪枝,而是有效战略能力的有害收缩,使得不同推理策略日益难以实现。为描述这一现象,我们通过轨迹级策略-更新交互定义策略,并构建了结合优化动力学与信息理论的统一理论框架。我们证明,主要RLVR目标逐步将概率质量集中于单一策略,而维持非平凡任务准确性则需要最低策略容量。这两者结果之间的冲突为灾难性崩溃提供了机制性解释。我们进一步推导出{镜像纠缠指数(MEI)}作为一个轻量级在线警示信号。为防止崩溃,我们提出了\textbf{Mesh Learning},它揭示了多种推理策略,防止任何单一策略主导优化。在AIME26、AIME25、MATH-500、GPQA和LiveCodeBench中,网格学习在Qwen和Phi模型家族中持续优于强基线,分别提升了最高13.4 pp和11.5 pp。这些结果确立了策略保持为稳定RLVR的关键原则。代码可在此 https URL 访问。
Understanding Enrichment in Reinforcement Learning
理解强化学习中的丰富化
- Authors: Jinwoo Kim, Shraddha Barke
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02846
- Pdf link: https://arxiv.org/pdf/2610.02846
- Abstract
When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over the length of a sample to an additive accumulation. We then apply our analysis of enrichment to interpreting the results of a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, both corrected and uncorrected, helps avoid collapse in sparse domains. Our main contribution is to understand, in general, how enrichment and correction can affect training, as opposed to claiming that either mode of operation is superior to unenriched RL.
- 中文摘要
当奖励稀疏时,带可验证奖励的强化学习(RLVR)通常通过提示或中间指导来生成更成功的推广。这种丰富偏差使策略梯度更新受到影响,除非通过重要性权重进行纠正,但现有方法省略或截断重要性权重以避免纠正方差过高。因此,修正丰富推广在RLVR中究竟获得了什么或失去了什么仍然不清楚。我们数学上展示了省略或截断的修正隐式地重新加权定义的奖励,并将由此产生的梯度误差分解为尺度、旋转和方差,以解释它们对学习的不同影响。为了使校正更实用,我们开发了一种新的序列蒙特卡洛(SMC)权重修正机制,在稳定性和粒子序假设下,能够调节样品长度上标准校正方差与加法累积的指数复合。随后,我们将富集分析应用于解释对Qwen3-1.7B在稀疏OpenMathReasoning带上的微调结果,建立了具体机制,说明富集(无论校正或未校正)如何帮助避免稀疏域的崩溃。我们的主要贡献是一般理解富集和校正如何影响训练,而不是声称任一操作模式优于未富集的强化学习。
Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning
开放团队多智能体强化学习中的交替正交学分分配
- Authors: Amit Thakur, Mukesh Singhal
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02847
- Pdf link: https://arxiv.org/pdf/2610.02847
- Abstract
Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action--turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered decomposition whose event-conditioned baseline removes the pure turnover component while preserving credit for actions that make the team robust to future replacements. We instantiate this idea with a permutation-invariant centralized critic over variable-size agent sets and event tokens, and derive both a counterfactual per-agent credit signal and a softly weighted interaction variant, TOCA-$\beta$, for high-variance control environments. Controlled diagnostic experiments show that TOCA improves return over event-aware MAPPO-style critics and that removing interaction credit substantially hurts performance. In a replacement-only Dynamic Spread benchmark, TOCA-$\beta$ achieves the best mean return at high turnover rates and improves over its no-interaction ablation. These results suggest that explicitly separating turnover from action credit is a useful principle for robust learning in dynamic cooperative teams.
- 中文摘要
开放团队多代理强化学习研究协作系统,其中代理在发作期间可能加入、离开或被替换。在此类环境中,团队返回变化既是因为代理选择了有用的行动,也因为活跃群体本身发生了变化。标准的集中批评者和共享优势通常将这两种效应混合成一个标量信用信号,允许幸存的代理因外部人员流动事件而获得奖励或惩罚。我们引入了流动正交信用分配(TOCA),这是一种面向开放团队的价值分解,区分行动效应、纯流动效应和行动--流动交互。在外部流动下,事件条件值允许一个中心分解,其事件条件基线去除了纯流动成分,同时保留了使团队对未来替代更具强韧性的行为的功劳。我们通过对可变规模代理集和事件代币的置换不变中心批判者实现这一理念,并推导出反事实的每个代理信用信号和一个软加权交互变体TOCA-$\beta$,适用于高方差对照环境。受控诊断实验表明,TOCA在事件感知型MAPPO式批评者中提高了回报率,且去除交互认可显著降低了性能。在仅替换的动态利差基准测试中,TOCA-$\beta$在高周转率下实现最佳平均回报,且优于其无交互消融。这些结果表明,明确区分周转与行动信用是动态合作团队中稳健学习的有用原则。
FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization
FARM:多任务无线网络优化的基础代理奖励模型
- Authors: Feiran You, Changxu Ni, Haozhe Ma, Jun Li, Hongyang Du
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.02947
- Pdf link: https://arxiv.org/pdf/2610.02947
- Abstract
Future wireless networks require learning agents to adapt across heterogeneous channel conditions, traffic patterns, quality-of-service (QoS) requirements, objectives, and operational constraints. Reusing decision knowledge across such tasks is challenging because conventional multi-task and transfer reinforcement learning methods primarily share or transfer policies, coupling transferable knowledge with task-dependent action mappings. This paper proposes FARM (Fundamental Agentic Reward Model for Multi-task Wireless Network Optimization), a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation. FARM introduces an Agentic Reward Model (ARM) that learns a task-conditioned reward prior from heterogeneous source-task trajectories and provides auxiliary guidance for task-specific policy optimization. In Stage I, ARM jointly models task conditions, temporal trajectory dependencies, and objective-dependent reward structures while each source task retains its own controller. In Stage II, the learned reward prior is frozen and reused to guide the adaptation of a target-specific controller for previously unseen tasks, without transferring source-task policies. Experiments on heterogeneous multi-access edge computing (MEC) tasks show that FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, compared with 16.1% for CRA Transfer, and reaches a 46.2% gain on the moderate-OOD FAR-M case. Further analysis shows that both Mamba and Transformer trajectory encoders support Reward-Space Transfer, while Mamba provides improved robustness as longer history dependencies are introduced.
- 中文摘要
未来的无线网络要求学习代理能够适应异构信道条件、流量模式、服务质量(QoS)需求、目标和操作约束。在此类任务中重复使用决策知识具有挑战性,因为传统的多任务和转移强化学习方法主要共享或转移策略,将可转移知识与任务依赖的动作映射耦合。本文提出了FARM(多任务无线网络优化基础智能奖励模型),这是一种奖励空间转移框架,将跨任务知识的重用从策略空间转移到轨迹级决策评估。FARM引入了一种代理奖励模型(ARM),该模型从异构的源-任务轨迹中学习任务条件奖励,并为任务特定策略优化提供辅助指导。在第一阶段,ARM联合建模任务条件、时间轨迹依赖性和目标依赖奖励结构,每个源任务保留自己的控制器。第二阶段,学习的奖励先验被冻结并重用,用于指导针对此前未见任务的目标特定控制器的适配,而无需转移源任务策略。异构多接入边缘计算(MEC)任务的实验显示,FARM在未见速率-延迟目标上较单任务SAC实现了29.8%的平均后期提升,而CRA转移为16.1%,在中等OOD FAR-M情况下提升了46.2%。进一步分析显示,Mamba和Transformer轨迹编码器均支持奖励-空间转移,而Mamba随着更长历史依赖的引入,提供了更好的鲁棒性。
RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction
RASPER:临床记录的奖励对齐摘要,用于EHR结局预测
- Authors: Arya Hadizadeh Moghaddam, Mohsen Nayebi Kerdabadi, Chen Chen, Dongjie Wang, Zijun Yao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02979
- Pdf link: https://arxiv.org/pdf/2610.02979
- Abstract
Unstructured discharge notes in Electronic Health Records (EHRs) often carry signal complementary to structured medical codes, holding patient-specific evidence that standardized cohort-level codes alone cannot capture. However, this evidence in notes is frequently buried in lengthy, noisy text that is not intentionally written with any specific clinical prediction in mind. Summarization is an obvious mitigation, but generic summaries, tuned for fluency rather than the outcome, routinely omit decisive evidence while retaining plausible but uninformative detail. To this end, we propose RASPER, a Reward-Aligned Summarizer for Prediction in EHR, that optimizes note summarization directly against the downstream clinical task. RASPER employs a tunable LLM-based summarizer to extract task-relevant evidence from discharge notes and trains it via reinforcement learning from prediction feedback, using a reward derived from the downstream predictor's loss. To ground the summarizer, a longitudinal encoder converts structured codes into soft prompts that incorporate each patient's clinical context into note summarization. By rewarding the quality of the resulting multimodal prediction, RASPER encourages the summarizer to retain patient-specific evidence that complements, rather than duplicates, information captured by structured codes. RASPER consistently outperforms strong baselines on both readmission prediction and medication recommendation across MIMIC-III and MIMIC-IV.
- 中文摘要
电子健康记录(EHR)中的非结构化出院记录通常携带与结构化医疗代码互补的信号,包含了单靠标准队列级代码无法捕捉的患者特定证据。然而,这些记录中的证据常常被冗长、杂乱的文本掩盖,这些文本并非有意为任何具体临床预测而写。摘要是显而易见的缓解措施,但以流利度而非结果为重点的通用摘要,常常遗漏决定性证据,同时保留合理但缺乏信息的细节。为此,我们提出了RASPER,一种用于EHR预测的奖励对齐摘要器,它能直接针对下游临床任务优化笔记摘要。RASPER采用可调的基于LLM的摘要器,从出院记录中提取任务相关证据,并通过预测反馈的强化学习进行训练,奖励来自下游预测器的损失。为使摘要器更为基础,纵向编码器将结构化代码转换为软提示,将每位患者的临床背景纳入笔记摘要中。通过奖励所得多模态预测的质量,RASPER鼓励摘要器保留患者特有的证据,这些证据与结构化代码捕获的信息相辅相成,而非重复。RASPER在MIMIC-III和MIMIC-IV的再入院预测和用药建议方面始终优于强基线。
hacktrace: behavior-supervised detection of reward hacking during code generation
Hacktrace:代码生成过程中的行为监督奖励黑客检测
- Authors: Hao Jiang, Xin Li, Annan Wang, Yichi Zhang, Weisi Lin
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03055
- Pdf link: https://arxiv.org/pdf/2610.03055
- Abstract
A coding agent can earn a passing grade by fixing its code, or by deleting the test that exposes the bug. Detecting such reward hacking requires recognizing attempted shortcuts, including those that fail. We release 173,561 annotated multi-turn coding trajectories from Qwen3-8B and show that supervising shortcut behavior independently of exploit success substantially improves detection. We introduce HACKTRACE, a behavior-supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enables monitoring before a turn is complete, without additional language-model tokens or passes. Combining this evidence with static features of the final files achieves a mean per-problem AUC of 0.997 with 8 ms of monitoring overhead, improving both accuracy and latency over monitors that run the model again on an honesty question and answer. The same generation states also provide an inexpensive monitoring signal for reinforcement learning. With strong GRPO penalties, HACKTRACE reduces the cheating share of passing solutions from 82-91% to 1-5%, while retaining honest, correct solutions and maintaining high detection accuracy as the policy evolves. Our results show that both the supervision target and the source of monitoring evidence matter for turning accurate detection into a useful training signal.
- 中文摘要
编码代理可以通过修正代码或删除暴露漏洞的测试来获得及格分数。检测此类奖励黑客需要识别尝试的捷径,包括那些失败的捷径。我们发布了173,561条来自Qwen3-8B的注释多回合编码轨迹,表明独立于漏洞成功监督捷径行为显著提升了检测能力。我们引入了HACKTRACE,一种行为监督监控器,能读取代理在生成代码时已计算的内部状态。重复使用这些状态使得在回合结束前即可进行监控,无需额外的语言模型标记或通过。将这些证据与最终文件的静态特征结合,实现了每问题的平均AUC为0.997,监控开销为8毫秒,提高了准确性和延迟,优于在诚实问答中重复运行模型的监控器。同一代状态还提供了廉价的强化学习监测信号。通过强有力的GRPO惩罚,HACKTRACE将通过解的作弊比例从82-91%降至1-5%,同时保持诚实、正确的解法,并保持高检测准确率,以应对策略的演进。我们的结果表明,监督目标和监测证据来源对于将准确检测转化为有用的训练信号至关重要。
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
HARPO:忠实且富有创造力的语言生成的幻觉感知强化学习
- Authors: Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.03063
- Pdf link: https://arxiv.org/pdf/2610.03063
- Abstract
Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
- 中文摘要
大型语言模型(LLM)容易产生幻觉内容,这会影响其在知识密集型任务中的可靠性。为了在不牺牲创造力的前提下应对这一挑战,我们提出了HARPO,一种旨在共同优化忠实度和创造力的强化学习框架。HARPO采用了通过可验证反馈训练的幻觉感知生成奖励模型(HA-GRM),以评估写作的忠实度和写作质量。选择激活机制(SAM)仅对HA-GRM判定为无幻觉的输出激活写作奖励,而数据课程则逐步将训练从创意写作转向以幻觉为中心的任务。在RAGTruth上,基于Qwen3-4B的HA-GRM获得了78.08%的响应级F1分数,而监督微调基线的66.37%。在Qwen2.5和Qwen3模型中,参数从1.7B到8B的实验显示,忠实生成和写作质量均有所提升。在Qwen3-4B上,HARPO将MultiHopRAG上HA-GRM判定的幻觉率从3.29%降至1.02%,同时将Arena-Hard-v2.0创意写作得分从16.95%提升至27.54%。
Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings
通过动态嵌入学习无作用时间序列的可转移策略
- Authors: Niklas Emonds, Georgia Koppe
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.03065
- Pdf link: https://arxiv.org/pdf/2610.03065
- Abstract
Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.
- 中文摘要
从无动作记录学习控制具有挑战性,因为干预效应未被观察到,策略可能利用重建动态中的错误。我们提出了一个基于层级的基于模型的强化学习框架,利用相关系统间的共享结构,从无动作记录中学习系统特定的控制策略。层级动力系统重建模型通过低维嵌入捕捉共享动态和个体变异。这些嵌入随后被重用用于参数化共享策略和价值网络,将重建动态的差异与控制差异联系起来。策略完全通过显式干预模型和加法潜在扰动的模拟训练。分段线性循环神经网络支持对受控动态的机械分析,而基于解码器的约束使干预的即时效果在观察空间中可解释,允许对一种模式进行干预,同时保护另一种模式免受直接操作。在Lorenz-63和双摆系统中,层级策略在独立训练策略上更有利于转移。在Lorenz-63上,它们的平均奖励优于重复规划,使用相同重建模型,表现与受控交互训练的方法相当,且仅嵌入推理后可推广到策略训练中不存在的系统。在神经行为记录中的应用显示,在受限神经扰动下,预测运动会被抑制。这些发现共同展示了共享动力学表征如何支持从无动作记录中生成可转移控制和机制假设。
Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning
超越单一视频:基于基准测试与主动证据寻求电子商务跨视频推理
- Authors: Jinghan Zhao, Yiman Hu, Liang Wu, Jian Xu, Bo Zheng
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03099
- Pdf link: https://arxiv.org/pdf/2610.03099
- Abstract
E-commerce videos are information-dense and frequently compared by consumers evaluating products and merchants assessing marketing strategies. However, existing multimodal models mainly focus on single-video understanding and have limited ability to compare information across videos. We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. Cross- video reasoning requires models to locate fine-grained evidence among many redundant frames and integrate visual details, speech, and on-screen text. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. To address the sparse credit assignment of reinforcement learning, we develop an offline trajectory rectification mechanism that identifies reasoning errors and missing multimodal evidence in RL-generated trajectories. The corrected trajectories provide supervised fine-tuning signals that reduce biases learned during RL. This mechanism supports a rectified bootstrapping pipeline in which initial RL exposes reasoning bottlenecks, supervised fine-tuning corrects them, and a final RL stage further improves the policy. AdSeek achieves 74.30 percent accuracy on the AdsCVR test split, outperforming its Qwen3-VL-8B-Instruct backbone by 27.90 percentage points. It also generalizes to the open- domain CrossVid benchmark, demonstrating effective active evidence gathering.
- 中文摘要
电商视频信息密集,消费者在产品评估和商家评估营销策略时经常进行比较。然而,现有多模态模型主要侧重于单视频理解,且在视频间信息比较能力有限。我们介绍了AdsCVR,这是首个电子商务跨视频推理基准,包含2483个视频和6110对问答对,涵盖六个推理维度。跨视频推理需要模型在众多冗余帧中定位细粒度证据,并整合视觉细节、语音和屏幕文本。因此,我们提出了AdSeek,一种智能框架,在多回合探索中动态选择视觉和音频工具,用主动证据采集取代静态均匀采样。为解决强化学习中稀疏的学分赋值,我们开发了一种离线轨迹纠正机制,识别强化学习生成轨迹中的推理错误和缺失的多模态证据。修正轨迹提供监督微调信号,减少强化学习中学到的偏差。该机制支持一个纠正引导流水线,初始强化学习暴露推理瓶颈,监督微调纠正,最终强化学习阶段进一步改进策略。AdSeek在AdsCVR测试分段中实现74.30%的准确率,比其Qwen3-VL-8B-Ininstruction骨干网高出27.90个百分点。它还推广到开放域CrossVid基准,展示了有效的主动证据收集。
How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning
如何在终身强化学习中寻找并重复使用持续适应的策略
- Authors: Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.03119
- Pdf link: https://arxiv.org/pdf/2610.03119
- Abstract
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.
- 中文摘要
在终身强化学习中,仅保留先前学过的策略不足以有效转移到新任务。有用的知识可能分布在多个先前策略中,且其相关性可能随着学习者经验的积累而变化。一种假说是,任务相似性可以在持续学习环境中有效利用,以发现并组合先前学过的策略。为测试此方法,自适应掩体选择与组合(AMSC)设计用于通过状态-动作-奖励样本中的非参数Wasserstein任务嵌入,估算在线体验的相似性。相似度评分的z分数归一化稀疏极大值可用于推导变量大小支持,以周期性选择和加权策略,以在学习新任务时形成先验。在CT图和MiniGrid上,AMSC实现了比评估模块组合基线更高的平均性能和前向转移,且不遗漏。Continual World的结果表明,识别相关先验知识及其层特异性组成可能需要额外的层级专属调整。消融显示,选择相关来源并确定重用强度是这些收益的核心。独立测量的成对迁移也与任务嵌入相似性呈正相关。这些结果表明,任务相似性可以作为选择和加权特定知识以终身强化学习重用的有效标准。
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
政策提炼的进展与崩溃:强化学习视角
- Authors: Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue Zhang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.03185
- Pdf link: https://arxiv.org/pdf/2610.03185
- Abstract
On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at this https URL.
- 中文摘要
策略上提炼(OPD)已成为语言模型训练后的重要方法。然而,尽管OPD能提升表现,但它也可能陷入过于冗长和重复的生成,而这些分歧结果背后的机制仍然不清楚。我们通过强化学习的视角解释这些结果:教师隐性奖励学生行为,即使是那些行为本身很少表现出来。从这个角度,我们的实验表明,OPD提升了表现,同时不扩展学生的能力。当隐性奖励模型可靠时,OPD使正确回答更容易被采样。相反,当偏好与质量不匹配时,就会发生奖励黑客:隐性奖励模型放大了过长、重复的学生推广,尽管它很少生成此类文本。基于这一诊断,我们发现在培训中掩盖不健康的反应并使用SFT初始化都能有效缓解崩溃。综合这些发现表明,OPD放大了教师隐性反馈偏好的学生行为,将关注点从教师生成的质量转移到了评估学生推广的可靠性。我们的代码可在此 https URL 获取。
Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
除非证据表明:教大语言模型调查员何时结案
- Authors: Tingzhu Bi, Ping Wang, Meng Ma
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.03190
- Pdf link: https://arxiv.org/pdf/2610.03190
- Abstract
Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.
- 中文摘要
事故、缺陷和停电调查以一个普通问答永远不会遇到的决定结束:迄今为止收集的证据是否足以结案。我们研究这一判决是为LLM调查员提供案卷证据,修正假设,然后要么以所读结论结案,要么保持开放并指出缺失内容。这种判断缺乏能力:未经训练的9B模型在97%的答案中夸大证据,前沿模型在84%的案例中仍高估91%,并结案了41个官方结论为“原因未明”的41个案件中的17个。测量也非常困难:病例来源在很大程度上预测其标签,而仅读取来源的规则在我们的测试案例中达到83.0的平衡准确率。因此,我们用三种测试来评估结案准确性:根据该规则及每个来源内报告的结案准确性;证据依赖,去除结论的依据,检查模型是否停止闭合;结论和差距质量,是模型断言和缺失的判断检查表。我们构建了Nautil、731个审计案例,涵盖航空、铁路、海事、化学安全和车辆缺陷报告及生产服务器事件,包含教师轨迹、非发行测试集和反事实证据版本。在这些轨迹上微调9B模型,使其结论遵循证据:移除理由使其结论率相较匹配对照降低26个百分点,夸大率从97%降至35%,正确且未夸大的结论从3%升至43%。仅奖励关闭决策的强化学习,将平衡准确率从69.2提升到83.3,与教师持平,同时将源内准确率从60.4提升到74.1,但会有所代价依赖证据。
Toward SLM-based agentic task-tool intent matching
迈向基于SLM的代理任务工具意图匹配
- Authors: Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Hervé Muyal, Marcelo Yannuzzi
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03213
- Pdf link: https://arxiv.org/pdf/2610.03213
- Abstract
Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
- 中文摘要
配备工具的人工智能代理利用工具调用访问数据并对外部系统进行操作。代理系统的横向增长增加了此类交互的数量,进一步推动了自动化、按呼叫进行监督的需求,这种监督能够在低延迟和/或本地运行。传统的授权方案可以确定代理是否被允许调用工具,但无法评估代理的底层认知,特别是工具选择是否代表满足任务意图的逻辑且相关的步骤。因此,允许的调用仍可能偏离任务意图:流氓代理可能偏离呼叫,或推动其他代理发出与任务意图不符的调用组合。因此,每一次调用都需要进行验证。本研究探讨了小型语言模型(SLMs)在这一目的上的适用性:SLM作为任务工具相关性分类器,独立地针对分配任务评估每个所选工具,并返回相关性信号以便后续强制执行。我们配备了一个新颖的数据集,包含跨越不同模型上下文协议(MCP)服务器的多工具任务,利用提示优化、监督式微调和通过GRPO的强化学习来优化和专门化SLM。
EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
EVOL:模拟器引导的进化专家综合,用于无部署学习路径推荐
- Authors: Geonwoo Bang, Dongho Kim, Moohong Min
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03273
- Pdf link: https://arxiv.org/pdf/2610.03273
- Abstract
Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.
- 中文摘要
用于学习路径推荐(LPR)的强化学习(RL)面临两个耦合障碍。首先,策略必须承诺一系列L概念且没有中间反馈,产生一个组合搜索空间,随着L的增长呈超指数增长,且仅在最后一步提供奖励。其次,专家学习路径将是稀疏奖励强化学习的自然解药,但它们不存在于教育数据中,因为学生日志记录的是学习者的行为,而非他们应做的事。我们通过导入机器人中基于模拟器的演示学习的配方来解决这两个障碍:知识追踪模拟器既用于通过进化搜索综合每位学习者的专家演示,也用于训练一个无部署策略,将这些演示提炼成前馈学习者。我们的框架EVOL通过一个非对称的actor-critic实现了该流水线,参与者承诺执行部署现实的盲计划,而criactor在训练期间利用特权模拟器状态。在三个数据集(ASSIST15、Junyi和EdNet;39-189个概念)和路径长度L = 5、10和20中,EVOL超过了涵盖启发式、顺序、强化学习、图增强RL和LLM增强方法的8个基线。我们还进一步比较了三种模仿策略(BC、AWR和DAPG),并表明最终性能由进化专家的质量决定,而非具体的模仿目标。
VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction
VenusRL:一个完全拆分的代理强化学习系统,具备优先调度和可扩展交互功能
- Authors: Mingjun Zhang (1), Yucheng Li (2), Menghao Zhang (2), Shuyong Zhu (1), Ping Zhang (3), Xiaohe Hu (3, 4), Jun Chen (3), Zhixin Wang (4), Xutong Wang (3), He Liu (3), Yanmin Jia (3), Shengrong Zhu (3), Peng Sun (5), Mingjie Zhang (3), Liming Liu (4), Jinlong Hou (4), Yuan Cheng (4), Yujun Zhang (1) ((1) Institute of Computing Technology, Chinese Academy of Sciences, (2) Beihang University, (3) Infrawaves, (4) Shanghai Innovation Institute, (5) Shanghai Zhifeng Co., Ltd.)
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2610.03286
- Pdf link: https://arxiv.org/pdf/2610.03286
- Abstract
Agentic Reinforcement Learning (RL) trains LLM agents through multi-turn interactions with external tool environments. Its multi-turn nature exposes two system-level bottlenecks unaddressed by existing agentic RL frameworks. First, end-to-end training throughput is constrained by the slowest trajectories to complete, yet optimizing per-GPU utilization alone scatters rollout progress across many groups, delaying the completion of enough groups to unblock the next training step. Second, tool sandboxes are statically over-provisioned by their declared memory ceilings, leaving most physical memory stranded while replicating near-identical state across sandboxes launched from the same prompt. We present VenusRL, a fully disaggregated agentic RL system that addresses both bottlenecks. VenusRL's priority-aware action scheduler uses length-prediction heuristics to identify sample groups whose completion is most likely to unblock the next training step, and pushes them ahead of others across batch admission, KV Cache residency, and cross-worker request orchestration. VenusRL's environment resource manager combines a memory-aware admission threshold with a template-keyed page-sharing pool, packing more sandboxes per node while preserving strict memory isolation via write-protected page table entry aliasing and copy-on-write. Across representative agentic RL workloads, VenusRL achieves up to 4.24x end-to-end training speedup over state-of-the-art baselines and reduces environment cost by up to 89%.
- 中文摘要
代理强化学习(RL)通过与外部工具环境的多回合交互来训练LLM代理。其多回合特性暴露了现有代理强化学习框架未解决的两个系统级瓶颈。首先,端到端的训练吞吐量受限于最慢的轨迹完成,但仅优化每GPU利用率就能分散部署进度,延迟足够组完成以解锁下一训练步骤。其次,工具沙盒因其声明的内存上限而静态过度配置,导致大部分物理内存被搁置,同时在同一提示符启动的沙盒间复制几乎相同的状态。我们介绍VenusRL,一个完全拆分的代理强化学习系统,解决了这两个瓶颈。VenusRL 的优先级感知动作调度器使用长度预测启发式方法,识别最有可能解锁下一训练步骤的样本组,并在批处理、KV 缓存驻留和跨工作者请求编排中领先其他组。VenusRL 的环境资源管理器结合了内存感知准入阈值与模板键分页池,在通过写保护页表条目别名和写时复制保持严格内存隔离的同时,每个节点加载更多沙盒。在代表性的代理强化学习工作负载中,VenusRL 实现了端到端训练速度高达 4.24 倍,且环境成本降低高达 89%。
Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning
双向Voronoi偏向的强化学习探索课程
- Authors: Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.03395
- Pdf link: https://arxiv.org/pdf/2610.03395
- Abstract
Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.
- 中文摘要
长期任务且奖励稀疏,是目标条件强化学习的探索瓶颈:从初始状态开始的策略很少达到目标,且没有学习信号。参考动作、手工设计的课程和塑造型奖励提供该信号,但需要演示或任务特定工程;自动起始状态和目标课程避免了这种情况,但通常只从一侧扩展,因此必须从该侧覆盖到目标的全部距离。我们提出了双向Voronoi偏向探索强化学习(BVER)课程,该课程从两端同时扩展。受双向RRT规划启发,BVER从目标向外扩展起始状态,从初始状态分布向外扩展目标,同时对未探索任务空间进行偏向,并引导它们相互引导,同时在两端训练一个目标条件化策略。在点质迷宫、四足箱子攀爬和机械臂环形挂杆转移中,BVER的学习速度超过所有无参考课程。在箱子攀爬中,它在0.4米箱子上达到95%的成功率,且在约65%的迭代次数中比最优秀的课程少65%,是唯一一个学会攀爬0.7米箱子的,并且形成了起始、目标和偏航的稳健策略变异。没有演示的情况下,它在0.4米箱和环形挂杆转移上接近参考课程的样本效率。消融显示,单从两端扩展都优于单向扩展。
Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning
超越熵:视频推理中的自我诊断多角色令牌优化
- Authors: Yudong Han, Yong Wang, Zaiquan Yang, Liang Lin, Chongyang Tao, Xiangxiang Chu, Liyuan Pan
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.03400
- Pdf link: https://arxiv.org/pdf/2610.03400
- Abstract
Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.
- 中文摘要
带有可验证奖励的强化学习显著推进了多模态推理,但仍受限于模糊的代币级信用分配。虽然高熵代币启发式鼓励可能性探索,但天真地将其推广到视频推理,往往会导致推理冗长,因为模型过度依赖高熵的视觉激活。依赖基于反事实的视觉代币定位来分配学分的替代方法,也倾向于过度优先考虑视觉探索,牺牲了决策性推理线索以推导答案,从而加剧了虚假视觉细微差别的干扰。此外,这些方法采用静态的反事实策略,在训练过程中未能与策略共同演化。本文介绍了DyCPO,一种共同进化框架,共同优化可靠的令牌选择和自适应反事实干预。它构建了一个多角色依赖指标,以平衡词符对比学习中的视觉探索与答案相关性挖掘,同时抑制仅探索的填充词和虚假的视觉噪声。DyCPO不依赖静态反事实先验,而是动态从模型自身成功与失败的推广中推导出反事实信号,实现自我诊断分析和优化目标与策略的共演化。复杂视频推理和通用视频理解基准测试的广泛实验显示了持续的性能提升,确立了DyCPO作为多模态强化学习的稳健令牌级学分分配范式。
OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation
OuroReward:文本到三维生成中强化学习的顺序奖励调度
- Authors: Bingyang Cui, Yujie Zhang, Yiling Xu, Yunfeng Guan
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.03423
- Pdf link: https://arxiv.org/pdf/2610.03423
- Abstract
Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneously through multiple reward aggregation, without explicitly modeling inter-dimension dependencies. This can cause imbalanced optimization and persistent interference among conflicting dimensions. To address this limitation, we propose OuroReward, an interference-aware sequential reward scheduling strategy for T23D RL. OuroReward first estimates pairwise dependencies among dimensions and constructs a cyclic optimization path that minimizes cumulative interference. By incorporating the tail-to-head dependency, the cycle captures global compatibility across the entire schedule. Then, OuroReward converts the cycle into a one-pass sequence, and starts optimization from the dimension with the lowest aggregate interference. Rather than assigning a fixed optimization budget to each dimension-wise reward, training adaptively determines when to advance to the next reward according to the remaining optimization headroom of the current one. We further introduce AdaSelect, an adaptive prompt selection strategy that identifies reliable and informative prompts aligned with the model's current capability. By focusing policy updates on these prompts, AdaSelect effectively improves training stability. Extensive experiments across different T23D models and RL algorithms demonstrate that our framework consistently improves generation quality across multiple dimensions.
- 中文摘要
文本到三维(T23D)生成的强化学习(RL)需要在多个质量维度(如语义对齐和纹理清晰度)上进行优化。现有方法通常通过多重奖励聚合同时优化这些维度,而未明确建模维度间的依赖关系。这可能导致优化不平衡和冲突维度间的持续干扰。为解决这一限制,我们提出了OuroReward,一种针对T23D强化学习的干扰感知顺序奖励调度策略。OuroReward首先估计维度间的两两依赖关系,构建一条循环优化路径以最小化累积干扰。通过引入尾到头依赖,循环捕捉了整个计划的全局兼容性。然后,OuroReward将循环转换为单次序列,并从聚合干扰最小的维度开始优化。训练不再为每个维度奖励分配固定的优化预算,而是根据当前奖励剩余的优化余量,自适应地决定何时推进到下一个奖励。我们进一步介绍了AdaSelect,一种自适应提示选择策略,识别与模型当前能力相符且具信息量的提示。通过聚焦这些提示的政策更新,AdaSelect有效提升了训练稳定性。跨不同T23D模型和强化学习算法的大量实验表明,我们的框架在多个维度上持续提升生成质量。
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
少测量,了解更多:自我监督测试时间特性获取
- Authors: Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03454
- Pdf link: https://arxiv.org/pdf/2610.03454
- Abstract
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-$k$, a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model's internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a stylized linear setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across task-agnostic and label-free acquisition baselines, ECHO-$k$ consistently improves budgeted downstream performance across diverse foundation-model backends. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.
- 中文摘要
多模态高维学习的最新进展使基础模型能够处理异构、大规模的数据。然而,在测试阶段,获取所有特征或模态的成本可能极高且常常冗余。因此,顺序选择信息型至关重要,但当下游任务或预测目标未知时,也存在挑战。为此,我们引入了ECHO-$k$,一种任务无关且自监督的模态习得原则:我们使用深度模型的内部预训练表示(例如基础模型)作为代理目标,总结跨模态信息。我们在风格化线性环境中提供理论保证,激励强化学习(RL)策略进行顺序模态选择。在任务无关且无标签的采集基线中,ECHO-$k$ 持续提升了不同基础模型后端预算下的下游性能。我们的方法提供了一条原则性地实现成本感知测试时间部署的路径,适用于任何测量昂贵或时间受限、下游任务事先未知的多模态系统。
Hierarchical Control via MPC-RL for Multi-Timescale Battery Systems
通过MPC-RL实现多时间尺度电池系统的分层控制
- Authors: Rasa Pourjam, Ehecatl Antonio del Río Chanona, Paulina Quintanilla
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.03508
- Pdf link: https://arxiv.org/pdf/2610.03508
- Abstract
Multi-timescale systems present a fundamental challenge, where fast operational decisions must coexist with long-horizon sustainability targets. In this work, we propose a new hierarchical control framework via Model Predictive Control (MPC) and Reinforcement Learning (RL) to separate decision-making on two distinct timescales. The high-level MPC optimizes long-horizon setpoints at the slow dynamic and on a fast timescale, a low-level pretrained RL agent tracks these setpoints in real time to maximize short-term objectives. RL is introduced to learn nonlinear control policies, without relying on model linearizations or requiring the heavy online computation from solving repeated optimal control problems. The framework is applied to a Battery Energy Storage System (BESS) operating in frequency regulation markets to balance fast profit opportunities (seconds) and slow battery degradation (weeks to months). The design employs a degradation-aware RL agent trained offline to generate safe long-horizon setpoints, and a degradation-unaware agent fine-tuned from it for fast runtime setpoint tracking. Compared to MPC baselines, the proposed approach successfully extends battery lifetime by 84% and increases operational profit by 34%.
- 中文摘要
多时间尺度系统面临根本性挑战,快速的运营决策必须与长期可持续目标共存。本研究提出通过模型预测控制(MPC)和强化学习(RL)提出新的层级控制框架,以分离两个不同时间尺度的决策。高级MPC在慢动态阶段优化长视野设定点,在快速时间尺度下,低级预训练强化智能体实时跟踪这些设定点,以最大化短期目标。强化学习旨在学习非线性控制策略,无需依赖模型线性化,也不需重复解决最优控制问题时的大量在线计算。该框架应用于频率调节市场中的电池储能系统(BESS),以平衡快速盈利机会(秒数)和电池性能缓慢(数周至数月)。该设计采用一个离线训练的具性能下降感知的强化智能体生成安全的长视距设定点,以及一个基于其微调的无性能性能智能体,实现快速运行时设定点跟踪。与MPC基线相比,该方法成功延长了84%的电池寿命,并提高了34%的运营利润。
Mastering Atari 2600 Games with Discovered Options
通过发现选项掌握Atari 2600游戏
- Authors: Erik M. Lintunen, Marlos C. Machado
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.03604
- Pdf link: https://arxiv.org/pdf/2610.03604
- Abstract
Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains, depend on handcrafted or quasi-symbolic representations, or offer little improvement over learning without options. We present Wayfarer, a general, domain-agnostic, online deep RL agent that discovers options through Laplacian representation learning from high-dimensional observations and leverages them for control. We show that the resulting options simultaneously improve exploration, accelerate credit assignment, and generalise effectively to unseen settings, enabling substantially faster learning of complex policies. Wayfarer achieves state-of-the-art performance among single-stream agents on the most challenging Atari 2600 games, with the largest gains in games that require long-horizon exploration and strategic behaviour, such as Montezuma's Revenge and Private Eye.
- 中文摘要
时间抽象,通常作为选项实例化,长期以来被视为加速信用分配、促进探索和实现强化学习(RL)泛化的机制。然而,开发在大规模高维领域中有效的通用选项发现方法仍是根本挑战。现有的选项发现方法要么局限于相对简单的领域,要么依赖手工制作或准符号表示,或者相较于无选项学习几乎没有改进。我们介绍Wayfarer,一款通用的、领域无关的在线深度强化学习代理,通过拉普拉斯表述学习从高维观测中发现选项,并加以控制。我们展示了这些选项同时提升探索性、加速信用分配,并有效泛化到看不见的环境,从而实现对复杂策略的显著快速学习。Wayfarer在最具挑战性强的Atari 2600游戏中,单流代理中达到了最先进的性能,尤其是在需要长期探索和策略行为的游戏中,如《蒙特祖玛的复仇》和《私家侦探》。
UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
UniIntervene++:一种高效现实世界强化学习的自适应干预代理
- Authors: Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.03620
- Pdf link: https://arxiv.org/pdf/2610.03620
- Abstract
Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \href{this https URL}{GitHub repository}.
- 中文摘要
在线强化学习(RL)使机器人策略通过物理交互改进,但其所需帮助会随着能力的演变而变化。基于离线估计或固定决策规则的现有干预策略因此可能与当前策略不匹配。为此,我们提出了UniIntervene++,一种适应性干预代理,学习在在线强化学习中自主执行与异构辅助行为之间的控制。具体来说,UniIntervene++首先制定了演化中的强化学习策略、轨迹修正和任务结构化的代码政策作为选项,在统一的半马尔可夫决策过程中学习它们的相对价值。基于此,能力自适应干预通过无辅助执行周期性探测强化学习策略,保持控制分配对其能力演变的响应性。最后,结合经验学习使辅助行为改进强化学习策略,其演变的结果反过来重塑未来的干预决策。通过这种方式,UniIntervene++ 共同决定何时干预、如何干预以及何时恢复控制权,随着强化学习策略的改进。在五个真实操作任务中,UniIntervene++ 的平均成功率为 89.67%,比所有基线高出至少 6 个百分点,同时将人工干预降至 0.77%,相对较最佳基线减少至少 94.6%。代码可在我们的 \href{this https URL}{GitHub repository} 中获得。
NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents
NeutronGym:面向大型语言模型代理的物理级中子仪器设计
- Authors: Lijie Ding, Changwoo Do
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Instrumentation and Detectors (physics.ins-det)
- Arxiv link: https://arxiv.org/abs/2610.03631
- Pdf link: https://arxiv.org/pdf/2610.03631
- Abstract
Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.
- 中文摘要
设计科学仪器测试语言模型代理是否能执行物理而非回忆,前提是评分不可争议。我们介绍NeutronGym,据我们所知,这是首个中子仪器设计可执行环境:代理通过验证工具构建仪器,McStas光线追踪他们构建的设备,层级解析阶梯对语法、运行时间、结构和科学进行评分,无需LLM裁判。程序类提供无限实例的固定布局实例,代理必须设定其设计参数,并保留参数范围;一个策划切片McStasBench,通过记忆探针和沙盒添加16个已发布仪器任务。7个模型最多复现16个中的7个,没有一个检索引用,也没有达到改进目标。环境也进行训练。对奖励进行强化学习,使Qwen3-8B在来自隐藏设计的族中保留实例的比例从11%提升到77%(第二个种子处为69%),超过未训练的Qwen3-32B,而该配方则在另外三个门控家族中各保留一个种子。分析说明了这个增益是多少。如果没有阶梯部分加分,它将崩溃60分。仅从奖励来看,训练模型只有在交付封闭形式物理时,才达到经典优化器在模拟预算下达到的水平(77%对81%,这一差距在此规模下不存在),而前沿模型仍能解决98-99%。获得可信结果意味着四个无模型基线能解决的任务设计失败,我们会发布找到它们的探针。
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
关键处署名:终端代理依赖感知策略优化
- Authors: Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03634
- Pdf link: https://arxiv.org/pdf/2610.03634
- Abstract
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
- 中文摘要
终端智能体在编码、调试及其他多步终端任务中受益于强化学习(RL)。在这些任务中,后续命令通常依赖于前命令产生的信息或中间结果。然而,现有的轨迹级和步级功劳分配方法并未显式追踪命令影响最终结果的读写依赖关系。因此,训练信号仍可能被分配给无关操作,削弱了从相关步骤的学习。本文提出依赖感知组策略优化(DepGPO),利用命令间的执行依赖来指导终端代理的功劳分配。具体来说,我们从执行轨迹构建命令依赖图,并从任务验证器检查的资源回溯。然后,我们将相关写入及其支持读段归功于这些路径,并用以在各步间重新分配轨迹优势。大量比较实验和消融研究表明,DepGPO能提升复杂终端任务的任务表现和训练稳定性。
On the Convergence of Success Conditioning for Policy Optimization
关于成功条件与政策优化的趋同
- Authors: Matthew Brun, Xu Andy Sun
- Subjects: Subjects:
Machine Learning (cs.LG); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2610.03642
- Pdf link: https://arxiv.org/pdf/2610.03642
- Abstract
Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, such a policy is obtained within $\mathcal{O}(\log(1/\varepsilon))$ iterations.
- 中文摘要
成功条件反射是一种在随机环境中改进决策策略的策略;它通过提高采取产生成功结果的行动概率来更新策略。成功条件反射在许多强化学习应用中很常见,但其限制行为和收敛率尚未被充分理解。本研究中,我们证明成功条件作用在一类广泛的马尔可夫决策过程(MDP)上收敛到最优策略。我们还推导了一些常见场景下的收敛率。对于贴现MDP,我们证明在$\mathcal{O}(1/\varepsilon^p)$迭代内收敛到$\varepsilon$最优策略,其中指数$p$依赖于问题数据。对于单周期MDP,这种策略可在$\mathcal{O}(\log(1/\varepsilon))$迭代内获得。
Planning to Learn
计划学习
- Authors: Ian Osband
- Subjects: Subjects:
Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.03667
- Pdf link: https://arxiv.org/pdf/2610.03667
- Abstract
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.
- 中文摘要
策略梯度方法是现代强化学习的核心,包括大型语言模型训练后学习。当它们难以使用时,常见的例子是探索、信用分配和动作采样噪声。分类法则没有这些。分类器是一种策略,其期望奖励 \emph{期望准确性},是赋予正确标签的概率,因为该标签已知,策略梯度是精确且平滑的。然而,精确策略梯度在交叉熵下会失去,即使是期望准确率。精确梯度是目光短浅的:它只根据当前购买的更新来评估一次更新,但每次更新也设定了下一次更新的起点,因此更新的价值取决于剩余学习量。从这个角度看,交叉熵是患者准确性,即一个例子如果其对数赔率以单位速度永远上升所承担的总误差,而精确策略梯度则是零视界极限。在剩余学习点截断该总数,得到视界损失,即一条线的变化,随着训练结束,从交叉熵向精确策略梯度移动。在简单的分配模型中,它可以被证明逃脱捕捉每个端点的陷阱。在MNIST和ImageNet上的ResNet-50、ResNet-101和ViT-S/16上,视界损失在平坦学习率下提高了交叉熵的顶1准确性,且增益随标签噪声增长。
Keyword: diffusion policy
CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization
CriticHack:在机器人政策优化下评估视觉奖励
- Authors: Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02527
- Pdf link: https://arxiv.org/pdf/2610.02527
- Abstract
Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five training runs raise task success by 10.2 percentage points and wrong-object failures by 10.9 points on 512 evaluation seeds, whereas five runs trained on the simulator's task-completion signal raise success without amplifying wrong-object failures (difference 9.2 points, 95% CI 5.6 to 13.0). The amplification recurs from a policy previously optimized against learned rewards, under the policy's native diffusion sampler, at matched distance from the initial policy, and across constrained-policy experiments with two critics and two optimizers. A tilt model explains when it occurs: under KL-regularized optimization, an outcome becomes more frequent whenever its expected reward under the initial policy exceeds the population average. Robometer separates successes from failures well overall (AUROC .81) but scores wrong-object failures slightly above successes (AUROC .37), so optimization raises both. The same model predicts the outcome shifts across 26 constrained settings (Spearman .89), including those in which task success falls, and Robometer's own published success-termination recipe inherits the error. A frozen outcome verifier redirects the same optimization toward the requested task.
- 中文摘要
学习到的视觉奖励模型越来越多地被用于优化机器人策略,但奖励模型对错误对象的执行评分可以与完成任务的执行评分一样高。我们证明,优化此类奖励可以放大这些错误对象的失败,而奖励和任务成功率都会上升,因此从业者通常监控的信号看起来是健康的。我们将扩散策略中的每个去噪参数与抽屉任务中的机器人计微调。从无奖励暴露的监督策略出发,五次训练运行在512个评估种子中使任务成功率提升10.2个百分点,错误对象失败提升10.9个百分点;而五次基于模拟器任务完成信号训练的运行则提升成功率,且不放大错误对象失败(差异9.2分,95% CI 5.6至13.0)。放大发生在先前针对学习奖励优化的策略中,在策略的原生扩散采样器下,在与初始策略匹配距离下,以及在两个批判者和两个优化者条件下的受限策略实验中重复出现。倾斜模型解释了其发生时间:在KL正则化优化下,当初始策略下的预期奖励超过总体平均时,结果变得更频繁。机器人计整体成功与失败区分较好(AUROC .81),但错误对象失败的得分略高于成功(AUROC .37),因此优化提升了两者。同一模型预测结果在26个受限设置中的变化(Spearman .89),包括任务成功的条件,Robometer自己发布的成功-终止配方继承了该错误。冻结的结果验证器将同一优化重定向到请求的任务。
AdaTempo: Learning Shared Relative Tempo from Demonstrations for Faster Robot Manipulation
AdaTempo:通过演示学习共享相对节奏以实现更快的机器人操作
- Authors: Jiale Cao, Yike Niu, Zhengrong Xue, Huazhe Xu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.02706
- Pdf link: https://arxiv.org/pdf/2610.02706
- Abstract
Visuomotor policies trained via imitation learning often inherit the unnecessarily slow timing of teleoperated demonstrations. Yet uniform speedup is unreliable because different phases of a manipulation task tolerate acceleration differently. In this work, we introduce AdaTempo, a self-supervised method that accelerates visuomotor policies by exploiting shared relative-tempo structure in demonstrations. AdaTempo establishes phase correspondence, aggregates the aligned relative tempo into a consensus, and maps it to a continuous speedup profile used to resample demonstrations into accelerated training trajectories. Training standard policies such as ACT or Diffusion Policy on these resampled trajectories directly embeds the desired tempo in the learned behavior, without runtime tempo selection or online retiming. Extensive evaluations show that AdaTempo achieves up to a $3.57\times$ speedup and yields a stronger success--speed trade-off than the original policies and representative acceleration baselines.
- 中文摘要
通过模仿学习训练的身体运动策略常常继承了远程操作演示不必要地缓慢的时序。然而,均匀加速不可靠,因为操作任务的不同阶段对加速的容忍度不同。本研究介绍了AdaTempo,一种自监督方法,通过利用演示中的共享相对节奏结构加速视觉运动策略。AdaTempo建立相位对应关系,将对齐的相对节奏聚合成共识,并将其映射到连续加速配置文件,用于将演示重新采样为加速训练轨迹。基于这些重采样轨迹的标准策略如ACT或扩散策略,直接将期望的速度嵌入学习行为中,无需运行时节奏选择或在线重新计时。广泛评估显示,AdaTempo实现了最高3.57美元/倍数的加速提升,并且比原始政策和代表性的加速基线更为成功——这是速度权衡。
Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation
等变视觉-触觉扩散策略用于接触丰富操作
- Authors: Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.03333
- Pdf link: https://arxiv.org/pdf/2610.03333
- Abstract
Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: this https URL
- 中文摘要
针对接触丰富操作的模仿学习需要高质量且获取成本高昂的专家数据。这使得学习样本高效策略成为关键问题。为此,我们提出了VISTA,一种工作空间级等变视觉触觉扩散策略,用于数据高效的接触丰富模仿学习。VISTA将视觉和触觉观察投射到球形标记中,通过置换等变球融合向视觉球面方向注入触觉接触线索,并利用端效器方向旋转融合谐波表示。所得表示条件为等变扩散策略以预测空间一致动作。在模拟和现实机器人环境中的大量实验表明,VISTA相比强视觉感知模拟学习基线,显著提升了数据效率。项目网站:此 https URL