生成时间: 2026-10-02 22:34:05 (UTC+8); Arxiv 发布时间: 2026-10-02 20:00 EDT (2026-10-03 08:00 UTC+8)
今天共有 76 篇相关文章
Keyword: reinforcement learning
Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
在多实例强化学习系统中整合公平性和可解释性
- Authors: Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.00035
- Pdf link: https://arxiv.org/pdf/2610.00035
- Abstract
Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of unfair predictions. This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction. MIL represents each student as a bag of weakly labeled interactions, while an RL agent selects informative instances for downstream classification. Two hypernetwork variants are evaluated to determine whether a user-defined preference scalar can continuously control the trade-off between predictive performance and Equalized Odds. The underlying RL-MIL baseline achieves strong classification performance, but both hypernetwork extensions exhibit mode collapse: changing the preference weight produces little systematic movement along the intended fairness-performance frontier. The failure is associated with objective dominance, weak gradient propagation through the conditioning mechanism, and interactions between dynamically generated parameters. The results show that fairness objectives can be incorporated into an interpretable RL-MIL pipeline, but preference conditioning alone does not guarantee controllable multi-objective behavior. Robust fair RL-MIL therefore requires explicit mechanisms for gradient balancing, objective separation, and stability analysis.
- 中文摘要
从教育互动数据预测学生表现需要既准确又足够透明的模型,以支持有意义的干预,而人口统计信息则增加了不公平预测的风险。本研究探讨了一个多目标框架,结合基于强化学习的多实例学习(RL-MIL)、对抗性偏见和偏好条件超网络,用于学生风险预测。MIL将每个学生表示为一组标记较弱的交互,而强化学习代理选择有信息的实例进行下游分类。评估了两种超网络变体,以确定用户定义的偏好标量是否能持续控制预测性能与均衡概率之间的权衡。底层的RL-MIL基线实现了强的分类性能,但两个超网络扩展都出现了模式崩溃:偏好权重的改变几乎不产生在预期公平性-绩效边界上的系统性移动。失败与目标优势、条件机制中的弱梯度传播以及动态生成参数之间的相互作用有关。结果表明,公平性目标可以被纳入可解释的RL-MIL流水线,但仅靠偏好条件并不能保证可控的多目标行为。因此,稳健的公平RL-MIL需要明确的梯度平衡、目标分离和稳定性分析机制。
Uncertainty-Aware RL-Controlled Adaptive 3D Mapping
不确定性感知的强化学习控制自适应3D映射
- Authors: Alpay Ozkan, Tunc Ozan Aydin, Marc Pollefeys, Jelena Trisovic, Daniel Barath
- Subjects: Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Image and Video Processing (eess.IV)
- Arxiv link: https://arxiv.org/abs/2610.00188
- Pdf link: https://arxiv.org/pdf/2610.00188
- Abstract
Voxel-based volumetric mapping is fundamental to 3D reconstruction, yet fixed-resolution grids remain inherently inefficient - wasting memory in uniform regions and losing detail in complex ones. Existing adaptive methods, such as MAP-ADAPT, partially address this by varying resolution based on geometry and user-defined semantic class lists, but these heuristics require expert tuning, lack generalization to unseen objects, and provide no explicit mechanism to control memory usage. We propose an adaptive framework that refines voxels based on semantic entropy, which captures label uncertainty, together with geometric curvature and texture richness as scene complexity cues, yielding principled resolution allocation without reliance on semantic taxonomies. To make the accuracy-memory trade-off explicit and user-controlled, we further introduce a reinforcement learning agent that learns voxel subdivision policies under a user-specified target memory budget, replacing hand-tuned thresholds with a single intuitive control parameter. The resulting multi-resolution TSDF achieves higher geometric accuracy, better semantic consistency, and improved memory-accuracy trade-offs compared to MAP-ADAPT and fixed-resolution baselines on both synthetic and real-world datasets. Our code and models are available at this https URL.
- 中文摘要
基于体素的体积映射是三维重建的基础,但固定分辨率网格本质上效率低下——在均匀区域浪费内存,在复杂区域失去细节。现有的自适应方法如MAP-ADAPT通过根据几何体和用户定义的语义类列表调整分辨率部分解决了这一问题,但这些启发式方法需要专业调优,缺乏对看不见对象的泛化,且不提供显式控制内存使用机制。我们提出了一种基于语义熵的自适应框架,该框架捕捉标签不确定性,结合几何曲率和纹理丰富度作为场景复杂度线索,实现原则性的分辨率分配,无需依赖语义分类法。为了使准确性与内存权衡显明且由用户控制,我们进一步引入了一个强化学习代理,在用户指定的目标内存预算下学习体素细分策略,用一个直观的控制参数替代手工调优的阈值。由此产生的多分辨率TSDF相比MAP-ADAPT和固定分辨率基线,在合成和现实世界数据集上实现了更高的几何精度、更好的语义一致性以及更好的内存准确性权衡。我们的代码和模型可在该 https 网址获取。
A Mobile Agent-Based Hierarchical Reinforcement Learning Framework for Energy-Balanced Data Collection and Wireless Charging in WSN
基于移动代理的分层强化学习框架,用于WSN中的能量平衡数据收集和无线充电
- Authors: Ali Heidaripour, Nastooh Taheri Javan
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2610.00264
- Pdf link: https://arxiv.org/pdf/2610.00264
- Abstract
Energy imbalance remains a key challenge in Wireless Sensor Networks (WSNs), as nodes near the base station deplete their energy faster due to heavy forwarding loads. While mobile agents (MAs) have been employed for either data collection or sensor charging, existing approaches lack adaptability and fail to integrate both functions under realistic hardware constraints. This paper introduces a unified mobile agent framework that performs both data collection and wireless charging sequentially under single-antenna limitations. The agent's decision-making is formulated as a two-layer Hierarchical Reinforcement Learning (HRL) problem, where the upper layer optimizes movement planning and the lower layer determines the appropriate service based on real-time network states. This hierarchical structure enables the agent to learn adaptive task scheduling policies without predefined rules. Extensive simulations demonstrate that the proposed method achieves up to 15% longer network lifetime and more balanced energy distribution compared with state-of-the-art mobile agent and deep RL approaches.
- 中文摘要
能量不平衡仍是无线传感器网络(WSN)中的关键挑战,因为基站附近的节点因高转发负载而更快耗能。虽然移动代理(MA)已被用于数据收集或传感器充电,但现有方法缺乏适应性,且在现实硬件限制下未能整合这两项功能。本文介绍了一个统一的移动代理框架,在单天线限制下顺序执行数据采集和无线充电。代理的决策过程被表述为两层分层强化学习(HRL)问题,上层优化移动规划,下层根据实时网络状态决定合适的服务。这种层级结构使智能体能够学习无需预设规则的自适应任务调度策略。大量模拟表明,所提方法相比最先进的移动代理和深度强化学习方法,网络寿命延长多达15%,能量分配更均衡。
DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
DriftOPD:用于单步VLA策略的序列级逆KL蒸馏
- Authors: Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.00317
- Pdf link: https://arxiv.org/pdf/2610.00317
- Abstract
Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.
- 中文摘要
视觉-语言-行动(VLA)模型越来越依赖于在远视界控制下生成短动作块的动作专家。虽然块级训练在机器人实例中很方便,但它优化了局部动作似然度,而不明确考虑长期任务的成功。序列级强化学习可以解决这一限制,但通常需要策略展开和闭环交互,这对真实机器人操作成本较高。我们介绍了DriftOPD,一个无需教师、无需展开的框架,用于连续VLA动作专家的序列级策略提炼。我们展示了序列级的反向Kullback-Leibler(KL)发散分解为块级反向KL项和一个未来潜在项,后者捕捉当前动作的长期视野效应。DriftOPD分别采用一步漂移目标和Q函数批评器,分别从离线演示中学到,实现序列层级优化,仅用离线数据和一步动作生成。在多种VLA架构中,模拟和实际操作中,DriftOPD通常优于现有的一步蒸馏基线,同时实现与多步教师策略相当的任务成功表现。这些结果表明,长期视野行为可以在无需在线互动或单独教师的情况下,有效地提炼为一步VLA动作专家。
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
最薄弱环节:用最坏情况的约束强化学习提炼LLM推理
- Authors: Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00332
- Pdf link: https://arxiv.org/pdf/2610.00332
- Abstract
Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest link. Motivated by this observation, we formulate reasoning distillation as a constrained reinforcement learning problem in which the task reward is maximized subject to a worst-case constraint on the teacher log-likelihood along every prefix of the trajectory. To avoid the prohibitive cost of dual Lagrangian solvers and the test-time teacher dependence of state-augmented methods such as Saute, we derive an unaugmented constrained MDP whose reward transformation preserves the hard-constraint semantics, admits a low-variance policy gradient decomposition into single-step and long-term terms, and provably satisfies the worst-case constraint almost surely in the penalty limit. Through extensive experiments on mathematical reasoning and code generation tasks, we demonstrate that our method significantly expands the accuracy-fidelity Pareto front. By matching the high Final Answer Correctness of pure RL and drastically reducing teacher constraint violations, we ultimately achieve the highest rigorous Reasoning Success Rate across all evaluated settings.
- 中文摘要
将大型语言模型(LLM)的推理能力提炼到更小的学生中,是高效部署的核心挑战。当前方法面临一个根本张力:纯粹为可验证的任务奖励(如通过GRPO)优化会导致奖励黑客行为,学生通过有缺陷的中间逻辑得出正确最终答案,而对教师施加软发散惩罚(例如基于KL的提纯)则稀释任务表现,关键是使学生在某些步骤通过高度教师共识来弥补严重的逻辑违规。我们认为这种平均与推理本质根本不符:思维链的有效性取决于其最薄弱环节。基于这一观察,我们将推理提纯提出为一种受限强化学习问题,其中任务奖励在教师对数似然的最坏情况下约束下最大化,且在轨迹的每个前缀上。为避免对偶拉格朗日求解器的高昂成本和状态增强方法如Saute对教师的测试依赖,我们推导出一个未增强的约束MDP,其奖励变换保持硬约束语义,允许低方差策略梯度分解为单步和长期项,并且几乎必然满足最坏情况约束,且在惩罚极限内几乎必然满足。通过大量数学推理和代码生成实验,我们证明了该方法显著扩展了准确性-保真度帕累托前沿。通过匹配纯强化学习的高最终答案正确性并大幅减少教师的约束违规,我们最终在所有评估环境中实现了最高的严谨推理成功率。
DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
DexPolicy:轨迹引导灵巧操作的计划探索
- Authors: Haoyu Wang, Siyuan Qian, Yanjun Li, Zeyu Zhang, Yandong Guo, Boxin Shi, Hao Tang
- Subjects: Subjects:
Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.00360
- Pdf link: https://arxiv.org/pdf/2610.00360
- Abstract
Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized PPO variant (FPO). Across five YCB objects and three training seeds, mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO). On a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials over three objects raise mean Target success from 25.0% to 85.0% (FPO), 10.0% to 63.3% (GRPO), and 8.3% to 43.3% (PPO), with one trained model per object-method condition. PPO component screening favors noise control over the tested optimizer contraction; the selected PPO schedule yields higher mean Target success than linear decay with the same endpoints on three tested objects. Training return, deterministic Target success, and tolerance to execution noise dissociate; schedules should therefore be judged by terminal task success under the intended execution conditions, per task and policy-optimization setting. Code: this https URL. Website: this https URL.
- 中文摘要
灵巧操作的强化学习(RL)必须发现手指与物体的接触,并精确控制该物体;服务于第一个目标的动作噪声可能会干扰第二个目标。在轨迹引导环境如ViViDex中,强化学习从人类视频中细化手部物体轨迹,我们的基线PPO运行在500万步后接近初始动作噪声,促使对探索尺度进行明确控制。DexPolicy将该尺度明确化为训练步骤的函数,从宽探索退火到狭窄探索,同时保持损失、架构、奖励和优化器设置固定。我们研究三种策略优化设置:PPO、无批评的GRPO延续和流参数化PPO变体(FPO)。在五个YCB对象和三个训练种子中,平均确定性目标成功率从49.4%升至68.1%(FPO),14.1%升至45.4%(GRPO),32.0%升至35.7%(PPO)。在RealMan RM75臂配Inspire/RH56手型上,360次试验对三个对象进行平均目标成功率从25.0%提升至85.0%(FPO),10.0%升至63.3%(GRPO),8.3%升至43.3%(PPO),每个对象方法条件下训练一个模型。PPO组件筛选更倾向于噪声控制而非测试优化器收缩;选定的PPO计划在三个测试对象上,在相同端点下产生更高的平均目标成功率高于线性衰减。训练回报、确定性目标成功率和执行噪声容忍度会分离;因此,调度应根据任务和策略优化设置,在预期执行条件下的终端任务成功情况来判断。代码:此 https URL。网站:此 https URL。
T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning
T2SPO:轨迹到步进策略优化,用于代理强化学习
- Authors: Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00388
- Pdf link: https://arxiv.org/pdf/2610.00388
- Abstract
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
- 中文摘要
强化学习使大型语言模型(LLM)智能体通过与环境交互学习多步行为。然而,许多交互任务中的奖励仅反映最终结果,仅提供有限的中间决策推动任务进展的指导。成功的训练轨迹包含可以为后续交互提供监督的中间状态。我们引入了轨迹到步骤策略优化(T2SPO),这是一种利用过去交互轨迹为策略学习提供步骤级反馈的方法。T2SPO从成功轨迹中推导出剩余距离目标,并将其与沿途访问状态的表示配对。基于这些例子,预训练的TabPFN回归器估计新部署每个状态下的剩余成功距离。跨连续状态的距离估计变化会为任务级监督提供代理步骤的辅助积分。随着训练的进行,新完成的轨迹刷新估计器的上下文,纳入新经验而不更新参数。在ALFWorld和WebShop上对1.5B和7B语言模型的实验表明,T2SPO持续提升整体任务成功率,优于GRPO。
MatrixReward: Reward from Rubric Matrix for Open-Ended Generation
MatrixReward:来自评分标准矩阵的开放式生成奖励
- Authors: Zihan Shen, Qi Liu, Zixuan Yang, Yiqun Chen, Chenglong Zhao, Xiaozhao Wang, Lei He
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.00389
- Pdf link: https://arxiv.org/pdf/2610.00389
- Abstract
Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.
- 中文摘要
开放式查询生成缺乏标准答案,因此需要有效的奖励机制。点分评分表只能提供有限的信息,说明同一提示下样本答案的相对质量;将多个评分标准判断合并为单一评分,也可能掩盖这些答案之间的差异。我们提出了MatrixReward,它通过比较每个评分标准下的每对抽样回答,构建出按评分标准的胜利率矩阵。每列矩阵的分布反映了该评分标准区分当前展开的强度,而列间的相关性则揭示评分标准重复性;这些统计数据共同产生依赖数据的评分标准权重。我们将这些权重与先前的评分标准权重结合起来。经过列归一化和加权后,观察到的每评分标准的最大值和最小值定义了正负的理想概况。每个推广距离这两个理想的距离决定了其相对接近性质量奖励。通过Qwen3-8B在四个开放式查询回答基准测试中评估,MatrixReward的平均得分为63.02,比最强基线高约2.0%。这些结果支持了这样一个观点:基于相对比较得出的矩阵可以用来更合理地构建开放式生成强化学习的奖励。
Whole-Body Aerial Grasping and Lifting via Partial Visual Observations
通过部分目视观察实现全身空中抓取和举起
- Authors: Jiaye Jin, Rui Jin, Xinhang Xu, Haotian Jin, Ruiyang Liu, Yi Wang, Jiayan Zhao, Kun Cao, Lihua Xie
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.00404
- Pdf link: https://arxiv.org/pdf/2610.00404
- Abstract
Aerial grasp-and-lift tasks require whole-body coordination across approach, acquisition, and lifting under partial target observations. Early approach failures can limit exposure to later task stages during training, while changing visibility complicates alignment and closure timing during execution. We present a recurrent teacher-student framework that learns a single policy in simulation to jointly command flight, arm motion, and gripper closure without an explicit task-phase input. A privileged teacher learns through reinforcement learning with a critical-state curriculum that exposes acquisition and lifting states before connecting them to normal approach trajectories. Its behavior is distilled into a recurrent visual student that replaces privileged target states with dual-view point clouds and proprioception, integrating observation history for closed-loop control. A dedicated closure objective supervises closure timing from sustained model-defined readiness sequences. Training and primary evaluation use a simulated acquisition-and-payload model with condition-triggered latching, virtual attachment, and wrench-based payload loading for short-distance lifting. Across 8,996 completed simulation episodes under this model, the frozen student achieves full-task success rates of 99.97%, 97.14%, and 95.84% under nominal, physics/control-randomized, and additional camera-randomized conditions, respectively. The nominal latch-count-weighted mean of per-seed 90th-percentile alignment errors at acquisition is 8.12 mm.
- 中文摘要
空中抓握与抬举任务需要在部分目标观察下,在进近、获取和举起之间进行全身协调。早期进场失败会限制训练中后续任务阶段的暴露,而能见度变化则使执行时的对齐和闭合时机变得复杂。我们提出了一个教师-学生的循环框架,在模拟中学习单一策略,在没有明确任务阶段输入的情况下共同指挥飞行、手臂动作和抓手闭合。特权教师通过强化学习学习,采用关键状态课程,先暴露获取和抬升状态,然后将其与正常进场轨迹连接起来。其行为被提炼为一个反复出现的视觉学生,用双视角点云和本体感觉替代特权目标状态,整合观察历史以实现闭环控制。专门的闭合目标通过持续的模型定义准备序列监督闭合时间。训练和初级评估采用模拟获取与载荷模型,结合条件触发锁定、虚拟挂钩和基于扳手的载荷加载以实现短距离举重。在该模型下完成的8996次模拟中,冻结学员在名义、物理/对照随机和额外相机随机条件下分别实现了99.97%、97.14%和95.84%的全任务成功率。获取时每种子第90百分位对准误差的名义卡扣加权平均为8.12毫米。
No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents
没有一种架构适用于所有人:跨环境层级红队代理评估
- Authors: Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, Ankit Shah
- Subjects: Subjects:
Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00557
- Pdf link: https://arxiv.org/pdf/2610.00557
- Abstract
Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates this http URL the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.
- 中文摘要
自主红队代理越来越多地通过制定战略和执行多阶段攻击来对AI驱动的网络防御进行压力测试。强化学习(RL)和大型语言模型(LLM)为这些智能体所需的规划和执行提供了互补机制,且先前工作将它们结合成混合层级结构。然而,给定架构通常在单一环境中开发和评估,因此观察到的优势是否反映了更强的决策机制,还是仅仅与特定环境的一致性。我们通过对两种同质层级红队架构的受控跨环境比较来弥补这一空白:一个带有强化执行器的RL规划器(RL+RL)和一个带有LLM执行者的LLM规划器(LLM+LLM)。我们在CybORG CAGE-4和Cyberwheel中,分别在两个网络尺度下、18种配置、统一的中断度量下,对专家自主防御者进行了评估。我们发现了明显的环境依赖性反转。RL+RL赢得了紧凑且奖励密集的CAGE-4(78.5%的中断成功率,而最强的LLM配置为18.0%)和100主机的Cyberwheel网络(81.0%对50.5%),而预训练的网络安全LLM代理则赢得了更大、升级门控的1010主机Cyberwheel网络(55.0%,RL为0.0%)。杀链分析解释了通过架构特定瓶颈来反转,汇总成功率。这个http URL是1010主机的Cyberwheel网络,RL发现并攻破主机,但在权限提升时停滞不前,而在CAGE-4中,LLM代理获得特权访问,但很少将其转化为运营影响。这些结果表明,单一环境中得出的结论可能无法泛化,混合规划者-执行者设计应以特定失败模式为动机,而非假设某一架构普遍优选。
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
让稀疏奖励发挥作用:多奖励强化学习的密度感知奖励聚合
- Authors: Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.00574
- Pdf link: https://arxiv.org/pdf/2610.00574
- Abstract
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. This reveals a residual batch-level signal imbalance and provides a basis for calibrating reward contributions. Based on this relation, we propose Density-Aware Reward Aggregation (DARA). We derive an inverse-square-root density correction that gives greater weight to signals from less frequently active rewards. DARA computes its weights from each rollout batch, adapting to changes in reward activity throughout training without modifying the underlying policy optimization objective. Experiments on tool calling and mathematical reasoning show that DARA learns the targeted behaviors faster than GDPO, reaching high format compliance in up to 26% fewer training steps on tool calling and near-saturated length compliance in up to 65% fewer steps on mathematical reasoning, while remaining competitive in final performance. Our code is available at this https URL.
- 中文摘要
多重奖励强化学习训练大型语言模型同时满足多个行为目标。GDPO中使用的奖励方向规范化保留了推广组内针对奖励的相对信息,但不同目标仍可能表现出学习进展不均。我们通过优势能量研究该行为,即奖励对一批优势的平方优势之和。在理想化GDPO归一化下,我们表明该能量与活跃组密度成正比:即在滚动组中奖励提供非零相对优势的比例。这揭示了批次级信号失衡的残余,并为校准奖励贡献提供了基础。基于该关系,我们提出了密度感知奖励聚合(DARA)。我们推导出一个反平方根密度修正,使得较低活跃奖励的信号权重增加。DARA 从每个部署批次计算权重,适应培训过程中奖励活动的变化,而不改变底层策略优化目标。工具调用和数学推理的实验显示,DARA 学习目标行为的速度快于 GDPO,工具调用训练步数减少多达 26%,数学推理中近乎饱和长度的合规,同时保持竞争力。我们的代码可在此 https 网址获取。
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
迈向利用大型语言模型实现分层网络防御:从规划到执行
- Authors: Harshith Doppalapudi, Nathaniel D. Bastian, Ankit Shah
- Subjects: Subjects:
Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00590
- Pdf link: https://arxiv.org/pdf/2610.00590
- Abstract
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.
- 中文摘要
通过强化学习(RL)训练的自主网络防御者通常绑定于其训练网络,限制了其随着网络规模变化而泛化的能力。分层式强化学习通过将战略定向与战术执行分离,降低决策复杂性,但并未消除这种再训练依赖。我们研究冻结的零样本大型语言模型(LLM)是否能在分层网络防御中提供无再训练的控制,以及随着LLM控制从规划到执行的扩展,性能如何变化。我们构建了一个控制器无关的规划者-执行者层级结构,规划者选择一个子网进行固定视野内防御,执行者选择该子网内的防御行动。利用高保真Cyberwheel环境及其内置的自动红队代理映射到MITRE ATT&CK框架,我们使用六个模型(参数范围从3B到70B)比较了RL+RL、LLM+RL和LLM+LLM配置,涵盖小型、中型和大型网络。仅用LLM替代规划器,随着网络规模增加,收益有限。相比之下,将LLM控制扩展到执行中,对于足够强大的模型会带来显著改进。例如,一个冻结的通用70B模型在三种网络尺度下,成功横向移动约为1%,攻击者影响接近零,且在相同模型权重下,而RL基线则针对每个尺度进行重新训练。我们的结果表明,足够强大的冻结大型语言模型可以在评估的网络尺度上保持强劲的防御性能,而无需针对特定任务的重新训练,同时也表明强有力的战术执行对于实现基于LLM控制的优势至关重要。
ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning
ALER:强化学习的自适应可学习体验重写
- Authors: Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00592
- Pdf link: https://arxiv.org/pdf/2610.00592
- Abstract
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least $0.82$ in all sixteen Endless T-Maze configurations and at least $0.99$ on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: this https URL.
- 中文摘要
在部分可观察强化学习(RL)中,后续观察可以使存储的信息变得过时,或改变其对下一个决策的意义。强化学习的内存架构和基准测试主要测试的是保持力,即在需要时保持信息不变的能力。我们形式化了另外两个要求。重写将决策相关内容设置为独立于旧值的值,经验融合则通过后续观察指定的规则转换旧内容。对于基于此类更新构建的任务,我们统计解决方案所需的内存状态,多个基线在需要更多状态的组合中达到最低成功率。我们引入了ALER(自适应可学习体验重写),这是一种将LSTM与槽内存配对的代理。一个独立寻址的Gumbel-Softmax写入,将权重集中在一个槽位上,覆盖该槽位,学习门则将检索的内容与策略和值头的重复状态融合。我们还引入了符文迷宫,这是三种环境,符文观察在矢量和像素观察下反转、取消、重置或重复隐藏提示的更新。在七个基线中,ALER在所有十六种无尽T-迷宫配置中至少达到0.82美元,五个符文T-迷宫组合至少达到0.99美元,并且在带有反转符文的四分支符文多走廊中平均成功率最高。在基于像素的符文MiniGrid内存中,十种配置中的八种平均成功率高于PPO-LSTM。项目页面:此链接 https URL。
Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
深入探索,更好地推理:扩散语言模型的分步风险敏感GRPO
- Authors: Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00661
- Pdf link: https://arxiv.org/pdf/2610.00661
- Abstract
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
- 中文摘要
扩散大型语言模型(dLLMs)通过去噪化序列或连续块生成文本,允许并行揭示多个代币。带可验证奖励的强化学习(RLVR)在这些决策中重复使用终端反馈,即使其条件语境发生变化。我们提出了逐步风险敏感GRPO(StepRS-GRPO),该方法在保留底层训练器的情况下,改变群优势转换在去噪状态间的风险系数。对于二元奖励,我们证明该转换恰好是对中心结果优势的瞬时和状态依赖性重新尺度。基于能力的校准建议系数尺度,而端点和插值消融则指导调度选择。在多个dLLM骨干和数学推理基准中,StepRS-GRPO提高了pass@1准确性和pass@k覆盖率,同时提高了答案多样性。在我们的消融研究中,质量匹配对照支持状态分配和调度方向的贡献,且在将优势的均方根(RMS)匹配到中心GRPO后,这些增益依然存在。推理-迹检诊断进一步表明,StepRS-GRPO带来的多样性增益超越了最终答案字符串。
Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
元多智能体强化学习,用于快速适应交互式策略,并应用于自动驾驶
- Authors: Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.00705
- Pdf link: https://arxiv.org/pdf/2610.00705
- Abstract
This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents' strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.
- 中文摘要
本文开发了一个元多智能体强化学习(meta-MARL)框架,以实现多智能体系统(MAS)中交互策略的快速适配。元强化学习(meta-RL)使智能体能够通过双级优化机制快速适应新的任务/环境。然而,现有的元强化学习通常侧重于单智能体系统。将这些框架和算法扩展到多智能体系统也面临额外挑战,因为任务不仅受环境特征,还受智能体战略交互影响。为解决这些挑战,我们将多智能体强化学习(MARL)问题建模为马尔可夫博弈(MGs),并开发了一个元MARL框架,用于在MG分布中快速交互式策略适应。定义了一个新概念,称为meta-NE,用于描述元MARL问题中的期望解概念。已建立基于梯度游玩的元MARL算法中meta-NE与静止点等价的充分条件。我们对自动驾驶任务的评估表明,所提meta-MARL方法比预训练MARL基线实现了更快的适应性,验证了我们框架的有效性。
Scalable Multi-Task Inverse Reinforcement Learning
可扩展多任务逆向强化学习
- Authors: Allen Tran, Jia Wan, Nathan Kallus, Aurélien Bibaut
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.00758
- Pdf link: https://arxiv.org/pdf/2610.00758
- Abstract
By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect agents' state occupancy. We propose a multi-task IRL method that pools data across multiple agents with different rewards in the same environment under a low-rank assumption. In addition to alleviating coverage requirements, so each task need not visit every state as long as others do, the method offers scalable evaluation of multiple tasks under new environments as computationally intensive planning scales with rank rather than the number of tasks. We provide finite sample guarantees on reward recovery and on policy learning in new environments. Experiments show our method is robust to limited coverage, recovers rewards on and off of each task's support, transfers to target environments at lower regret than baselines, with its computational advantage over per-task methods widening as tasks grow.
- 中文摘要
通过学习可转移奖励,逆强化学习(IRL)使得在修改环境下的代理能够进行反事实评估。这种转移对覆盖率有严格要求,因为目标环境会影响代理的状态占有率。我们提出了一种多任务IRL方法,在同一环境中,在低秩假设下,将数据池化于多个代理之间,奖励不同。除了减轻覆盖要求,使每个任务不必像其他任务那样访问所有状态外,该方法还提供了在新环境中对多个任务进行可扩展评估,因为计算密集型规划随排名而非任务数量增长。我们对奖励恢复和新环境中的策略学习提供了有限样本保证。实验表明,我们的方法对有限覆盖具有鲁棒性,能在每个任务支持下恢复奖励,迁移到目标环境时遗憾率低于基线,随着任务增长,其相较单任务方法的计算优势也随任务增长而扩大。
Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability
从有限时间可达性学习目标达成的准几何
- Authors: Daisuke Yamada, Travis Pence, Vikas Singh
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.00778
- Pdf link: https://arxiv.org/pdf/2610.00778
- Abstract
In goal-conditioned reinforcement learning (GCRL), quasimetric learning models goal-reaching costs as quasimetric distances, connecting local constraints to global value geometry. Its local constraints, however, should reflect the direction- dependent effects of control composition over a finite horizon together with environmental feasibility. We propose ReQRL, which constrains the critic's value gradients through finite-horizon reachability. Drawing on state-constrained optimal control, we decouple dynamical reachability from boundary geometry, estimating both from data. On OGBench, our method outperforms or rivals existing quasimetric approaches and other offline GCRL methods.
- 中文摘要
在目标条件强化学习(GCRL)中,拟度量学习将目标达成成本建模为准距离,将局部约束与全局值几何连接起来。然而,其局部约束应反映控制组合在有限视野上的方向依赖效应及环境可行性。我们提出了ReQRL,它通过有限视界可达性来约束批评者的值梯度。基于状态约束的最优控制,我们将动态可达性与边界几何解耦,两者均从数据中估算。在OGBench上,我们的方法优于或可与现有的拟度量方法及其他离线GCRL方法媲美。
Event-Triggered Practical Fixed-Time Integral Reinforcement Learning for Unknown Nonlinear Systems
未知非线性系统的事件触发实用固定时间积分强化学习
- Authors: Tien Dat Vu, My Nguyen Bach, Minh Doan
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.00800
- Pdf link: https://arxiv.org/pdf/2610.00800
- Abstract
This paper develops an event-triggered fixed-time integral reinforcement learning framework for optimal control of unknown nonlinear systems. An integral data-driven identifier is first used to reconstruct the unknown dynamics, after which an inverse-optimal formulation is employed to construct a fixed-time running cost. A learning law satisfying the practical fixed-time property is then derived. Previously collected data, or data obtained during a finite excitation interval, are stored in an experience-replay buffer and incorporated into the weight-update law. This avoids the persistent-excitation condition, which is often difficult to satisfy in practical operation. To reduce communication and control updates, an event-triggered mechanism is introduced. The paper shows that, under the event-triggered implementation, the closed-loop system still achieves practical fixed-time stability, while the proposed triggering rule guarantees the exclusion of Zeno behavior. Finally, a nonlinear example is presented to verify the theoretical results developed in the paper.
- 中文摘要
本文开发了一种事件触发固定时间积分强化学习框架,用于对未知非线性系统的最优控制。首先使用积分数据驱动标识符重建未知动态,随后采用逆最优公式构建固定时间运行成本。随后推导出满足实用固定时间性质的学习定律。先前收集的数据,或在有限激发区间内获得的数据,被存储在经验回放缓冲区中,并纳入权重更新定律。这避免了持续激发条件,而该条件在实际操作中往往难以满足。为减少通信和控制更新,引入了事件触发机制。论文表明,在事件触发实现下,闭环系统仍实现了实际的固定时间稳定性,而提出的触发规则保证排除Zeno行为。最后,提出一个非线性示例以验证论文中提出的理论结果。
SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
SHARPO:能动强化学习的分段级学分作业
- Authors: Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.00838
- Pdf link: https://arxiv.org/pdf/2610.00838
- Abstract
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
- 中文摘要
代理强化学习(RL)训练大型语言模型(LLM)以应对长时间多步交互。然而,单一局部错误可能导致任务失败,而轨迹级奖励则有限地指导个人决策获得功劳。为解决这一限制,我们引入了分段级事后诸葛优势重权重策略优化(SHARPO),这是一种在面向环境的分段层面细分层面细化群相对策略优化(GRPO)的信用分配机制。受现有策略自提纯(OPSD)方法启发,SHARPO计算每个分段内师生对数概率差距,并利用所得信号计算GRPO优势的有界乘数。该乘数在分段内所有令牌共享,允许不同分段间的信用变化。凭借Qwen2.5-7B-Instruct,SHARPO在ALFWorld和WebShop基准测试(包括GRPO、SDAR、RLSD和StepOPSD)上表现优于现有基线。
Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
学习多重时间尺度的目标条件化强化学习
- Authors: Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.00849
- Pdf link: https://arxiv.org/pdf/2610.00849
- Abstract
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
- 中文摘要
现有的离线目标条件强化学习(GCRL)方法在处理长视野任务时遇到困难。折约会使远距离状态之间的价值差异减少,直到它们低于函数近似误差,导致代理没有排序状态的信号。时间抽象将k个环境步骤视为单一转移,在长距离恢复该信号,但没有单一固定k能适用于所有状态-目标距离:大k保持长时间距离的价值差异,同时消除邻近状态间的区别,小k则相反。我们明确了这一权衡,引入广义隐式时间抽象(GITA),它对k的单一值函数进行了条件化。GITA通过聚合多个k值的优势加权监督来训练一个策略,因此赋予状态-目标对更大正面优势的尺度对其更新贡献更大。GITA无需在局部分辨率和长程信号之间做选择;它保留了两者,而无需承诺每一个k。在OGBench上,GITA的表现优于广泛的离线GCRL基线,所有任务的平均成功率比HIQL提高了25个百分点(相对提升73%)。它也比最强的固定k方法OTA提升了7个百分点(相对提升14%)。
Cross-Benchmark Transfer from RL on Agentic Coding Tasks
从强化学习在代理编码任务上的跨基准转移
- Authors: Sushant Mehta, Logan Ritchie, Edwin Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
- Arxiv link: https://arxiv.org/abs/2610.00890
- Pdf link: https://arxiv.org/pdf/2610.00890
- Abstract
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
- 中文摘要
编码代理常常在最后一公里失败:他们构建了大部分功能却放弃了需求,只测试其实现已处理的案例,破坏本应保持不变的行为,或验证未被检查的假设。我们探究对专家构建的代理编码任务进行强化学习(RL)是否弥合了这一差距,以及代理所学的内容是否能超越训练分布。我们对Kimi K2.7代码进行了后训练,这是一个1T参数(32B活跃)的开权重专家混合模型,仅RL就覆盖了1700个任务:其中1000个仓库任务通过隐藏的通过测试和通过测试对现有行为的通过测试进行评分,700个终端任务由专家编写的隐藏验证器评分。奖励是目标检查中通过的比例,如果有任何通过测试失败,则降至零。在32级LoRA适配器上进行GSPO的一个时代,在我们评估的六个外部基准测试中均提升pass@1,涵盖三个代理框架:SWE-Bench Pro(60.1至64.8)、DeepSWE(31.0至43.4)、终端-Bench 2.1(67.4至82.0)、终端-测试台3(1.4至12.1)、终端-测试台4(0.0至7.6)和SWE-Marathon(5.0至25.0)。在五个独立任务集(Terminal-Bench 4修订了Terminal-Bench 3)中,改进显著(p < 0.001),且在训练数据收集后发布的三组测试中依然显著(p = 0.004);该模型在两种未曾用于训练的约束框架下也均有提升。DeepSWE和终端-工作台3的中位数轨迹在代理步长中缩短了24-35%。基础模型失败的DeepSWE运行大多是近距离未中,训练模型新解决的任务中,配对轨迹显示其避开了上述四种失败模式。
Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning
先磨砺再适应:无数据进入状态磨砺,用于测试时强化学习
- Authors: Zhanming Zhang, Vinoth Selvendran
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.00903
- Pdf link: https://arxiv.org/pdf/2610.00903
- Abstract
Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides noisier self-supervision and may spend much of a limited adaptation budget merely concentrating probability mass before reliably expressing capability it already possesses. We propose \textbf{entry-state sharpening}: use data-free training \emph{before} TTRL to prepare a general-purpose checkpoint in a state that subsequent label-free adaptation can exploit more efficiently. The idea is not tied to one training recipe; different data-free objectives can move the same base model to different entry states. Across five data-free checkpoints derived from Qwen3-4B and evaluated under an identical 15-step TTRL protocol, entry policy entropy strongly rank-orders endpoint conversion efficiency, a reliability-to-reachability measure (Spearman $\rho=-0.90$; $\rho=-0.99$ after controlling for entry reachability). The contrast across objectives is striking: R-Zero remains diffuse at $3.39$ nats and finishes below the untuned base in 6/6 matched comparisons across MATH, GPQA, and AMC, whereas SPIRAL reaches $0.07$ nats and achieves the highest post-TTRL accuracy on MATH and GPQA despite its self-play stage using no math training data. An in-domain label-free self-distillation intervention further shows that the entry state can be deliberately sharpened. These results motivate treating checkpoint preparation as a \emph{state-control problem}: use data-free training to improve TTRL readiness, with entry entropy as a label-free control signal and reachable capability as the constraint.
- 中文摘要
测试时间强化学习(TTRL)通过基于自身样本的监督,调整无标签测试问题上的语言模型。这使得检查点的\emph{entry状态}具有重要性:扩散策略提供更响亮的自我监督,且可能花费有限的适应预算大量集中概率质量,然后可靠地表达其已有能力。我们提出\textbf{entry-state sharpening}:使用无数据训练\emph{before}TTRL准备一个通用检查点,使其后续无标签适应更高效地利用。该理念不局限于单一训练配方;不同的无数据目标可以将同一基础模型移动到不同的进入状态。在基于Qwen3-4B的五个无数据检查点中,采用相同的15步TTRL协议评估,进入策略熵强烈排序端点转换效率,这是一种可靠性到可达性的衡量指标(Spearman $\rho=-0.90$;控制进入可达性后$\rho=-0.99$)。目标间差异显著:R-Zero在MATH、GPQA和AMC的6/6匹配比较中保持为3.39美元,低于未调谐基准;而SPIRAL在自游阶段未使用数学训练数据,仍达到0.07美元NATS,且在MATH和GPQA上实现了最高的TTRL后准确率。一种域内无标签的自我蒸馏干预进一步表明,进入状态可以被有意地加深。这些结果促使将检查点准备视为\emph{状态控制问题}:利用无数据训练提升TTRL准备度,以无标签控制信号为输入熵,约束为可达能力。
eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing
eRLT:通过动作相关令牌路由实现高效的VLA强化学习
- Authors: Dehao Huang, Jianbang Liu, Jianpan Gao, Chao Tang, Zilang Cen, Zedong Dan, Jiaheng Wang, Tingguang Li, Yue Wang, Hong Zhang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.00913
- Pdf link: https://arxiv.org/pdf/2610.00913
- Abstract
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
- 中文摘要
视觉-语言-动作(VLA)模型为机器人操作提供了强的行为先验,但高效适应后游任务仍具挑战性。近期工作通过在线强化学习(RL)适应冻结的VLA来应对这一挑战,其样本效率取决于行为者和批评者所使用的状态表示质量。现有方法要么使用VLA无关的视觉编码器,要么通过固定压缩内部VLA表示来构建此类表示。这两种设计都未明确提取任务特定动作相关VLA特征,这些特征对下游动作细化和动作值估计最为有效,因此限制了样本效率。为解决这一限制,我们引入了eRLT,通过在冻结VLA的标记和层级中路由任务特定动作相关信息构建有效的状态表示。具体来说,学习到的路由标记在多个深度动态聚合视觉语言特征,而轻量级层路由器则将这些摘要合并为固定维的强化学习令牌。路由模块通过专家演示初始化捕捉预测专家行为的特征,随后利用在线互动中的批评反馈进行细化,进行动作价值估计。在七个LIBERO和RoboTwin任务中,eRLT使平均归一化学习曲线AUC比代表基线提升了多达23.7%。真实机器人实验中,USB连接器插入和主板排线插入的AUC分别提升了108.9%和46.7%,相较于最强基线。
Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
适配器丛林:分拆RLVR预算总比集中它好
- Authors: Jonathan Williams, Esin Tureci Karthik R. Narasimhan
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.00991
- Pdf link: https://arxiv.org/pdf/2610.00991
- Abstract
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
- 中文摘要
对抽样完成的多数投票是测试时间扩展的主要支柱,而带可验证奖励的强化学习(RLVR)则是使每个完成更优的主力。标准流程包括两者:用RLVR训练一个策略,然后多次采样并投票。我们证明了这种组合是有损的。投票只能推翻投票者不共享的错误,RLVR则优化策略,使其样本越来越多地犯同样错误。每种方法每个问题的完成度正好为160美元,使用完整RLVR预算训练单个LoRA适配器,提升了我们测试的每个模型的单样本准确率(1.5美元至8美元)。然而,在四个模型中有三个,多数票比未训练的基础模型低,高达4.8美元。训练过程中损害逐渐增加:投票错误相关性逐渐增强,多数投票准确率在早期达到峰值,随后下降最多7.0美元点。原因是集中度,而非RLVR本身。我们将相同的数据和训练预算分配到$K$的LoRA适配器上,每个适配器在自己随机不相交的碎片上训练,结果称为适配器丛状。丛状组在所有$16(型号,$K$)设置中投票超过完全训练的适配器,而在$K{\geq}4$时,它们与基础模型保持在$0.8点以内或更高。一个在丛林成员步数处提前停止的适配器,是一个强对照,能匹配丛林的低$K美元。对于$K{\geq}8$,丛状组保留更多RLVR单样本增益,并在八个设置中有六个投票超过该控制。集中成本也随着投票数增加而增加:从16美元到160美元,灌木丛对完全训练过的适配器的领先优势从1.3美元扩大到3.3美元。当计划是采样和投票时,RLVR预算更适合广泛使用而非深度使用。
Calibration-risk routing for controlled world-model adaptation
受控世界模型适应的校准风险路由
- Authors: Yifan Zhang, Liang Zheng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01001
- Pdf link: https://arxiv.org/pdf/2610.01001
- Abstract
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.
- 中文摘要
基于模型的强化学习(MBRL)可以利用模拟经验,但模拟器到目标的转变会带来模型选择问题:在有限的目标数据下,纠正模拟器和直接拟合目标都可能失败。我们引入模型修正世界模型(MC-WM),将初始目标数据拆分为不相交的拟合、选择和校准分区,并部署标准化校准风险较低的系列。学习到的置信信号和确定性效度基于权重一步的想象策略更新,而无需重写物理奖励。我们评估了3个受控多关节动态(MuJoCo)中540个独立报告运行单元;在部署前的工件门后重复了一个精确路由单元,共完成541次执行。
FutureWorlds: Learning Robotic World Models from Alternative Futures
未来世界:从另类未来学习机器人世界模型
- Authors: Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.01019
- Pdf link: https://arxiv.org/pdf/2610.01019
- Abstract
Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: this https URL.
- 中文摘要
机器人世界模型预测动作条件化的未来场景,为理解行动结果提供了基础。然而,将替代预测转化为有用的学习信号仍具挑战:相似的候选者限制了信息质量的比较,而分歧轨迹则需要持续维护其各自历史。我们介绍了FutureWorlds,一个统一候选构建、历史维护和相对质量学习的框架。基于多模态离散自回归模型,FutureWorlds在强化学习中使用多样化束搜索构建平衡置信度与多样性的候选未来。候选特定有界记忆保留场景状态,并确保生成和策略评分使用匹配历史。我们进一步提出了MemSPO(记忆条件搜索引导策略优化),将视频轨迹奖励转换为群体相对优势以优化世界模型。在RT-1、BridgeV2和RoboCasa上,FutureWorlds分别将32帧预测的LPIPS降低了14.78%、20.84%和9.12%,相较于每个数据集中最强基线。在固定评估配置下,仅有200次MemSPO更新进一步提升生成质量,并支持训练期后的持续预测。内存消融、解码灵敏度分析和光流评估表明,这些提升不仅限于视觉质量,还能实现更准确的运动预测和更一致的对象状态。项目页面和代码:此 https URL。
Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
带参与奖杯的探测器:随机奖励强化学习作为大型语言模型能力探测器
- Authors: Yu Mao, Lei Yu, Zining Zhu, Yusheng Zheng, Haohang Li, Freda Shi, Yutong Yin, Zhaoran Wang, Jingcheng Niu
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.01066
- Pdf link: https://arxiv.org/pdf/2610.01066
- Abstract
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.
- 中文摘要
我们将虚假奖励悖论与模型的可达性联系起来,并提出随机奖励强化学习(RL)作为探测企业的有用工具,回应了关于探究性能究竟揭示模型何物的长达十年的争论。对于即使是随机奖励也能提升大型语言模型(LLMs)性能的惊人发现,有两种主流解释:一种将收益归因于强化学习中的特定机制;另一种归因于数据污染。我们的结果引发了另一种观点:虚假奖励强化学习可以探究模型的可达性,即在特定约束下,进一步训练能从当前状态达到超出当前表现的水平。例如,在合成算术上具有相同准确率(3.5%)的两个OLMo检查点,在相同正确性奖励强化学习下,在最佳运行时分别达到8.5%和55%。在训练前和训练中期对OLMo检查点进行分析,可以发现三种不同的训练响应模式:早期即使正确答案得到奖励,RL的提升也很有限;在预训练后期,奖励正确答案变得有效,而随机奖励依然较弱;而进入训练中期后,即使是随机奖励也能带来巨大收益。类似的排序也出现在这些检查点的数字掩蔽监督微调(SFT)分析中,表明该模式并非特定强化学习机制特有的。此外,带有随机奖励的强化学习为训练能让LLM做出什么提供了独特的视角,因为其奖励信号无法提供哪些答案是正确的信息。通过询问训练在没有正确性反馈的情况下能达到什么,它解决了基于解码性探测中一个核心问题的标签泄漏问题:成功的探测是揭示了模型的能力,还是学会了任务本身。
Improving Math Reasoning through Value-guided Informative Search
通过价值引导的信息性搜索提升数学推理能力
- Authors: Shaohuai Liu, Yuning Wu, Haoran Liu, Enzo Jia, Devin Chen, Kai Wei
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01080
- Pdf link: https://arxiv.org/pdf/2610.01080
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent work introduces search into RLVR rollouts to increase trajectory diversity, but diversity alone does not ensure that the search-induced rollout policy improves upon the current policy. To address this gap, we propose APIVIS, a training-time framework that adapts finite-budget Gumbel search to chunk-level mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, allowing improvements found by search to produce informative relative rewards. It further applies selective supervision to search-improved tokens, preserving a learning signal when uniform group rewards render GRPO ineffective. We show that exact value-guided selection improves the expected verifier reward at each searched state and that this guarantee extends to the complete rollout policy, with a corresponding approximate guarantee under bounded value-estimation error. Experiments on widely recognized mathematical reasoning benchmarks and different model scales demonstrate substantial improvements over competitive search-based methods, validating the effectiveness of APIVIS.
- 中文摘要
带可验证奖励的强化学习(RLVR)显著提升了大型语言模型的数学推理能力。近期工作在RLVR推广中引入搜索以增加轨迹多样性,但仅凭多样性并不能保证搜索诱导的推广策略优于现有策略。为弥补这一空白,我们提出了APIVIS,这是一个训练时间框架,将有限预算的Gumbel搜索调整为块级数学推理。APIVIS结合了每个推广组内的直接响应和搜索响应,使搜索发现的改进产生有信息的相对奖励。它还进一步对搜索改进的代币应用选择性监督,当均匀的组奖励使GRPO无效时,仍保留学习信号。我们表明,精确价值引导选择提升了每个搜索状态下的预期验证者奖励,且这种保证延伸至完整的推广策略,并在有界值估计误差下获得相应的近似保证。在广泛认可的数学推理基准和不同模型尺度上的实验显示,APIVIS相较于竞争性基于搜索的方法有显著提升,验证了APIVIS的有效性。
MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending
MASkillBlender:通过技能融合实现多类人机车操控的去中心化全身协调
- Authors: Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang, Chengfeng Jia, Rong Su, Junjie Fu, Guanghui Wen
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2610.01102
- Pdf link: https://arxiv.org/pdf/2610.01102
- Abstract
Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
- 中文摘要
协调多人形机车操作前景看好,但由于高维全体控制、去中心化决策和可扩展性,仍具有挑战性。尽管近期强化学习方法改善了单人生物整体控制,但将其扩展到多人形环境仍不易,且通常需要大量奖励工程或任务特定设计。我们提出MASkillBlender,一种通用的多智能体强化学习框架,旨在实现分散式多人生物整体协调。通过学习共享的去中心化高层策略,通过可重复使用预训练的单人生物技能,MASkillBlender仅通过任务级奖励实现协调行为,无需任务特定动作引用。为提高学习效率,我们进一步引入了基于排列的数据增强策略,应用于齐次多人形系统,理论上表明排列样本在齐次马尔可夫博弈表述下保持原始样本的策略梯度方向。我们评估了MASkillBlender在两种类人形身体上的多类人协调任务。模拟结果表明,所提出的框架始终能实现强劲的任务性能,并实现不同任务和类人生物身体间的协调行为。
Does Scaling Reinforcement Learning Really Require More Training?
扩展强化学习真的需要更多培训吗?
- Authors: Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01133
- Pdf link: https://arxiv.org/pdf/2610.01133
- Abstract
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
- 中文摘要
缩放推理通常在强化学习(RL)或推理上花费更多计算。我们证明,完成的强化学习训练历史可以产生比优化器访问的检查点更强的策略。我们称之为策略空间缩放:扩展固定强化学习历史可访问的可部署策略集,而无需扩展训练或增加每次查询的推理计算。我们用SURGE(通过特征空间融合实现无梯度的强化化增强仪)实例化。SURGE结合了同一强化学习运行中的两个检查点:一个高精度锚点和一个产生较短响应的竞争性捐赠点。它将两个检查点表示为共享初始化的变更,然后对锚点的更新进行谱析,以保留其主导组件并纳入供体的互补组件。在固定的锚点更新保留量目标后,SURGE通过权重确定块大小,无需测试候选策略。我们评估了两个1.5亿数学推理历史DeepSeek和Nemotron,以及一个7B编码历史OLMo。SURGE在使用比锚点更少的推理标记的情况下,提高了两个输入检查点的基准平均准确率。在DeepSeek AIME24上达到54.17%,而本地测量的最大50.83%;在OLMo HumanEval+上达到83.7%,低于82.8%。这些提升超过了观察到的训练曲线。几何控制支持了强化学习更新结构的重要性,超越了仅仅权重位移或令牌减少。每个构建模型作为单一策略运行。我们的发现将存储的强化历史定位为可复用的缩放资源:训练运行所提供的能力不必止于最佳检查点。
My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning
我的错:自我诊断作为自我演进能动强化学习中的学分分配
- Authors: Yihua Zhu, Qianying Liu, Weixu Qiao, Xuan Ren, Weiwei Xu, Wenbo Li, Wei Wang, Ruijia Chen, Xinmiao Luan, Yin Luo, Hao Huang, Xiang Zheng, Hidetoshi Shimodaira
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.01161
- Pdf link: https://arxiv.org/pdf/2610.01161
- Abstract
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
- 中文摘要
代理强化学习(RL)已成为训练大型语言模型代理进行多步任务的强大方法,但依赖终端结果奖励在长期任务中产生了两个学分分配问题。首先,相同结果的推广组不提供终端奖励的学习信号。其次,终端奖励仅提供全轨迹反馈,难以识别哪些决策导致失败。近期研究补充了来自轨迹分析的更细粒度信息,如自然语言对中间决策和错误的反思。然而,自然语言诊断难以直接用于学分分配:其错误声明可能不可靠,且无法量化每个错误对学习的影响程度。我们提出了自我诊断引导的终端学分再分配(FAULT),将诊断错误转化为以终端结果为锚定的明确阶级学分。FAULT检查诊断证据,并学习任务结果中的相对错误成本。在培训过程中,策略和自我诊断者协同进化,而错误成本则在线更新,从近期结果中更新。在ALFWorld上,FAULT从相同结果组恢复学习信号,信号覆盖率达到95%,GRPO为41%,GiGPO为72%,同时更好地将特定错误步骤归功于特定步骤。在两个模型尺度上,FAULT表现强劲。在长期ALFWorld和WebShop任务上有所改进,同时在短期基于搜索的质量保证中保持竞争力。
Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
依赖感知奖励塑造用于能动强化学习
- Authors: Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01207
- Pdf link: https://arxiv.org/pdf/2610.01207
- Abstract
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at this https URL.
- 中文摘要
在用强化学习训练大型语言模型时,终端奖励几乎无法指导哪些步骤重要。分配步骤积分的常见方法忽略了这样一个事实:建立在未纠正错误基础上的工作是浪费的,而独立工作仍然有效。只有最终成功/失败奖励,失败事件中的每个步骤即使取得了进展,未来总奖励为零。我们提出了依赖感知奖励塑造(DARS),它将任务进展表示为通过前提关系连接的谓词,并在依赖图上分配步骤级的信用。注释者标记每一步的证词,用于验证、失效或修复。验证过的谓词根据距离最近破损前提的图距离进行折现,而独立谓词则不受影响。修复会根据剩余错误更新权重;失效谓词需要重新验证以恢复信用。固定电位将这些注释转换为每步有符号的奖励。通用的奖励和注释接口使DARS能够集成多种推理和代理训练方法,如GiGPO和ARPO/AEPO,而无需更改其推广策略或优化器。在1.5B至8B的五个任务家族和模型中,DARS比使用相同预算和工具训练的GiGPO提升成功率高达10分(ALFWorld),提高了WebShop任务得分和Search-R1质量保证准确率,补充了AEPO基于熵的AIME24/25训练,并用Python解释器补充了AEPO基于熵的训练,并在1.7B和4B的受控无工具推理比较中超过OmniOPD。消融显示,步级信用、依赖衰减和图拓扑均有贡献。在ALFWorld上,精简的8B注释器与API标注符匹配,使DARS能够高效运行,无需前沿评判。代码可在该HTTPS网址获取。
Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization
通过动态归导优化实现三维网格生成的流量匹配强化
- Authors: Zhen Zhou, Zhiwei Ning, Puhua Jiang, Sheng Zhang, Yifei Tang, Jie Yang, Xintong Han, Wei Liu, Chunchao Guo
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.01233
- Pdf link: https://arxiv.org/pdf/2610.01233
- Abstract
Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without explicitly specifying a target velocity field toward preferred samples. In 3D generation, constrained by pretrained model capabilities, rollout diversity, and reward-distribution complexity, directly applying these RL methods yields limited gains in geometric quality. We introduce a forward-process RL method \textbf{Dynamic Homing Optimization (DHO)}, which reformulates negative-trajectory optimization as positive-sample attraction-guided dynamic homing. Specifically, Minimum-Cost Attractive Matching (MAM) assigns each negative sample a distinct positive target, and Time-Aware Dynamic Correction (TDC) then redirects its trajectory toward the target using a remaining-time-aware corrective velocity. Building on asynchronous online DHO, we develop \textbf{Flow3D-Pro}, an image-to-3D geometry generation framework. Experiments show that DHO outperforms representative DPO-, GRPO-, and NFT-style objectives in 3D generation, while Flow3D-Pro produces higher-quality 3D geometry than existing mesh generation methods.
- 中文摘要
流匹配是三维生成的核心,但实际上其强化学习(RL)方法大多是从二维视觉生成中改编而来。当应用于负轨迹时,代表性的DPO、GRPO和NFT风格目标主要将预测速度偏离对应方向,而未明确指定目标速度场指向首选样本。在三维生成中,受预训练模型能力、滚动多样性和奖励分布复杂性限制,直接应用这些强化学习方法在几何质量上获得有限提升。我们引入了前向过程RL方法\textbf{动态归向优化(DHO)},将负轨迹优化重新表述为正样本吸引引导动态归巢。具体来说,最小成本吸引力匹配(MAM)为每个负样本分配一个不同的正目标,时间感知动态修正(TDC)则利用剩余时间感知的修正速度将其轨迹重新定向目标。基于异步在线DHO,我们开发了\textbf{Flow3D-Pro},一种图像到三维几何生成框架。实验显示,DHO在三维生成中优于代表性的DPO、GRPO和NFT风格目标,而Flow3D-Pro则能生成比现有网格生成方法更高质量的三维几何。
Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
NISQ系统上混合量子强化学习的上下文感知错误缓解编排
- Authors: Bisma Majid, Shabir Ahmed Sofi, Mir Mohammad Yousuf
- Subjects: Subjects:
Machine Learning (cs.LG); Quantum Physics (quant-ph)
- Arxiv link: https://arxiv.org/abs/2610.01253
- Pdf link: https://arxiv.org/pdf/2610.01253
- Abstract
Quantum Reinforcement Learning (QRL) integrates reinforcement learning with parameterized quantum circuits and is a promising approach to combinatorial optimization. On Noisy Intermediate-Scale Quantum (NISQ) devices, however, decoherence, gate imperfections, and measurement errors reduce policy quality and make learning less reliable. Existing error mitigation techniques are generally applied as fixed corrections that do not adapt to changing noise conditions or to the evolving state of training. This work presents Adaptive Policy-Guided Error Mitigation (APGEM) as a context-aware orchestration layer of the hybrid quantum-classical training loop that dynamically selects the most suitable mitigation strategy during QRL training. APGEM evaluates Zero-Noise Extrapolation (ZNE), Probabilistic Error Cancellation (PEC), Clifford Data Regression (CDR), and Readout Error Mitigation (REM) using policy-level indicators, including quantum-state fidelity, policy entropy, cumulative reward, and approximation ratio, and integrates the selected strategy directly into the reinforcement learning loop. The framework is evaluated on the Capacitated Vehicle Routing Problem (CVRP), a representative NP-hard problem in urban logistics, under a range of NISQ noise models and noise levels. APGEM consistently outperforms conventional static mitigation methods, reaches approximately 94% of the utility of an oracle strategy, maintains higher quantum-state fidelity as noise increases, and produces more stable learning behaviour throughout training. Ablation studies show that the framework learns context-aware mitigation policies that adapt to different noise environments and circuit execution conditions. These findings demonstrate that integrating adaptive error mitigation into the learning process substantially improves the robustness and reliability of QRL on NISQ hardware.
- 中文摘要
量子强化学习(QRL)将强化学习与参数化量子电路集成,是一种有前景的组合优化方法。然而,在噪声高的中尺度量子(NISQ)设备上,退相干、门不完美和测量误差会降低策略质量,使学习可靠性降低。现有的误差缓解技术通常作为固定修正应用,无法适应噪声条件变化或训练状态的演变。本研究将自适应策略引导错误缓解(APGEM)介绍为混合量子-经典训练循环中的上下文感知编排层,在QRL训练过程中动态选择最合适的缓解策略。APGEM利用策略级指标(包括量子态保真度、策略熵、累计奖励和近似比)评估零噪声外推(ZNE)、概率误差消除(PEC)、Clifford数据回归(CDR)和读出误差缓解(REM),并将所选策略直接整合进强化学习循环。该框架基于城市物流中代表性的NP难问题——容量车辆路由问题(CVRP)进行评估,适用于多种NISQ噪声模型和噪声水平。APGEM持续优于传统静态缓解方法,约达到oracle策略的94%效用,随着噪声增加保持更高的量子态保真度,并在整个训练过程中产生更稳定的学习行为。消融研究表明,该框架学习能够适应不同噪声环境和电路执行条件的上下文感知缓解策略。这些发现表明,将自适应错误缓解整合进学习过程,显著提升了 NISQ 硬件上 QRL 的鲁棒性和可靠性。
ColoACT: Multi-Cue Action Chunking for Smooth Autonomous Colon Navigation on a Self-Propelled Endoscopic Robot
ColoACT:多提示动作分块技术,实现自控内窥镜机器人上的平稳自主结肠导航
- Authors: Jian Hu, Shujing He, Leixin Chang, Zongze Li, Ding Huang, Chaoyang Shi, Chengzhi Hu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.01258
- Pdf link: https://arxiv.org/pdf/2610.01258
- Abstract
Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4\% and 72.5\% in straight and curved segments, respectively, and achieves 70\% success in 90-degree turns and 60\% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: this https URL.
- 中文摘要
自主结肠镜导航可以减轻操作员负担及环状结构形成或组织损伤的风险,但由于可变形的解剖结构、弱纹理和镜面内镜视觉效果,以及接触丰富的粘弹性相互作用,仍具挑战性。现有方法要么依赖几何驱动的管道,这些管道高效且可解释,但因手动工程特征和切换逻辑而脆弱;要么采用基于学习的策略,其推断深度/几何在薄弱纹理和镜面高光下可能变得时间不一致或过于平滑,而仿真训练的变体(如深度强化学习)可能还可能存在模拟到真实的差距。我们提出了ColoACT,一种自主导航系统,集成基于RGB-D-E的动作块变换器策略(ColoACT策略),用于紧凑型自走式斜齿轮内窥镜机器人(BGER)。ColoACT策略通过估计的相对深度和基于梯度的伪高程图增强RGB,以增强褶皱脊的显著性及其他高频几何线索,并通过预测重叠动作块并通过时间系联融合实现对BGER的平滑连续控制。在不同的\textit{ex-vivo}猪结肠(约60厘米)中,我们的系统在直线和曲线段分别实现了85.4%和72.5%的成功率,在90度转弯和双弯序列中成功率为60%,在具有挑战性的三弯段中进一步证明了可行性。项目页面可在以下网址访问:此 https URL。
PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
PROMO:四足机器人的偏好条件多目标强化学习
- Authors: Amr Mousa, Rifny Rachman, Neil Karavis, Michele Caprio, Richard Allmendinger
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.01260
- Pdf link: https://arxiv.org/pdf/2610.01260
- Abstract
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at this https URL.
- 中文摘要
四足行走需要平衡指令跟踪、稳定性和能效等冲突目标,而传统强化学习(RL)则将这些优先级硬编码为训练时的固定标量奖励。我们提出了PROMO(偏好条件多目标强化学习),这是一种语义多目标方法,使这一权衡成为单一移动策略的明确运行时输入。PROMO在保持具体行动先验固定的同时,对部署偏好的策略进行条件,从而将操作员意图与可行步态生成所需的奖励塑造项分离。与固定目标控制器、多目标基线和独立训练的专家相比,PROMO通过单一可部署策略实现目标专精和稳健。在模拟中100个抽样偏好中,67项行为在精确帕累托优势下不被支配,平均偏好-目标相关系数为0.843,显示出广泛的帕累托覆盖和可预测的偏好响应。同一策略将零射击转移到Unitree Go2,仅偏好变化即可使比能量降低最多30.4%,位置误差减少38.7%,峰值身体姿态偏差降低59.0%,相较于平衡偏好。这些结果确立了偏好条件多目标强化学习作为自适应腿部移动的实用运行界面,其作用超越了离线帕累托集构建。开源代码和视频可在此 https URL 获取。
IQS-BO: In-Context Query Selection for Bayesian Optimisation
IQS-BO:贝叶斯优化中的上下文查询选择
- Authors: Luca Geminiani, Nadja Klein
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.01269
- Pdf link: https://arxiv.org/pdf/2610.01269
- Abstract
Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogate and maximising an acquisition function at every evaluation step. In-context approaches based on Prior-data Fitted Networks (PFNs) amortise part of this cost by pre-training transformers on functions drawn from synthetic priors. PFNs4BO amortises the surrogate but still relies on a numerically maximised acquisition function, while FIBO performs BO fully in-context by sampling optimiser locations from a learned density, which fixes the decision rule and admits no surrogate. Learned acquisition functions score a finite candidate set with a trained network, but, lacking a label for the query, learn the score by reinforcement learning on previously solved tasks. We propose IQS-BO, a PFN that learns the query decision by supervised learning on synthetic priors. In a single forward pass, IQS-BO predicts the probability that each candidate maximises the objective over the set, and we show that the minimiser of its objective is the posterior probability of this event. The model can be pre-trained without a surrogate for fully in-context BO, or take the predictions of a fixed probabilistic surrogate as additional input, amortising only the decision step. Our method proposes queries at a fraction of the cost of acquisition-based methods, while either matching or outperforming standard BO with Gaussian processes (GPs) and available in-context methods on synthetic and real-world benchmarks. Finally, we propose a mixture prior for pre-training PFNs which combines samples from GPs with functions exhibiting warped inputs, isolated narrow optima, or plateaus that are poorly modeled by stationary kernels common in GP surrogates. We show that pre-training on this prior can lead to improved optimisation performance.
- 中文摘要
贝叶斯优化(BO)是优化昂贵黑箱函数的强大框架,但通常需要在每个评估步骤重新拟合代理并最大化获取函数。基于先验数据拟合网络(PFN)的上下文内方法通过对合成先验函数进行预训练变换器来摊销部分成本。PFNs4BO摊销代理,但仍依赖数值最大化的获取函数,而FIBO通过从学习密度中抽样优化位置完全上下文执行BO,这固定了决策规则且不允许替代。学习得的获取函数在训练网络下对有限候选集进行评分,但由于查询没有标签,通过对已解决任务的强化学习来学习得分。我们提出了IQS-BO,一种通过监督学习合成先验来学习查询决策的PFN。在一次前向传递中,IQS-BO预测每个候选对象在集合上最大化目标的概率,我们证明其目标的最小值是该事件的后验概率。该模型可以在完全上下文内BO的前提下进行预训练,无需替代,也可以将固定概率代理的预测作为额外输入,仅摊销决策步骤。我们的方法提出查询成本为基于获取方法的一小部分,同时在合成和现实基准测试中,能匹配或超越标准BO的高斯过程(GP)和上下文内方法。最后,我们提出混合预训练PFN的方案,将来自GP样本的样本与表现出扭曲输入、孤立狭最优或被GP替代中常见的平稳核模型化的平台的函数结合起来。我们证明,对该先验进行预训练可以提升优化性能。
PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading
PPO-HRAP:采用混合体制感知政策进行风险控制交易的近端政策优化
- Authors: Duong Hien Chi Kien, Thanh Trung Huynh
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Trading and Market Microstructure (q-fin.TR)
- Arxiv link: https://arxiv.org/abs/2610.01325
- Pdf link: https://arxiv.org/pdf/2610.01325
- Abstract
Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
- 中文摘要
交易中的强化学习常常难以平衡上行参与与回撤控制。仅利润政策可能崩溃,转向对向上漂移资产的被动长期曝险,而激进的风险惩罚奖励在波动期可能过于防御。本文提出PPO-HRAP混合型策略,结合了近点策略优化与可解释的先验机制。代理观察市场特征和组合状态变量,获得组合对数回报、VIX条件回撤-增加罚款、目标暴露偏差和周转成本的奖励,并执行PPO行为者输出与制度衍生目标暴露的混合操作。在2020-2022年SPY测试窗口中,PPO-HRAP实现了27.62%的总回报、8.48%的年化回报、0.6447的夏普比率、0.8588的索蒂诺比率和0.4592的卡尔玛比率,同时将买入持有的最大回撤从34.10%降至18.47%。在五个SPY种子中,PPO-HRAP保持稳定,平均总回报为0.2725美元\pm 0.0109美元,平均夏普比率为0.6219美元\pm 0.0565美元。QQQ和DIA上的单次跨资产测试进一步显示,所提方法在三项报告资产的总回报和夏普比率均排名第一。这些结果表明,将学到的操作与先前的波动率意识机制结合起来,是改善风险调整交易行为的实用方式,尽管当前政策仍导致高周转率,跨资产的稳健性仅限于单次运行证据。
Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?
在近端策略优化中重复使用过去样本:何时以及如何有效?
- Authors: Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01399
- Pdf link: https://arxiv.org/pdf/2610.01399
- Abstract
Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.
- 中文摘要
在策略内深度强化学习方法中,近端策略优化(PPO)因其在不同应用领域持续强劲的实证表现,已成为事实上的标准。然而,策略内方法本质上具有样本效率低效:当前策略下收集的新数据仅用于少量更新后即被丢弃。非策略方法通过经验重放避免了这种低效率,实现显著的样本效率提升,但代价是训练不稳定性或大量调优。这促使了混合策略的兴起,这些策略通过非策略数据重用来增强PPO。现有的PPO样本重用变体显示出相较普通PPO的样本效率提升,但系统性研究中关于何时、在哪些场景下及程度上的帮助仍存在不足。本研究中,我们通过在多重要性加权框架内实例化两个变体,研究样本重用在PPO中的有效性。两者都保留了核心PPO机制,仅重用近期迭代窗口的样本,从而将数据重用的影响与其他因素隔离开来。这两种变体称为wPPO-you和wPPO-BH,分别使用原版重要权重或平衡启发式修正权重。对两者,我们推导出策略改进的下界,为其损失提供理论基础。我们利用这些下界实证研究数据重用何时以及如何提升样本效率或PPO在连续控制任务中的最终性能。
DRL-driven RAN Slicing Management: A V2X-oriented Approach In Multi-service Scenarios
基于DRL驱动的RAN切片管理:多服务场景下的面向V2X方法
- Authors: Daniel E. Garcia-Fernandez, Pablo Vera-Soto, Sergio Fortes, M. Martinez, I. de-la-Bandera, M. L. Luque, A. Mendo, J. Ramiro, Raquel Barco
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2610.01424
- Pdf link: https://arxiv.org/pdf/2610.01424
- Abstract
The integration of Vehicle-to-Everything (V2X) communications is driving a profound transformation in vehicular connectivity, expected to significantly enhance traffic efficiency and safety. However, the stringent requirements of V2X services, particularly ultra-low latency and high reliability, present significant technical challenges. 5G's Network Slicing emerges as a key enabler by providing tailored virtual networks that ensure isolation and adaptability for heterogeneous services. This work proposes an intelligent Radio Access Network (RAN) slicing management framework specifically designed for scenarios where safety-critical V2X and high-capacity eMBB slices coexist. In such complex environments, harmonizing conflicting traffic requirements demands continuous, data-driven optimization. To achieve this, the proposed framework leverages an advanced Deep Reinforcement Learning (DRL) approach which dynamically optimizes resource allocation in real time. The framework is empirically validated on a real 5G Standalone (SA) network, where experimental results demonstrate that the DRL-driven approach successfully balances both objectives, outperforming traditional static and proportional allocation strategies by minimizing SLA violations while ensuring high resource utilization for eMBB slices.
- 中文摘要
车辆与所有(V2X)通信的整合正在推动车辆连接的深刻变革,预计将显著提升交通效率和安全。然而,V2X服务的严格要求,尤其是超低延迟和高可靠性,带来了重大技术挑战。5G的网络切片成为关键推动力,提供定制化的虚拟网络,确保异构服务的隔离性和适应性。该工作提出了一个智能无线接入网(RAN)切片管理框架,专门设计用于安全关键的V2X和高容量eMBB片共存场景。在如此复杂的环境中,协调冲突的流量需求需要持续的数据驱动优化。为此,所提框架采用先进的深度强化学习(DRL)方法,实时动态优化资源分配。该框架在真实的5G独立(SA)网络上进行了实证验证,实验结果表明,基于DRL的方法成功平衡了这两个目标,通过最小化SLA违规,同时确保eMBB切片的资源利用率高,优于传统的静态和比例分配策略。
Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
重新思考基于概率的强化学习——从后期集中
- Authors: Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01458
- Pdf link: https://arxiv.org/pdf/2610.01458
- Abstract
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
- 中文摘要
无验证者强化学习与基于概率的奖励为训练大型语言模型在缺乏外部验证者时的通用推理任务提供了有前景的方式。然而,这些奖励的可靠性,尤其是在长视野推理中,仍然缺乏充分探索。这项工作识别出一种依赖长度的概率奖励失败模式,我们称之为后验集中现象(PCP)。我们表明,基于推理追踪的参考答案的概率通常会随着追踪变长而缩小到低方差区间。这一现象导致奖励几乎无法区分,在基于GRPO的设置下,使基于概率的策略优化变得不稳定且效率低下。基于此,我们提出了带有专注感知后奖励(RLCPR)的强化学习,这是一个无验证者限制的强化学习框架,以明确考虑PCP以提升优化稳定性和令牌效率。它包含两个组成部分:不确定性感知数据抽样,减少生成前易集中的展开;以及关注注意力的正则化,后验奖励崩溃时惩罚不必要的长轨迹。大量实验表明,除了更高的代币效率外,RLCPR在七个基准测试中(包括广域和数学推理挑战)中,表现高达最先进的无验证者强化学习基线4.0%。
Sharpening Tax in Post-Training
培训后税收的提升
- Authors: Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01509
- Pdf link: https://arxiv.org/pdf/2610.01509
- Abstract
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
- 中文摘要
关于大型语言模型(LLM)训练后强化学习(RL)的一个新假说是,它仅仅是提升基础模型的现有行为,从而提高单次准确率,但代价是解决方案覆盖率的降低。尽管这种权衡在数学和编码任务中已被观察到,但这种权衡不必扩展到代理任务,因为多回合工具的使用和交互可能需要在训练后新获得的能力。我们令人惊讶的发现是,配备轻推理束的预训练LLM可以作为有能力的代理。尽管准确率(pass@1)远低于后训练LLMs,但只要有足够的测试时间预算,它们往往在解覆盖率(pass@K)上超过后续训练的对应者。我们进一步分析了其底层机制,表明后训练会将任务推向两个极端:总是解决或永远未解决,从而提高采样效率和一致性,但代价是解的覆盖率。为衡量成本,我们提出了锐化税(Sharpening Tax),这是一种诊断指标,用于量化训练后测试时间可扩展性损失。在来自四个族群和三个代理基准测试(共42个案例)的14对基础/后训练模型中,税在大多数环境中普遍存在,可通过少数几次推广估算,且与其他指标相关性良好。最后,我们提出了后置调律群抽样(PTGS),这是一种简单的即插即用贝叶斯采样器,可根据提示的估计难度调整采样温度。在两种能动环境中的强化学习训练中应用PTGS,税费低于固定温度基线,在重复抽样下解决更多任务,同时提高单次样本的准确性。
Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
决策泰坦:离线强化学习中的长期记忆测试时训练
- Authors: Jude Waide, Robert Lieck
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01513
- Pdf link: https://arxiv.org/pdf/2610.01513
- Abstract
Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
- 中文摘要
长期依赖关系仍然是人工智能领域顺序决策面临的主要挑战:RNN存在梯度消失和基于向量的隐藏状态表达有限,而基于Transformer的模型则受限于注意力的二次级缩放。近期研究提出了通过测试时间训练(TTT)框架来解决这一问题,该框架通过梯度下降在训练和测试时间的两个阶段,将情景记忆存储在神经网络的参数中。该方法在自然语言处理领域取得了成功,但据我们所知,尚未应用于强化学习(RL)领域,也尚无研究分析该记忆的实际功能。本文探讨了TTT框架通过增强TTT层的决策变换器(Decision Titan)来实现离线强化学习的潜力。我们分析了模型在X-Maze环境中的性能和属性,X-Maze是T-Maze的扩展,旨在测试顺序记忆,并研究记忆机制如何通过可视化门值随时间进行学习。我们的关键发现是,Decision Titan能够学习范围为上下文窗口20倍的长期依赖关系,推广到训练数据长度的1.7倍,但关键是时间泛化依赖于所用时间嵌入,而学习长期依赖的能力取决于相关信息的编码方式。
Towards Optimal Policy Improvement
迈向最佳政策改进
- Authors: Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin Böhmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01566
- Pdf link: https://arxiv.org/pdf/2610.01566
- Abstract
Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced MDP, characterizing planning with an explicit or implicit model as a path towards optimal policy improvement. Because practical methods commonly solve such induced problems through iterative improvement in the form of greedification, we take steps towards optimal greedification under the central practical constraint of approximate evaluation. We formulate greedification under this constraint as probabilistic decision-making under uncertainty and derive a novel operator that is optimal with respect to the resulting objective. Empirically, the operator and its practical gradient-based approximations improve aggregate performance across GumbelAlphaZero, SAC, ReBRAC and Generalized Policy Iteration, in experiments spanning discrete and continuous actions, model-based and model-free, online and offline RL.
- 中文摘要
实用强化学习(RL)算法通过迭代策略改进在近似评估存在下学习求解马尔可夫决策过程(MDP)。我们从基本原理研究策略改进,定义最优策略改进为在指定约束下,在一次更新中产生最佳策略。我们证明,限制在一组状态的最优改进等同于求解诱导MDP,通过显式或隐式模型将规划表征为通往最优策略改进的路径。由于实用方法通常通过贪婪化形式的迭代改进来解决此类诱导问题,我们在近似评估这一核心实用约束下迈向最优贪婪化。我们将该约束下的贪婪化表述为在不确定性下的概率决策,并推导出一个相对于最终目标最优的新算子。从经验上看,该算符及其实用的基于梯度的近似方法在 GumbelAlphaZero、SAC、ReBRAC 和广义策略迭代等实验中提升了汇总性能,涵盖离散与连续动作、基于模型和无模型、在线和离线的强化学习。
Continual Reinforcement Learning with Neuroevolution
神经进化持续强化学习
- Authors: Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
- Subjects: Subjects:
Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01583
- Pdf link: https://arxiv.org/pdf/2610.01583
- Abstract
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
- 中文摘要
尽管有许多关于强化学习(RL)在持续任务变化下可塑性丧失的原因和解决方法的研究,但目前尚无任何强化学习方法能够在适应与遗忘之间始终达成良好平衡。这里我们转向另一种优化范式——神经进化(NE):通过突变和选择直接在权重空间中搜索神经网络群体的算法。在各种环境和变化中,策略范围从几百个参数到百万参数网络不等,我们比较了进化策略(ES)和遗传算法(GAs)与最先进的持续强化学习变体和基于群体的强化学习。ES最稳定地实现了良好的稳定性与可塑性权衡,而GA是最可塑性的方法,但遗忘比ES更多。为此,我们研究了每种方法解的回报景观。ES找到最宽的邻域,即扰动策略仍能解决任务的权重空间区域,连续任务邻域间重叠大小与方法的稳定性-可塑性权衡相关。通过新颖性搜索奖励GA中的行为多样性,使群体更具可塑性,代价是遗忘。最后,强化学习中常见的可塑性丧失症状不会转移到NE。总体而言,这些结果确立了NE作为持续任务变化下强化学习的竞争替代方案,并表明在权重空间扰动下训练可能是持续学习更广泛有用的机制。
ReCo: Response-Consistent Locomotion with Policy-Aware MPC for Legged Manipulation
ReCo:与政策感知MPC配合响应一致的腿部操控
- Authors: Kuankuan Sima, Yichao Gao, Chenxi Gu, Kefan Zhao, Lin Zhao
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.01612
- Pdf link: https://arxiv.org/pdf/2610.01612
- Abstract
Continuous legged manipulation requires accurate end-effector tracking while the base keeps walking. Combining reinforcement learning (RL) with model predictive control (MPC) suits this task: the learned policy provides robust locomotion, while MPC coordinates the base and arm to compensate for tracking errors. However, MPC can compensate only for base motion that it can predict, and a learned policy's command response varies with gait phase, contact, and payload. We present ReCo, a framework that couples response-consistent locomotion with policy-aware MPC for legged manipulation. Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics. An identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion. On the simulation benchmark, ReCo reduces position and orientation root-mean-square error (RMSE) by 28.7% and 27.4% relative to the best baseline for each metric. Real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion.
- 中文摘要
连续的腿部操作需要在基座继续行走的同时精确地跟踪端部执行器。结合强化学习(RL)与模型预测控制(MPC)适合此任务:学习策略提供稳健的移动,MPC协调基座和臂以补偿跟踪误差。然而,MPC只能补偿其能预测的基准运动,且学习策略的指令响应随步态阶段、接触点和有效载荷变化。我们介绍ReCo框架,将响应一致性的移动与策略感知的MPC结合,用于腿部操作。响应塑造训练策略在随机动力学中一致且重复地响应指令。一个已识别的闭环响应模型允许MPC联合规划移动指令和手臂运动。在模拟基准测试中,ReCo相较于各指标的最佳基线,分别将位置和方向均方根误差(RMSE)分别降低了28.7%和27.4%。真实世界实验展示了机载连续腿部操作,基底和机械臂运动协调一致。
Iterative Policy Refinement through Semantic Rollout Analysis
通过语义推广分析进行迭代策略优化
- Authors: Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01652
- Pdf link: https://arxiv.org/pdf/2610.01652
- Abstract
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
- 中文摘要
结构化策略通过引入任务特定的归纳偏见提升模拟学习的效率、鲁棒性和可解释性,但现有结构生成方法要么依赖大量人类输入,要么依赖LLM编码的静态领域知识,这可能与专家演示不一致。我们提出了一个闭环框架,通过LLM引导的策略推广分析,迭代优化结构化策略。通过将部署记录为语义有意义的表格数据,并提示LLM生成诊断分析代码,我们的方法识别策略结构中的次优性,并在无需人工指导的情况下迭代修正。赛车和开门任务的实验表明,我们的方法在模仿学习性能上比零样本LLM生成结构提升多达15%,且实现相同强化学习性能所需的计算量减少75%。这些结果表明,表格式推广分析为LLM生成的策略结构与专家演示的对齐提供了有效的反馈信号,我们可以利用它自动生成良好的策略结构。
MiLoop: Selective Memory Propagation for Neural Combinatorial Optimization
MiLoop:选择性记忆传播用于神经组合优化
- Authors: Changliang Zhou, Yuanyao Chen, Rongsheng Chen, Zhiyun Lin, Zhenkun Wang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01685
- Pdf link: https://arxiv.org/pdf/2610.01685
- Abstract
Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.
- 中文摘要
建构性神经组合优化(NCO)已成为一种有前景的范式,它能够逐步学习构建组合优化问题(COP)的解,减少对手工规则的依赖,实现快速推断。虽然许多动态嵌入的方法具有良好的推广效果,但它们通常在每一步都通过深度注意力堆栈从零重建子问题表示。该类别中许多高效能方法依赖解标或伪标注来高效训练,或在强化学习(RL)期间通过激进的搜索空间剪枝。为解决这些局限性,我们提出了基于内存环路(Memory-in-the-Loop,MiLoop)的构造框架,利用选择性内存传播中已具备的多步计算。每次部署都为学习提供解决方案质量的反馈,同时传播历史表示,从而使浅层策略能够在不依赖外部解标签或训练时搜索空间剪枝的情况下学习有效的动态嵌入。具体来说,MiLoop将当前嵌入与注意力层之前的历史记忆融合,并在之后应用自适应门控更新。更新后的表示既支持当前决策,也支持逐步重用。跨越四个COP的广泛实验表明,MiLoop在1亿至1000万节点的实例中持续产出高质量解决方案,彰显其强大的泛化能力。
Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
架构采样:通过计算多样性在冻结视觉语言模型中进行测试时间尺度
- Authors: Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01687
- Pdf link: https://arxiv.org/pdf/2610.01687
- Abstract
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
- 中文摘要
测试时间缩放通常通过从冻结模型中抽样多个响应来寻求更好的答案,而传统温度采样则在同一固定计算路径上生成所有候选数据。我们引入了架构采样,这是一种无训练的方法,通过重复使用解码层中选定的块,通过不同的前向计算生成候选样本。改变块的位置和重复次数引入计算多样性,而无需更新模型权重或添加辅助参数。在五个Qwen检查点和十二个多模态基准测试中,架构抽样在相同九个候选方案预算下平均pass@9比标准路径温度采样提升6.58个百分点。重用早期层能带来最显著的提升,候选覆盖率的提升即使在贪婪解码下依然持续。最终候选样本在无标签测试时间强化学习的推广中,词汇重叠较低,准确性也提升。这些发现将我们架构抽样的优势扩展到候选对象覆盖之外,展示了从模型自身输出中学习的更有效性。
Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning
具有可执行验证器的函数结构化强化学习数学推理
- Authors: Zihan Liu, Xurong Xie
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01729
- Pdf link: https://arxiv.org/pdf/2610.01729
- Abstract
Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at this https URL.
- 中文摘要
算法数学推理需要可靠的分解、计算和聚合。最终答案奖励对中间错误提供有限指导,而成功执行并不保证数学正确性。本研究提出了函数结构化图强化学习(FSG-RL),将子问题图和Python实现与多验证器反馈连接起来。该策略首先通过监督微调(SFT)学习从函数图生成代码。随后使用答案门槛奖励和跨度级学分分配优化策略。该框架还支持教师监督和结构化记忆。从Primary School Math 8K(GSM8K)、MathQA、MATH和Omni-MATH中策划的基准测试将公共函数图与私有验证规范配对。在统一评估协议下,GRPO将最终答案的准确率从43.25%提升至67.50%,完全解题成功率从32.25%提升至52.25%。在教师监督下的持续强化学习(RL)带来了额外的提升。这些提升不仅仅体现在生成正确格式化的代码,还支持验证者引导的强化学习用于数学推理。代码可在此 https URL 获取。
Q-Learning for Reachability in MEC-Free MDPs
无MECMDP的Q-可达性学习
- Authors: Lu-Chin Chang, Suguman Bansal
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
- Arxiv link: https://arxiv.org/abs/2610.01781
- Pdf link: https://arxiv.org/pdf/2610.01781
- Abstract
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
- 中文摘要
可达性规范的强化学习(RL)是顺序决策的基础。先前的工作确立了渐近收敛至最优策略的方法,但仅通过必须显式估计底层马尔可夫决策过程(MDP)转移概率的基于模型的方法。我们提出了Quasar,这是首个无模型算法,对无非终端最大端分量(MEC)的MDP片段实现渐近可达性,是每个MDP依标准MEC商约简为的构建模块。我们的算法遵循经典的Q学习方法,利用时间差分更新收敛到最优策略,而无需学习转移概率。最终学习器将内存占用从O(|S|^2|A|)该算法要求O(|S||A|)。在标准化的定量验证基准集上,我们的算法以比之前基于模型的先进方案少几个数量级的样本量收敛到最优策略。这些结果共同迈向可达性学习的实际部署,以及随之而来的规范引导强化学习。
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
iADD:提升扩散政策优化中的对齐与多样性
- Authors: Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01789
- Pdf link: https://arxiv.org/pdf/2610.01789
- Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
- 中文摘要
基于强化学习的扩散模型后训练,如去噪扩散策略优化(Dezaising Diffusion Policy Optimization,DDPO),在奖励函数下优化反向扩散过程。然而,当前的奖励优化方法以牺牲多样性和质量为代价。本文通过细致的理论考量和方法设计,提供了更好的权衡。我们分析理论框架,数学证明扩散模型的更新可能对多样性有害,这与之前研究中提出的结论相反。此外,我们提出了基于坚实理论基础的增量费曼-Kac训练方法,以实现迄今为止最佳的比对与多样性权衡。我们进行了大量实验,并在三种不同任务中将方法与相关扩散策略优化方法进行比较,并为每个组成部分提供强力消融,从而验证了比对和多样性的显著性能提升。
LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
LineupRL:通过字幕到系列识别实现时间序列字幕的可验证强化学习
- Authors: Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01800
- Pdf link: https://arxiv.org/pdf/2610.01800
- Abstract
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
- 中文摘要
时间序列字幕是时间序列理解的一个基础步骤,也可以作为信号与自然语言之间的桥梁。监督式微调(SFT)依赖于更大模型的字幕,且不能超过其质量。强化学习(RL)可以,但其奖励是为其他模式和任务设计的,且在时间序列领域的开放式生成中难以转化。我们通过提出LineupRL(一种可验证奖励的强化学习(RLVR)流水线来解决这个问题,其奖励是字幕到系列识别。奖励模型是一个冻结的大语言模型(LLM)验证器,它将生成的字幕和候选时间序列作为原始值读取,而非图表,并且必须从多个干扰因素中选择描述的时间序列。匹配对验证者的需求远低于写题或评判字幕,因此现成的LLM可以提供奖励。在两个字幕基准测试以及预测和重建中,预测者仅看到字幕时,LineupRL在所有指标上都优于SFT和RL基线。由LineupRL训练的3B视觉语言模型(VLM)在参数的1/24方面也优于72B视觉语言模型,后者源自SFT基线的字幕。我们的案例研究显示,LineupRL抵抗奖励黑客攻击,且其训练的字幕员既能追踪趋势,还能命名关键点的数值。
Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
更快的协调流程:一步在线多代理流程策略
- Authors: Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01882
- Pdf link: https://arxiv.org/pdf/2610.01882
- Abstract
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
- 中文摘要
多智能体强化学习(MARL)提供了一个强大的框架,通过与环境的交互学习协调行为。开发MARL策略需要在复杂和多模态动作分布的表达建模与高效训练和执行之间取得平衡。生成策略,尤其是基于扩散的策略,能够忠实捕捉复杂和多模态行为,但昂贵的迭代采样阻碍了其在在线多智能体环境中的可扩展性。我们提出了一个通过单步流模型(OMAF)结合表达式生成策略与高效一步动作生成的在线MARL框架。OMAF采用基于Transformer的流策略捕捉复杂的协调行为,而其近似路径分数代理则提供了一条原则性的同步流策略优化路径。为实现稳定和采样高效学习,我们进一步开发了一种将软最大Q值估计与联合流策略目标耦合的联合优化方案,用于协调策略学习。通过消除迭代抽样,OMAF大幅降低了训练开销,同时不牺牲策略表达性。MPE和MAMuJoCo的10个标准任务中广泛实验表明,OMAF始终实现优越性能,与基线方法相比,回报率高达3.4倍,样本效率提升10.5倍。这些结果验证了OMAF作为在线MARL表达高效且计算高效一步流程策略范式的有效性。
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
基于选择的结构化推理:迈向高效的多模态搜索代理
- Authors: Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01892
- Pdf link: https://arxiv.org/pdf/2610.01892
- Abstract
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: this https URL.
- 中文摘要
多模态代理通常在每个动作前生成自由形式推理。对于小型模型,有限的模型容量可能导致推理冗长,对动作生成缺乏有用指导,同时产生大量推理成本。为应对这一挑战,我们引入了基于选择的结构化推理(SSR)框架,将推理重新表述为选择而非开放式生成。SSR将重复的高级推理表示为预先指定的、可重用的自然语言候选。在每个环节,模型根据当前上下文的似然从这些推理候选中选择,无需辅助任务首。使用预指定推理迹可实现并行评分,教师强制预填充通过共享上下文KV缓存同时计算候选人内及跨候选的标记似然。我们利用2B和4B模型在七个多模态搜索基准测试中评估SSR。在多个强化学习目标和监督式微调中,SSR在不牺牲任务性能的前提下实现了显著的效率提升。SSR的平均成功率可与同规模领先的搜索代理媲美,同时将每回合推理延迟降低超过90%,总每题模型推理延迟降低28-54%。项目页面:此 https 网址。
Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
异步LLM后培训:群质量上限与收敛分析
- Authors: Qijia He, Ruinan Jin, Jun Luo, Shaofeng Zou, Yingbin Liang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01896
- Pdf link: https://arxiv.org/pdf/2610.01896
- Abstract
Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(\epsilon^{-4})$ to $O(\epsilon^{-2})$ as $\epsilon\to0$, where $1+\epsilon$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
- 中文摘要
异步强化学习(RL)提高了大型语言模型训练后训练的效率,但引入了由早期策略产生的陈旧展开。关于这种陈旧如何影响收敛及其影响的理论理解仍然有限。我们推导出一个收敛界限,明确描述梯度估计量第二矩与偏置之间的权衡。对于轨迹级重要性加权估计器,我们的分析表明,一旦第二矩被均匀控制,延迟通过裁剪或重缩放引入的偏置进入界限。基于这一见解,我们提出了一种新型群质量上限GRPO(GMC-GRPO)方法,该方法在一类共享共同第二矩保证的加权估计量内最小化基于比率的偏差界限。我们为异步GMC-GRPO建立了收敛保证,并证明与TIC-GRPO相比,它将四阶延迟项的阈值依赖性从$O(\epsilon^{-4})$提升到$O(\epsilon^{-2})$,作为$\epsilon\to0$,其中$1+\epsilon$为比值阈值。在局部策略重叠下,依赖延迟项随$G减小 ^{-2/5}$ 调整步长后,其中 $G$ 为组大小。对于固定行为和当前策略,组重新缩放引入的偏置在 $G\to\infty$ 时也会消失,而轨迹层次裁剪带来的偏置可能持续存在。跨 Qwen3 模型和推理基准测试的实验显示,GMC-GRPO 在较大推展延迟下,在稳定基线中表现最佳。
Same Reward, Different Skills: When Multimodal RL Learns to Look
同样的奖励,不同的技能:当多模态强化学习学会观察时
- Authors: Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Yiqiang Chen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01908
- Pdf link: https://arxiv.org/pdf/2610.01908
- Abstract
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
- 中文摘要
带有可验证奖励的强化学习(RLVR)即使在训练中没有视觉信息也能提升视觉语言基准得分。在测试图像时,盲训练模型在3B处恢复了大约一半的真实图像增益,在7B中恢复了近五分之四。长时间的实像训练会削弱基准提升的基础性。这两个发现都暴露了同一空白:提示中的图像不是学习信号中的图像。我们的设计规则——视觉可解析性——要求视觉证据是正确答案的必要条件,且任务必须保持可学习性。我们在反事实坐标场景中测试,问题保持固定且目标从未被命名,因此正确答案需要在图像中找到目标。通过标准GRPO和正确性与格式奖励,7B模型在保持的场景中,找到目标(发现)的准确率从0.425提升到0.875,且在未训练过的问题类型上有所提升。两个控制组定位增益来源。用灰色画布替换测试图像会使发现降至零;在匹配的第30步和四个种子中训练灰色画布,即使用真实图像测试模型,基本也不会带来任何收益。所学技能可应用于独立于训练语料库构建的接地任务。回答训练问题的说明,加上相同的图像奖励和预算,可减少近三分之二的收益。改变奖励所需的内容改变了强化学习。
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
绘制RAG格局:效率、防御、交互性与推理的四轴分类法
- Authors: Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01936
- Pdf link: https://arxiv.org/pdf/2610.01936
- Abstract
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
- 中文摘要
大型语言模型(LLMs)在许多任务中展现出卓越的流利度,但仍受限于静态、参数受限的知识以及易受信息幻觉影响。检索增强生成(RAG)通过将外部检索纳入生成过程,将模型输出基于可验证且最新的来源,解决了这些问题。此前的调查主要聚焦于核心RAG架构和标准流程,而近期研究则探讨了超越这些基础设计的更广泛挑战和能力。本综述对当代RAG的发展进行了整合且结构化的分析,将该领域组织成四轴分类法:提升检索效率、强化鲁棒性和安全性、支持用户驱动和交互式工作流,以及支持多步或复杂推理。我们规范了RAG框架的关键组成部分,并回顾了涵盖密集检索和稀疏检索、融合策略、嵌入优化以及基于强化学习的检索策略的方法,强调这些进展如何影响实际部署和系统设计。我们还综合了评估实践、领域特定应用以及如朴素、高级和模块化RAG等架构变体。最后,我们概述了与检索质量、可靠性、域适应、可扩展性和可解释性相关的持续挑战,并识别构建更可靠、适应性和透明RAG系统的机会。
Do Your Own Research: Learning to Forecast by Learning to Search
自己做研究:通过学习搜索来学习预测
- Authors: Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.01955
- Pdf link: https://arxiv.org/pdf/2610.01955
- Abstract
Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
- 中文摘要
基于结果的强化学习可以训练语言模型预测现实世界事件,但先前的预测工作要么在训练前冻结研究上下文,要么只在测试时部署代理研究,因此收集证据的技能从未受奖励影响。我们引入了一个代理预测环境、数据集和基于2100+个已解决Polymarket问题的框架;代理在推出时获得自己的上下文(网页搜索、页面阅读和财务时间序列,均受分层泄露过滤限制,仅限于每个问题截止前发布的信息),我们用单一时期GRPO在Brier评分奖励下训练Qwen3.5-35B-A3B(3B活跃参数)。训练改变了代理与信息的交互方式:校准提升30-40%,每次部署搜索尝试次数从3.8降至2.25次,随着证据纪律的学习。在相同框架下,针对四个前沿模型进行评估,训练出的政策也领先于所有基于证据预测测试的前沿模型,包括Claude Opus 4.5(软布赖尔0.254对0.256,n=265),推理成本约为推断成本的5%,且在最难的问题上,即群体尚未决定的问题,其优势最为宽广。我们将环境、数据集和每次部署记录作为可重复使用的时间预测代理工具发布。
Token-Level Video Reinforcement Learning
令牌级视频强化学习
- Authors: Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.01973
- Pdf link: https://arxiv.org/pdf/2610.01973
- Abstract
Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.
- 中文摘要
视频生成的强化学习(RL)通常为整个采样视频分配一个标量奖励。然而,视频并非完全有缺陷:有些视觉代币可能已经满足提示,而另一些则需要修正。标量奖励无法定位错误,导致优化扰乱满意的代币,同时降低真正需要改变的代币。我们介绍代币级视频强化学习(Token-Level Video Reinforcement Learning,简称TVRL),这是一个通过优化奖励获得代币级信用的框架。我们的关键见解是,冻结视觉语言模型的答案似然同时提供两种信号:其输出对视频级奖励有贡献,而视频输入梯度的幅度则揭示哪些生成的视频代币对该评分影响最大。我们通过将提示衍生的问题奖励平均为一个群体相对优势,并使用分离、问题条件化的代币-积分映射,在截断策略比率内重新加权稠密去噪-转移对数概率,从而实现TVRL的群体相对策略优化。在VBench-2.0下,TVRL总体得分为57.69,比基础模型高出3.60分。TVRL还将三个SDE抽样器(SAGE、Flow和Dance)匹配GRPO基线提升2.68-3.15分,在四个奖励模型(VideoAlign、VideoScore2、UnifiedReward2和Qwen3.5-9B)提升1.33-3.15分。
Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos
贝尔曼遇见柳雅普诺夫:通过掌握混沌进行无监督强化学习
- Authors: Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02012
- Pdf link: https://arxiv.org/pdf/2610.02012
- Abstract
Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.
- 中文摘要
强化学习(RL)是训练智能体的强大范式,但其成功依赖于人类工程师的领域专长,他们为每个新任务设计信息性奖励信号。无监督强化学习旨在通过内在动机(IM)减少这种工程化:即由智能体与环境互动中产生的奖励信号。然而,现有的强化学习目标涉及信息变量的选择,这重新引入了该领域试图消除的领域专业知识。我们引入了前向CIP(F-CIP),这是一种基于RL原生的可控信息生产(CIP)目标表述,仅由系统动态定义,无需此类选择。我们证明F-CIP与强化学习兼容,并展示了其在现有算法中的有效性。用F-CIP训练代理,可以无监督地发现原始行为,如平衡和保持可控性,这些对于更复杂的机器人行为至关重要。结合简单的前进速度奖励,我们的方法产生了协调的步态,如跳跃和奔跑,这些动作需要奖励工程来学习。
On Language Drift during RLVR Post-Training
关于RLVR培训后期的语言漂移
- Authors: Michael Sullivan, Alexander Koller
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.02015
- Pdf link: https://arxiv.org/pdf/2610.02015
- Abstract
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
- 中文摘要
LLM推理模型的最新进展---主要由可验证奖励强化学习(RLVR)后训练范式驱动---使其能够完成极其复杂的任务。然而,随着能力的提升,LLM在其思维链(CoT)中也越来越多地表现出语言漂移的迹象:不寻常、非标准且看似无意义的语言使用。尽管该方法有充分文献支持---且可能影响CoT的可监控性---目前语言漂移的原因尚不充分。本文指出语言漂移发生的条件:理论上证明RLVR优化压力允许无界语言漂移,而监督微调则不允许。随后我们实证证明语言漂移特指在RLVR中新颖推理任务中产生---i. e.当目标行为无法从基础模型中提取出来时。最后,我们证明,在不限制预期奖励的情况下限制语言漂移是不可能的,这意味着CoT可监测性无法在前沿训练后RLVR期间的表现上有所提升。
Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
通过自适应Tversky策略优化实现可控多标签视频安全检测
- Authors: Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne
- Subjects: Subjects:
Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.02019
- Pdf link: https://arxiv.org/pdf/2610.02019
- Abstract
The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at this https URL .
- 中文摘要
基于视频的社交媒体快速发展增加了用户接触有害内容的频率,带来了对可靠自动视频安全检测的需求。尽管近期的视觉语言模型(VLM)展现出强大的视频理解能力,但现有有害视频检测系统面临两个关键限制:它们通常将安全检测简化为二元分类,忽视了不安全视频的多标签特性;以及依赖静态训练目标,无法支持可控的精确-召回权衡,尽管所需操作点可能因审核流程和不安全类别而异。为弥补这些不足,我们提出了自适应特沃斯基策略优化(ATPO),这是一个针对多标签视频安全检测(Multi-VSD)的强化学习框架。ATPO引入了自适应特沃斯基奖励(ATR),在训练过程中动态调整假阳性和假阴性惩罚,实现可控的精度回忆权衡。在 SafeWatch-Bench 和 XD-Violence 上的实验显示,ATPO 大幅提升了多标签性能,将 SafeWatch-Bench-Real 上的 Jaccard 指数从 40.66 提升至 75.44。此外,ATR 实现了精确回忆操作点的可靠引导,支持具有异构策略需求的部署场景。代码和检查点可在此 https URL 提供。
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
SPHERE:通过LLM增强的空间偏好学习和人机循环强化学习实现自适应VR室内场景生成
- Authors: Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
- Arxiv link: https://arxiv.org/abs/2610.02023
- Pdf link: https://arxiv.org/pdf/2610.02023
- Abstract
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study ($N=42$) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: this https URL
- 中文摘要
虽然大型语言模型(LLMs)推动了3D室内场景合成,但当前流水线未能在会话间保留用户特定的偏好,使沉浸式创作成为重复且体力消耗巨大的过程。我们介绍SPHERE,一种自适应的虚拟现实生成框架,将孤立的综合转化为连续的人机共创。SPHERE从自然的多模态交互(语音和控制器编辑)中提取持久空间偏好。为确保几何对空间扭曲的韧性,它将这些原始编辑抽象为层级约束,建模局部功能和全局拓扑上下文。此外,人机在环中强化学习机制根据用户最终编辑的场景动态更新检索策略。一项混合设计用户研究($N=42美元)和离线消融显示,SPHERE显著减少了纠正编辑和物理需求,防止了对浅层对象特征的偏向,从而产生几何弹性、与配置文件对齐的布局。最终,SPHERE展示了如何通过捕捉展示的空间逻辑实现受控的空间适应,建立了一个可靠、受控的人机协作框架,用于沉浸式创作。项目页面和源代码将可在以下网站获取:此 https URL
CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
CARM:用于大型语言模型强化学习的取消感知响应掩蔽
- Authors: Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02039
- Pdf link: https://arxiv.org/pdf/2610.02039
- Abstract
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
- 中文摘要
近年来,强化学习(RL)在大型语言模型(LLM)训练后迅速采用,数学推理和代码生成取得了显著提升。然而,在实际系统中,策略更新以及部署引擎与训练引擎之间的差异可能导致抽样响应偏离策略。序列级掩蔽通过决定整个响应是否应贡献于优化来解决这一不匹配。一种常见的掩蔽规则使用采样标记概率比的长度归一化几何平均值。其带符号的对数比率可以跨位置相互抵消,从而隐藏了显著的双向策略漂移。我们提出了\emph{消去感知响应掩蔽}(CARM),这是一种序列级掩蔽,在平均前取每个令牌对数比率的绝对值,防止相反概率变化相互抵消。我们证明,被接受的响应满足在指定区间外的抽样标记比率比例及其超出边界的平均对数距离上的联合界限。数学推理和代码生成的实验表明,CARM在AIME 2024/2025/2026及以后阶段的平均mean@16值比几何平均掩蔽提升了最多3.13美元百分点,并且在四个代码基准测试中平均pass@1比最强评估基线提高了2.88美元。这些发现支持CARM作为一种理论基础且有效的大型语言模型强化学习中响应级非策略控制方法。
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
同态优势算符:在完全同态加密约束下稳定强化学习
- Authors: Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2610.02074
- Pdf link: https://arxiv.org/pdf/2610.02074
- Abstract
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to a unique recursive error phenomenon known as the Bellman drift. This article introduces the Homomorphic Advantage Operator (HAO), a stabilization framework designed to prevent polynomial approximation divergence in FHE-based deep RL. HAO adapts the zero-mean centering projection from advantage-based value estimation directly to temporal-difference (TD) targets. This linear projection annihilates the uniform state-value baseline that drives the Bellman drift, maintaining per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping. The proposed HAO framework was evaluated using a three-tier experimental methodology, including a tabular Markov Decision Process (MDP), an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features. The results demonstrate that the proposed HAO strictly bounds network pre-activations within the safe polynomial approximation domain. The proposed HAO RL agents achieved 0% boundary breaches across all random seeds used, whereas regularization alone (L2 weight decay and gradient clipping) breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes. Finally, HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients.
- 中文摘要
保护隐私的机器学习在智能系统中对拥有机密数据的云端部署带来了重大挑战。全同态加密(FHE)为安全计算提供了一种令人信服的解决方案,能够保持云计算的数据机密性。然而,将FHE应用于强化学习(RL)需要用多项式近似替代非线性运算,而多项式近似由于一种独特的递归误差现象——贝尔曼漂移,导致灾难性的分歧。本文介绍了同态优势算子(HAO),这是一个旨在防止基于FHE的深度强化学习中多项式近似发散的稳定框架。HAO将基于优势的价值估计中的零均值中心投影直接调整到时间差分(TD)目标。该线性投影消除了驱动Bellman漂移的均匀状态值基线,保持每状态动作排名,同时无需额外的非线性乘法深度,避免昂贵的密文自助。所提出的HAO框架采用三层实验方法评估,包括表格马尔可夫决策过程(MDP)、使用真实CKKS密码操作的加密CartPole环境,以及具有密集连续特征的20节点物流路由基准测试。结果表明,所提出的HAO严格限制网络预激活在安全多项式近似域内。所提出的HAO RL代理在所有随机种子中实现了0%的边界突破,而仅通过正则化(L2权重衰减和梯度裁断)在5个种子中有3个突破边界,且不稳定基线在83.8%的事件中达到了这一点。最后,HAO代理在表格域中将最佳策略精度提升18.0个百分点,且在为截断梯度加入DP-SGD式高斯噪声时保持稳定。
Finetuning with Sampling: SFT Learns Better Than You Think
采样微调:SFT学习效果比你想象的更好
- Authors: Aayush Karan, Sitan Chen, Yilun Du
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02140
- Pdf link: https://arxiv.org/pdf/2610.02140
- Abstract
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
- 中文摘要
向前沿模型引入新能力一直是后训练的目标,后训练主要采用监督微调(SFT)和强化学习(RL)。传统观点认为,强化学习能够在不失去现有能力的情况下强推广新任务,而SFT则容易出现泛化薄弱和灾难性遗忘。同时,SFT可以从非策略专家数据中学习,而强化学习必须依赖模型通过反复抽样找到成功轨迹的能力。在我们的研究中,我们试图利用策略内学习的优势,同时利用非策略数据中包含的特权信息。然而,我们没有修改学习目标以适应这些数据,而是调整数据分布以更好地适应学习者。我们引入了一种马尔可夫链蒙特卡洛(MCMC)抽样算法,该算法在给定一个微调参考模型时,逐步将非策略轨迹转化为更符合策略的路径。在科学技能习得、数学推理和开放式专业知识等任务中,我们的采样算法使SFT能够与主流的后训练技术竞争,通常推广得更好,且遗忘的基础不强。此外,经过微调的模型表现出强的分布性能,能够学习超越基础模型分布的提升。在更高层次上,我们的方法将采样视为模型原生操作符,塑造数据以促进可学习性,作为后训练栈中通用的基元提供了更广泛的实用价值。
Faynt: Scaling and Optimizing Policies for Competitive Melee
Faynt:为竞技近战调整与优化策略
- Authors: Ali Janati, Nikita Kuzmin, Rohit Swamy, Charles Niu
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02144
- Pdf link: https://arxiv.org/pdf/2610.02144
- Abstract
We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.
- 中文摘要
我们引入了Faynt,这是一套1000万和7500万参数的变形金刚策略,适用于《任天堂明星大乱斗:近战》,每个策略用一个检查点控制全部26个角色。经过强化学习(RL),10M在其支持阵容中对14个专业和多角色发布的游戏中赢得了240场(98.4%),并且每场比赛均保持胜率。这些对手保留了21帧或24帧的动作延迟;Faynt没有增加延迟,我们尚未单独分析这一差异的影响。在一项针对私人提供的零延迟Slippi-AI模型的单独评估中,10M在两种条件条件下赢得了全部68场比赛。我们研究了架构、优化、缩放和超参数转移,以指导约84万次人类回放的预训练。后训练结合了基于排名和结果的课程、75M到10M的提炼,以及仅限于Fox镜像匹配的强化学习。在初始的152局基准测试中,监督下的10M赢得了69.7%的比赛,而预训练的75M赢得了45.4%,尽管整体上控制者预测损失更高。用于监督检查点选择的加权验证损失与四个预训练和监督策略的胜率排序一致。经过监督后训练后,两种模型每分钟承受的伤害更低,早期领先优势更大,且在失去第一条生命后赢得的频率更高。在记录的游戏状态下,10M的优化推断平均为5.2毫秒,NVIDIA T4的75M平均每决策为5.2毫秒,75毫秒为8.7毫秒,不包括模拟器的执行和通信。我们将权重开源,包括基准测试套件和自动化模型锦标赛平台。
When Do Intrinsic Rewards Lead to Exploration?
内在奖励何时会引发探索?
- Authors: Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02159
- Pdf link: https://arxiv.org/pdf/2610.02159
- Abstract
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.
- 中文摘要
内在奖励旨在通过赋予代理经验价值来引导强化学习中的探索,例如通过预测误差或学习进展。然而,最大化这些奖励不必产生最有信息量的体验。我们提出了一个正式的探索标准,通过策略获得的反事实信息比较政策:他们的历史在替代策略下替代经验的效果如何。我们构建了一个单一的简单环境,其中指定的基于计数、预测误差、赋权和信息获取目标具有最大化策略,但这些策略在获取反事实信息方面是帕累托次优的。我们解释了这些失败,并建立现有内在奖励成功鼓励最优探索的条件。我们还构建了一个目标,当探索在严格符合标准时,赋予更高的价值。
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
AutoCompact:学习何时在长视野编码代理中对上下文进行压缩
- Authors: Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.02163
- Pdf link: https://arxiv.org/pdf/2610.02163
- Abstract
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
- 中文摘要
编码代理通过漫长的代码检查、搜索、编辑和测试轨迹来解决仓库级软件工程任务。随着任务推进,早期探索变得乏味,因此管理上下文不仅仅是避免溢出:代理必须决定何时压缩、保持何种工作状态以及如何继续。我们引入了AutoCompact,它训练编码代理作为策略的一部分做出这些决策。为了收集训练数据,我们运行基础代理执行编码任务,并使用法官审查其压缩决策、总结和压缩后的动作。有缺陷的输出会被修正后的替换,然后在环境中执行,因此每个轨迹都从修正后的决策中延续。我们利用这些轨迹进行监督微调,然后通过强化学习和任务成功奖励共同优化编码和压缩。在SWE-bench Verified和SWE-PolyBench Verified上的实验显示,AutoCompact相比基础模型分别提升了9.2%和5.0%的通过率。这些改进在所有评估的推理预算中均适用,包括256K上下文窗口且永不溢出,以及16K窗口的溢出触发回退压缩。
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
OmniSeek:多回合视听推理的原生工具集成
- Authors: Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.02181
- Pdf link: https://arxiv.org/pdf/2610.02181
- Abstract
We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
- 中文摘要
我们介绍OmniSeek,一种智能框架,将全能大型语言模型(Omni-LLM)转化为具备原生工具功能的主动多回合推理代理。OmniSeek不再被动地一次性处理整个视听序列,而是将证据获取纳入推理过程:它动态决定是观看还是聆听,以及在哪个时间窗口内检索不同模态中稀疏但关键的证据。通过迭代多回合协议,检索到的原始音频或视觉片段被附加回上下文中,以支持后续推理。为冷启动此功能,我们构建了一个数据引擎,综合OmniTraj-170K,这是一个多跳思维链轨迹语料库,包含交错的音频和视觉证据。我们首先在这些轨迹上监督模型,培养多回合工具使用行为,然后通过两阶段强化学习和可验证的奖励进一步优化策略。此外,我们引入了视听必然性目标,明确奖励依赖于两种模态推理的成功轨迹,避免单一模态的捷径。跨广泛基准测试的实验表明,OmniSeek能够学习自适应跨模态证据寻求,并持续提升视听推理表现。
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
通过强化稀疏自编码器特征,生成建模本质无序的蛋白质区域
- Authors: Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.02189
- Pdf link: https://arxiv.org/pdf/2610.02189
- Abstract
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at this https URL.
- 中文摘要
内在无序蛋白区(IDR)在转录调控、信号转导和亚细胞定位等细胞过程中扮演核心角色,但其功能设计仍具挑战性。基于结构的设计方法不易应用于IDR,现有蛋白质语言模型训练于全长蛋白序列,因此学习的先验偏向于折叠结构域。本文介绍IDiom,一种自回归蛋白语言模型,基于IDiom-DB训练,该数据集包含来自AlphaFold数据库的5400万个预测IDR。IDiom生成多样序列,重现自然IDR的组成、模式、基序和预测无序。为控制功能相关序列模式,我们还引入了带有稀疏自编码特征的强化学习(RL-SAE),这是一种后训练方法,奖励生成激活特定特征集的序列。在八个IDR设计任务中,RL-SAE序列平均激活30个目标特征中的90%,而激活引导则激活率为24%。我们证明,RL-SAE相比引导和监督微调能提升生成IDR的预测亚细胞定位和转录活性,并使不同生物功能相关特征能够在单个序列中组合。因此,IDiom和RL-SAE通过显式控制功能相关序列特征,实现可解释和可组合的IDR设计。更广泛地说,RL-SAE可扩展到其他蛋白质设计环境,在可解释特征提供有用设计目标时。代码可在此 https 网址获取。
HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation
HiPhy:物理上可行多原理视频生成的层级对齐
- Authors: Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.02197
- Pdf link: https://arxiv.org/pdf/2610.02197
- Abstract
Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
- 中文摘要
视频生成模型已实现了卓越的视觉真实度,具备成为通用世界模拟器的强大潜力。尽管取得了这些进步,它们仍未能生成符合物理定律的视频。在现实环境中,多个物理原理必须在同一视频中协同工作,问题更加明显;例如,“一个气球上升,蒸汽从锅中升起”需要浮力和流体动力学以连贯且同时展开。然而,现有方法大多忽视多原则交互,专注于每个视频中的单一原理。我们提出了HiPhy(层级物理对齐),这是一种强化学习框架,通过双层目标将视频生成建立在物理定律基础上:局部强制执行单个物理原理的时间动力学,并全局确保整个场景的物理和语义一致性。为支持多原则生成,我们构建了一个5万个提示词的数据集,并引入了涵盖多种共发物理事件的提示基准MultiPhyBench。我们的实验显示,HiPhy显著优于以往方法和基线,显著改善了各种基准测试中的物理常识和语义对齐,在涉及多个同时物理原理的场景中,竞争方法退化最为显著。
FERPO: Forward Entropy-Regularized Policy Optimization
FERPO:前向熵正则化策略优化
- Authors: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.02198
- Pdf link: https://arxiv.org/pdf/2610.02198
- Abstract
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
- 中文摘要
多种持续控制中的在线强化学习方法利用学习批评者的动作梯度改进策略。然而,批评者通常被训练为预测收益,准确的价值预测不一定能产生准确的行动导数,可能导致策略更新不可靠。我们提出了前向熵正则化策略优化(FERPO),这是一种基于策略的最大熵强化学习算法,利用批判值进行策略改进,且不对动作区分批评者。FERPO从由熵和Kullback-Leibler(KL)发散正则化的策略改进目标推导出最优目标动作分布。然后我们通过利用自归一化重要性抽样(SNIS)估计的前向KL目标,将该行为体拟合到该目标。通过限制目标分布偏离展开策略,KL正则化有助于保持这些重要权重的良好表现。与可能偏向目标分布模式子集的反KL目标不同,前向KL目标鼓励覆盖多个高价值模式,从而促进探索。MuJoCo Playground和ManiSkill上的实验和消融显示出竞争性能和样本效率的提升。计算基准测试还显示演员更新速度快于相对熵路径策略优化(REPPO)。
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
KaliBench:Kali Linux 上网络安全工具使用的细粒度基准测试,提供无运行时且可验证的奖励
- Authors: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2610.02206
- Pdf link: https://arxiv.org/pdf/2610.02206
- Abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
- 中文摘要
LLM越来越多地应用于网络安全工作流,期望将分析师的意图转化为工具调用。然而,现有评估侧重于基于知识的评估或端到端的代理任务,并未直接衡量LLM生成实际网络安全工具可执行命令的能力。这一差距至关重要,因为网络安全操作依赖严格的命令行接口(CLI),其中轻微的语法错误、错误的标志——值绑定或参数顺序错误都可能使执行失效。我们介绍KaliBench,这是一个精细的基准测试和数据集,用于Kali Linux上自然语言到CLI翻译,包含8,504对查询命令,涵盖1,642个工具,跨越23个能力维度和5个安全阶段。KaliBench 通过基于手稿的流水线构建,具备确定性规范化和别名感知评估,实现工具选择和参数构建的精确且可重复的评估。为确保语义正确性和可执行性,我们开发了一个多阶段验证流水线,结合基于大型语言模型的验证、沙箱终端执行和人工参与优化。基于这些细粒度、确定性信号,KaliBench 进一步实现无运行时间的可验证训练奖励。在三种评估模式和24种通用及安全导向开放权重模型配置中,没有任何开放权重模型在无限制环境中的精确命令准确率超过42%,凸显了在没有明确工具提示的情况下准确使用基于CLI的网络安全工具的困难。我们还进一步证明,基于KaliBench的可验证奖励的监督微调和强化学习显著提升了8B模型,并实现了与685B模型相当的性能。
Keyword: diffusion policy
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
iADD:提升扩散政策优化中的对齐与多样性
- Authors: Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.01789
- Pdf link: https://arxiv.org/pdf/2610.01789
- Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
- 中文摘要
基于强化学习的扩散模型后训练,如去噪扩散策略优化(Dezaising Diffusion Policy Optimization,DDPO),在奖励函数下优化反向扩散过程。然而,当前的奖励优化方法以牺牲多样性和质量为代价。本文通过细致的理论考量和方法设计,提供了更好的权衡。我们分析理论框架,数学证明扩散模型的更新可能对多样性有害,这与之前研究中提出的结论相反。此外,我们提出了基于坚实理论基础的增量费曼-Kac训练方法,以实现迄今为止最佳的比对与多样性权衡。我们进行了大量实验,并在三种不同任务中将方法与相关扩散策略优化方法进行比较,并为每个组成部分提供强力消融,从而验证了比对和多样性的显著性能提升。