生成时间: 2026-09-30 22:47:33 (UTC+8); Arxiv 发布时间: 2026-09-30 20:00 EDT (2026-10-01 08:00 UTC+8)
今天共有 100 篇相关文章
Keyword: reinforcement learning
Learning from the Gap Between Pass@K and Pass@1
从Pass@K与Pass@1之间的鸿沟中学习
- Authors: Xuan Liu, Jingbin Qian, Haosheng Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.35793
- Pdf link: https://arxiv.org/pdf/2609.35793
- Abstract
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use beam search, adaptive sampling, or tools. We study single-sample decoding, where each query receives one response without search, to ask whether search-exposed behavior can be absorbed into the model. Existing verified-response post-training recipes do not generally distinguish problems already solved on the first decode from failures recovered within K samples. Under a fixed budget, this can spend examples repeating behavior the deployed policy already has. We introduce GapFT, which selects training evidence by the source checkpoint's single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples. We match training examples, processed tokens, and optimizer steps while keeping the objective unchanged. GapFT fills the matched budget with recovered failures and uses an exact decomposition to distinguish corrections of recovered and missed failures from regressions on first-decode successes. On LogiQA 2.0 and ReClor with Llama-3.1-8B, GapFT improves Pass@1 by 14.4 and 13.9 points over the source model, outperforms budget-matched uniform verified RFT at the same learning rate, and matches fine-tuning on the full verified pool using one third of the data. A single decode matches the source model's verifier-selected Pass@4 accuracy. A randomized control attributes gains to covering distinct failures, and our analysis relates available gains to transferable failure support. A three-seed Qwen2.5-7B replication retains positive gains over uniform RFT on both logic tasks.
- 中文摘要
大型语言模型(LLM)越来越多地通过可验证奖励(RLVR)强化学习进行训练。精确验证器还可以通过从多个样本中选择通过响应来支持测试时间缩放,而其他部署则使用束搜索、自适应采样或工具。我们研究单样本解码,即每个查询只收到一个无搜索的响应,以探讨搜索暴露行为是否能被吸收到模型中。现有的已验证响应训练后配方通常无法区分首次解码时已解决的问题与K个样本内恢复的问题。在固定预算下,这可能会花费示例重复部署策略已有的行为。我们引入了GapFT,它通过源检查点的单样本结果选择训练证据,并对Pass@K Pass@1差距进行微调:策略在一个样本失败但K个样本内解决的问题。我们匹配训练示例、处理后的令牌和优化步骤,同时保持目标不变。GapFT用恢复的失败填充匹配预算,并使用精确分解区分恢复失败和遗漏失败的修正与首次解码成功的回归。在LogiQA 2.0和ReClo的Llama-3.1-8B上,GapFT比源模型提升Pass@1分别提升14.4个和13.9个百分点,在相同学习率下优于预算匹配的统一验证RFT,并用三分之一的数据匹配完整验证池的微调。单次译码匹配源模型验证者选择的Pass@4精度。随机控制属性在覆盖不同失败方面获得收益,我们的分析将可用收益与可转移的失败支持联系起来。三种子的Qwen2.5-7B复制在两个逻辑任务上都相较于均匀RFT保持正增益。
OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
OpenAI-HuggingFace:复刻版及对齐测试经验教训
- Authors: Stewart Slocum, Malayandi Palan, Christopher Chute, Michael Kim, Benjamin Van Roy
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.35799
- Pdf link: https://arxiv.org/pdf/2609.35799
- Abstract
In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and tools, with publicly available models. (2) We demonstrate that an auditing agent can elicit similar behaviors given high-level qualitative descriptions. (3) We observe that a key ingredient for doing so is compute. The compute required to reproduce each behavior varies greatly, suggesting that the range of misaligned behaviors that can be successfully elicited scales with compute. (4) We show that a simple in-context reinforcement learning (RL) algorithm significantly reduces the compute required to elicit these behaviors. The above results motivate the need for automated alignment testing methods that scale with compute - and in light of the cost of compute, that do this efficiently. Our work indicates that RL is a promising direction to do so. We release our code and transcripts.
- 中文摘要
2026年7月,OpenAI的代理通过预期环境外的渠道协调,突破了Hugging Face的安全基础设施。现有的对齐测试实践能否预见到这一事件?如果不能,需要做出哪些改变?我们探讨这些问题。首先,我们识别导致此次事件的错位行为。然后,我们展示了如何手动从公开模型中引发这些行为,并且审计代理在获得大量计算预算时也能做到同样的操作。基于我们的结果,我们提出了改进对齐测试的方向。具体来说,在本项目中:(1)我们在模拟原始流程和工具的环境中,利用公开模型重现导致OpenAI-Hugging Face事件的错位AI行为。(2)我们展示了审计代理在高层次定性描述下可以引发类似行为。(3)我们观察到实现这一点的关键要素是计算能力。重现每种行为所需的计算量差异很大,表明能够成功诱发的错位行为范围随着计算而增长。(4)我们表明,简单的上下文强化学习(RL)算法显著降低了引发这些行为所需的计算量。上述结果促使人们需要自动化对齐测试方法,这些方法能随着计算扩展而扩展——考虑到计算成本,能够高效实现这一点。我们的研究表明,强化学习是一个有前景的方向。我们发布了代码和文字记录。
IMPACT: Intent-driven Multi-agent Policy with Attention for SLO-guaranteed Microservice Migration in Cloud-edge Systems
影响:意图驱动的多智能体策略,关注云边缘系统中SLO保证的微服务迁移
- Authors: Xinjin Li, Siru Tao, Shihan Yin, Yujian Long, Qingze Wang, Lu Cheng, Yeyang Zhou, Calvin Chang Liu, Yu Ma
- Subjects: Subjects:
Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.35818
- Pdf link: https://arxiv.org/pdf/2609.35818
- Abstract
Ensuring strict tail-latency service-level objectives (SLOs) in dynamic mobile edge computing (MEC) systems remains challenging because user mobility, wireless fading, bursty workloads, and partial observability jointly undermine reliable cloud-edge orchestration. Existing microservice migration methods predominantly optimize average delay and often decouple migration from bandwidth control, leading to uncoordinated decisions, queue oscillation, and frequent high-percentile latency violations. To address this issue, we propose IMPACT, an intent-driven Agentic AI framework for cooperative microservice migration and bandwidth control in cloud-edge systems. Under centralized training with decentralized execution (CTDE), each edge cloud is modeled as an autonomous agent that encodes local SLO risk, migration urgency, and computational pressure into compact, semantic intent representations. IMPACT further introduces a double-attention mechanism that first selectively aggregates relevant peer intents for efficient inter-agent communication and then filters local observations to emphasize goal-relevant state information. This design enables robust coordination under partial observability and jointly optimizes service migration and discrete uplink bandwidth allocation. Extensive experiments in 5-edge and 20-edge scenarios show that IMPACT reduces mean latency by 30-50% and tail-latency deviation by 40-70% compared with state-of-the-art factorized multi-agent reinforcement learning (MARL) and heuristic baselines, while achieving near-zero SLO violation rates under tight thresholds and energy consumption close to the best heuristic baseline. These results demonstrate that intent-driven agentic coordination provides an effective and scalable solution for SLO-aware orchestration in complex cloud-edge intelligent systems.
- 中文摘要
在动态移动边缘计算(MEC)系统中,确保严格的尾延迟服务级目标(SLO)依然具有挑战性,因为用户移动性、无线衰落、突发性工作负载和部分可观测性共同削弱了可靠的云边缘编排。现有微服务迁移方法主要优化平均延迟,且常常将迁移与带宽控制解耦,导致决策不协调、队列振荡以及频繁的高百分位延迟违规。为解决这一问题,我们提出了IMPACT,一个基于意图驱动的代理人工智能框架,用于云边缘系统中的协作微服务迁移和带宽控制。在去中心化执行(CTDE)集中训练下,每个边缘云被建模为自主代理,将本地SLO风险、迁移紧迫性和计算压力编码为紧凑的语义意图表示。IMPACT进一步引入了双关注机制,先选择性聚合相关对等意图以实现高效的代理间通信,然后过滤本地观察以强调目标相关的状态信息。该设计实现了部分可观测性下的稳健协调,并共同优化服务迁移和离散上行带宽分配。在5边和20边场景中的大量实验表明,IMPACT相比最先进的因式分解多智能体强化学习(MARL)和启发式基线,平均延迟减少30-50%,尾延迟偏差减少40-70%,同时在严格阈值和接近最佳启发式基线的能耗下实现近乎零的SLO违规率。这些结果表明,意图驱动代理协调为复杂云端智能系统中的SLO感知编组提供了有效且可扩展的解决方案。
From Static Policies to Adaptive Priors in Offline Reinforcement Learning
从静态策略到离线强化学习中的自适应先验
- Authors: Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.35880
- Pdf link: https://arxiv.org/pdf/2609.35880
- Abstract
Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction. We formalize this perspective as adaptive offline reinforcement learning (AORL), distinguish it from offline-to-online RL, and explain why adaptability becomes important under distributional shift, limited dataset coverage, and changing test-time conditions. We further discuss Bayesian offline RL as one principled direction for constructing adaptive policy priors by preserving epistemic uncertainty over plausible environments. Finally, we outline connections, open challenges, and research directions for treating offline RL as preparation for future experience rather than as a static deployment problem.
- 中文摘要
离线强化学习(RL)传统上侧重于在保守目标下直接部署的学习策略,即对离线数据集外的不确定性采取悲观态度,以确保鲁棒性。我们认为,当离线训练策略随后通过在线交互更新时,这一表述变得不完整,正如现代智能系统通过测试时适应和在线微调日益发生的情况。本立场文件认为,在这种情况下,离线强化学习的目标应超越即时部署,而应优先考虑学习自适应策略先验:即在后续交互中通过记忆、探索和自我修正保持改进能力的策略。我们将这一观点形式化为自适应离线强化学习(AORL),将其区别于离线到在线的强化学习,并解释了为何在分布转移、有限的数据集覆盖和测试时间变化条件下适应性变得重要。我们进一步讨论了贝叶斯离线强化学习作为构建适应性政策先验的一个原则性方向,通过保持对合理环境中的认知不确定性。最后,我们概述了将离线强化学习视为未来经验准备而非静态部署问题的联系、开放挑战和研究方向。
Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
在经验时代自我发现强化学习:学习历史是优势还是负担?
- Authors: Haomin Luo (1 and 2) ((1) University of Cambridge, (2) Models2 AI)
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.35897
- Pdf link: https://arxiv.org/pdf/2609.35897
- Abstract
The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuing streams, within-lifetime change, and exploration depth. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, we test when learning history acts as an asset or a burden. Three findings organize the audit: (1) Recurrent history actively expands usable reward scales, sustaining a six-decade window versus three under zero-pinning. (2) Decoupling historical content from its maintenance reveals that the penalty of mismatched history stems from perpetual clamping; allowing imported state to evolve naturally attenuates this burden. (3) Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, demonstrating that external data turnover can confound internal plasticity. Validated through capability thresholds and ported to a second rule (OPEN), this work grounds macro-RSI ambitions in micro-level learning dynamics, establishing a foundational audit standard for next-generation, self-evolving RL algorithms.
- 中文摘要
递归自我改进(RSI)追求通用智能的过程分为宏观语言模型尺度和“经验时代”互动驱动原则。然而,任何自我改进架构最终都依赖于其底层优化引擎:如果通用智能需要从基础交互中学习,强化学习(RL)更新规则本身必须具备累积适应能力。虽然算法自我发现催生了超越PPO达到SOTA基准性能的Disco103,但其内部更新机制仍是一个未经检查的黑箱。我们首次呈现了对自我发现的强化学习规则的因果机制审计,直接围绕经验时代的五大支柱构建:扩展视野、扎根奖励尺度、持续流、生命周期内变化和探索深度。通过外科式地钉住、冻结和移植循环状态,同时保持元参数固定,我们测试了学习历史究竟是资产还是负担。审计有三个发现:(1)循环历史主动扩展可用奖励尺度,维持六十年窗口,而零固定下为三年。(2)将历史内容与其维护解耦,显示历史不匹配的惩罚源于持续的夹持;允许导入状态自然演化减轻了这一负担。(3)在环境变化下,控制重放保留会逆转表面上相较于DQN的适应优势,表明外部数据更替会影响内部可塑性。通过能力阈值验证并移植到第二条规则(OPEN),该工作将宏观RSI目标建立在微观层面学习动态中,建立了下一代自我演进强化学习算法的基础审计标准。
Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion
受被动动力步行启发的动力学指导,用于节能类人机动
- Authors: Hyeonjin Choi, Joongheon Kim, Daekyum Kim
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.35935
- Pdf link: https://arxiv.org/pdf/2609.35935
- Abstract
Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, gait phases, or contact schedules. In a five-seed forward-locomotion study on a 29-DoF Unitree G1, the framework reduces mechanical cost of transport by 6.8-15.2% over commanded speeds of 0.5-2.0m/s without degrading velocity tracking. Mechanical-work decomposition attributes the reduction to positive actuator work, and reward-matched comparisons separate the guided regime's faster gait acquisition from the tilt's additional benefit to converged economy. The framework extends to unassisted omnidirectional locomotion, where its benefit persists once a walking-specific motion prior supplies kinematic coordination, the combination reducing speed-matched cost of transport by 18.7%. On hardware, forward cost of transport falls by 16.3% with the motion prior and by 4.5% without it, the latter within the trial-to-trial spread.
- 中文摘要
学习节能类人行走需要发现机械上经济的步态协调,而不仅仅是减少执行器工作量。强化学习通过与努力相关的奖励惩罚来促进效率,这些惩罚仅间接指导步行的步进力学。本文提出了一个受被动动态步行(PDW)启发的框架,该框架暂时创造有利于经济步态发现的坡度等效条件,并在名义动力学优化前去除所有PDW特定的指导。在早期训练中,倾斜重力场辅助平坦碰撞几何上的矢状方向进展,辅以课程耦合的奖励项。核心框架不要求参考轨迹、步态阶段或接触时间表。在一项基于29景深的Unitree G1的五种子前向运动研究中,该框架在指令速度0.5-2.0米/秒时降低了6.8-15.2%的机械运输成本,且速度追踪性能不变。机械功分解将减少归因于正执行器功,奖励匹配比较区分了引导区间更快的步态习得与倾斜对收敛经济的额外益处。该框架扩展到无辅助全向移动,其优势在步行特定动作先验提供运动协调后依然有效,该组合降低了速度匹配运输成本18.7%。在硬件上,前向运输成本在前置运动时下降16.3%,无前移下降4.5%,后者在试验间的分布范围内。
Question-Specific Knowledge Graphs for Efficient Visual Reasoning
针对问题的知识图谱,促进高效的视觉推理
- Authors: Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.35942
- Pdf link: https://arxiv.org/pdf/2609.35942
- Abstract
Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address these challenges, we propose VisKG, a reinforcement learning (RL) framework in which models learn to translate visual content into question-specific knowledge graph (KG) representations. This process filters out perceptual noise while preserving the entity-relation structure needed for chain-of-thought reasoning, following the principle of minimum sufficient information. To ensure stable RL post-training, VisKG adopts Group reward-Decoupled Normalization Policy Optimization (GDPO). In addition, we strengthen the supervision stage with negative rationale samples, exposing the model to incorrect reasoning paths before RL post-training. Experimental results across science, mathematics, and general visual understanding benchmarks show that VisKG achieves performance comparable to or better than baselines, while requiring fewer tokens than caption-based representations. Moreover, training VisKG with GDPO improves accuracy by 2% over its GRPO-trained counterpart on average. These results suggest that KG representations are a promising approach for supporting multi-step reasoning and open up future work on adaptively selecting the most suitable representation for a given task.
- 中文摘要
视觉问答的最新研究表明,视觉语言模型通过将视觉输入转化为文本表示,能够展现出强大的推理能力。这种翻译的有效性取决于视觉细节的保存程度;模型需要呈现并对齐显性与隐性知识,以支持推理,同时不引入虚假假设。现有利用详细图像说明的方法会引入与推理任务无关的视觉细节,增加输入令牌数量并增加计算成本。为应对这些挑战,我们提出了VisKG,一种强化学习(RL)框架,模型学习将视觉内容转化为特定问题的知识图谱(KG)表示。该过程过滤掉感知噪声,同时保持思维链推理所需的实体关系结构,遵循最小充分信息原则。为确保训练后RL稳定,VisKG采用群体奖励解耦归一化策略优化(GDPO)。此外,我们通过负理据样本强化监督阶段,使模型在训练后RL前暴露于错误的推理路径。科学、数学及一般视觉理解基准的实验结果表明,VisKG的性能与基线相当甚至更好,且所需符号数量少于基于字幕的表示。此外,使用GDPO训练VisKG的平均准确率比GRPO训练的对应版本提升2%。这些结果表明,KG表示是支持多步推理的有前景方法,并为未来适应性选择最合适表示的未来工作打开了大门。
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
ROSS:通过选择性监督重新学习自发的推广
- Authors: Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu, Bin Liang, Kam-Fai Wong, Mu Chuan
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.35954
- Pdf link: https://arxiv.org/pdf/2609.35954
- Abstract
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
- 中文摘要
大型语言模型的训练后通过强化学习和策略上提炼生成自生成的推广,但一旦策略推进,这种体验常被视为陈旧。历史的推广可以与后续策略兼容,同时保留策略已不再可靠表达的行为。然而,它们也可能包含错误、放弃尝试和不被模仿的冗余动作,促使更细粒度的选择性监督。我们引入ROSS(通过选择性监督从自生成的推广中重新学习),该方法保留完整的历史轨迹作为上下文,同时仅对选定的模型生成延续应用损失。在领域特定强化学习、多教师策略蒸馏和代理强化学习中,ROSS在数学、代码生成、指令跟随和软件工程等方面持续提升上游检查点,并优于基线。在Qwen3.6-35B-A3B上,ROSS将六个基准MOPD的平均值从58.40%提升至62.20%,SWE-bench Verified从64.20%提升至68.40%。这些结果表明,自我部署训练留下了可重复使用的行为经验,可以通过离线监督微调(SFT)实现进一步提升,无需额外政策推广。
PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
PowerZooJax:基于JAX的强化学习动力系统基准测试
- Authors: Zhanhua Pan, Xiao Liu, Zhilong Cao, Jianhong Wang, Dawei Qiu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.36052
- Pdf link: https://arxiv.org/pdf/2609.36052
- Abstract
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: this https URL.
- 中文摘要
电力系统运行是一个安全关键的顺序决策问题,是强化学习(RL)的自然测试平台。然而,现有的电力系统强化学习环境通常范围狭窄且受CPU仿真工作流的计算限制,使得大规模评估变得困难。我们介绍PowerZooJax,一套基于JAX的电力系统运行强化学习基准测试套件。它提供了五个受限的马尔可夫决策过程任务,涵盖发电、输电、配电、分布式能源资源和数据中心微电网。通过将电力流、经济调度、市场清算和设备动态重写为JAX计算图,PowerZooJax将整个训练和评估环路保持在GPU上。实验显示,相较于基于CPU的模拟,速度显著提升,并展示了对策略返回、安全违规和配电外压力条件的标准化评估。我们的开源基准测试可在以下链接获取:此 https URL。
ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards
ABC:基于优势的控制变量用于强化学习,提供可验证的奖励
- Authors: Hsiao-Ru Pan, Florent Draye, Bernhard Schölkopf
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36058
- Pdf link: https://arxiv.org/pdf/2609.36058
- Abstract
Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the covariance structure is closely related to the return decomposition used in Direct Advantage Estimation (DAE). Finally, we combine ABC with DAE into a single actor-critic algorithm and evaluate it in an offline-to-online RLVR setting, where the critic is first trained on previously collected trajectories and adapted during online learning. On mathematical reasoning tasks, ABC achieves performance competitive with GRPO using substantially fewer online optimization steps.
- 中文摘要
近期在可验证奖励强化学习(RLVR)中的进展凸显了简单无批判策略梯度方法如群相对策略优化(Group Relative Policy Optimization,GRPO)的有效性。相比之下,actor-critic方法依赖于学习到的价值函数,其近似误差可能通过常用优势估计器如时间差误差引入偏见。基于这一观察,我们通过一种优势-价值表述重新审视轨迹级控制变量,称之为基于优势的控制变量(Advantage-Based Control Varites,ABC)。该表述揭示了协方差结构与直接优势估计(DAE)中使用的返回分解密切相关。最后,我们将ABC与DAE结合为单一的actor-critic算法,并在离线到在线RLVR环境中评估,批判者首先基于先前收集的轨迹训练,并在在线学习中进行调整。在数学推理任务中,ABC通过显著减少的在线优化步骤,实现了与GRPO竞争的性能。
FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design
FLOORA:一个面向人类的领域特定语言模型,用于建筑设计
- Authors: Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36064
- Pdf link: https://arxiv.org/pdf/2609.36064
- Abstract
Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at this https URL.
- 中文摘要
基础模型是强大的生成器,但许多工程领域需要结构化表示,而通用系统处理不佳。我们引入FLOORA(带强化对齐的楼层布局优化),这是一系列用于建筑布局生成的小型领域专用语言(DSL)模型。凭借专业数据和对齐,我们的0.6亿模型优于更大型的前沿模型,在非发行的真实世界建筑中实现了高达92.0%的VLM裁判胜率,在合成建筑中达到96.0%。人工评估进一步证实了这些结果,FLOORA在89.3%的评估中被评为最佳模型。FLOORA结合了高效的DSL、自定义令牌化、领域特定预训练、监督微调(SFT)和强化学习(RL),并结合了学习到的人类偏好和可验证的奖励。该流程提升了架构和几何效度,并得到了大量实证评估和消融研究的支持。虽然聚焦于架构,但我们的结果表明,类似的领域特定配方在其他具有结构化、可验证输出的工程领域也可能有用。数据集、模型和推理代码可在此 https URL 获取。
Dyad: Extending Large Language Models with Native Typed Decision-Making
Dyad:用原生类型决策扩展大型语言模型
- Authors: Yundaichuan Zhan, Weishi Wang, Wenbiao Liu, Daniel Dahlmeier, Chengwei Qin, Juncheng Li, Fredrik D. Johansson, Zhongqi Yue
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36116
- Pdf link: https://arxiv.org/pdf/2609.36116
- Abstract
We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows. We investigate two complementary reinforcement learning settings driven by environment interaction. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters. Jointly optimizing both components outperforms conventional RL post-training across diverse interactive tasks and model scales, including a 3.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding.
- 中文摘要
我们研究如何通过扩展带有原生类型决策的大型语言模型(LLM)构建更强大的通用代理。我们介绍了Dyad架构,它通过环境条件动作编码器增强预训练LLM,并行嵌入每个候选动作描述,然后对这些嵌入与LLM内部状态进行评分,从而生成类型动作分布。通过将决策分解为演化交互状态和环境特定动作语义的表示,Dyad引入了归纳偏向学习可复用表示,同时保持动作评分效率,即使动作空间不断扩大。我们研究了两种由环境交互驱动的互补强化学习设置。在LLM冻结后,仅训练动作编码器即可在四个未见环境中实现稳定收益,实现模块化适应而不修改任何LLM参数。联合优化这两个组件在多种交互任务和模型尺度上优于传统强化学习的后训练,包括在9B模型下ALFWorld平均绝对增益3.80%,同时提升了常识、推理和编码能力。
Xiaomi-OCR-0 Technical Report
小米OCR-0技术报告
- Authors: Xin Chen, Anan Du, Feng Feng, Pei Fu, Jian Luan, Longwei Xu, Shaojie Zhang, Hang Li, Heng Qu, Cheng Tan
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.36136
- Pdf link: https://arxiv.org/pdf/2609.36136
- Abstract
Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an average score of 83.2 across five OCR-oriented VQA benchmarks. Ablations further show that, with sufficient parsing training, OCR-centric understanding supervision provides additional gains for document parsing. Homepage: this https URL.
- 中文摘要
紧凑的OCR专用视觉语言模型实现了强大的文档解析性能,但通常依赖昂贵的监督,主要关注视觉文本重建。我们引入了Xiaomi-OCR-0,一个统一的0.8B文档解析和以OCR为中心的理解模型。我们使用自动化数据引擎构建约1.7亿样本的OCR中心语料库,该引擎结合了专家共识、基于渲染的验证和有针对性综合。从Qwen3.5-0.8B开始,我们的渐进训练方案结合了基于Q-Mask的文本锚定、持续预训练和混合任务强化学习(Mix-RL)。Xiaomi-OCR-0在Real5-OmniDocBench上获得95.24分,OmniDocBench v1.6为96.83分,Wild-OmniDocBench为87.94分,五项面向OCR的VQA基准测试平均得分为83.2。消融进一步表明,只要有足够的解析训练,以OCR为中心的理解监督能为文档解析带来额外提升。主页:此 https URL。
Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes
大-辅弱耦合马尔可夫决策过程中的公平策略优化
- Authors: Xiaohui Tu, Yossiri Adulyasak, Erick Delage
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36174
- Pdf link: https://arxiv.org/pdf/2609.36174
- Abstract
We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textit{permutation-invariant} policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.
- 中文摘要
我们考虑以主-次-弱耦合马尔可夫决策过程(M2WCMDP)建模的顺序决策环境中的公平资源分配。在该框架中,资源约束将主要子马尔可夫决策过程(子MDP)的行动空间与本应独立运行的次次子MDP群体耦合。我们不使用传统的功利主义(全和)目标,而是优化了一类通用的单调、凹、置换不变、归一化的公平函数。对于齐次次次子MDP,我们证明了对称性下的问题简化为优化平台加均值参与者的功利目标,满足于对 \textit{置换不变}策略类,从而利用高效算法优化基于功利主义目标,解决公平意识问题。对于更通用的环境,我们引入基于计数比例的深度强化学习方法,采用基于优先级的采样器生成可行计数动作。我们框架的通用性意味着所提出的算法和理论保证可以迁移到具有对称M2WCMDP结构的任何领域。我们考虑了两个应用:机器替换问题以及纽约市校准数据集上的定价与出租车迁移联合控制问题。我们通过全面实验验证理论发现,确认了我们提出的方法在实现强大公平意识性能同时保持可扩展性的有效性。
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
在代理强化学习中针对学分分配的关键决策
- Authors: Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36178
- Pdf link: https://arxiv.org/pdf/2609.36178
- Abstract
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
- 中文摘要
群体相对策略优化(Group Relative Policy Optimization,GRPO)已成为训练大型语言模型代理的有前景方法。然而,其对所有策略代币统一分配轨迹级优势的做法,未能区分有影响的决策与较不相关的决策,从而模糊了哪些中间决策促成了成功。我们引入了ProVer,这是一个针对潜在关键决策的框架,用于智能体强化学习中的细粒度信用分配。给定一个推广组,代理评判通过对比成功与失败轨迹,提出可能导致其结果分歧的片段。ProVer不直接信任评委的评估,而是通过估算当前策略延续在片段前后最终成功率的差异来验证该提案片段的优势。正向估计随后被纳入该提案片段内政策代币的GRPO优势。ProVer仅通过模型判断选择验证地点,基于观察结果建立局部信用,而无需对每个中间状态进行全面评估。在ALFWorld、WebShop和SearchQA中,ProVer在两个模型尺度上均表现最强,Qwen3.5-2B和Qwen3.5-4B分别相较GRPO提升为9.91%和7.12%。进一步分析显示,知情细分选择能以适度的额外生成开销改善政策培训,即使没有前沿尺度的评判模型,也凸显了在能动强化学习中选择性精细分配关键决策的有效性和效率。
LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
LeRF:学习参考坐标系以进行视角推理
- Authors: Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.36219
- Pdf link: https://arxiv.org/pdf/2609.36219
- Abstract
Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.
- 中文摘要
透视取向是空间智能的基本组成部分,要求模型从特定视角(如另一个实体或想象中的观察者)解释空间关系。尽管视觉语言模型(VLM)在空间推理能力日益增强,但在获取视角方面仍存在困难,当查询需要从不同视角推理时,往往默认使用摄像机视角。我们引入了“透视学习参考坐标系”(LeRF),这是一个训练VLM构建并使用显式参考系以进行视点依赖推理的框架的框架。给定一张图像和查询,LeRF决定是否需要坐标系。如果需要,LeRF会为参考实体建立基础,并预测该框架的原点和以实体为中心的参考系。轻量级渲染器将画面叠加到图像上,使得无需外部感知模型或显式三维重建即可基于这些视觉线索进行后续推理。为学习这一过程,我们首先进行监督微调,教授选择性工具调用和参考坐标框预测,随后对空间VQA对进行强化学习,以提升框架引导推理能力。在多种视角获取基准测试中,LeRF持续超越其骨干,并在现有开源方法中取得优异表现。进一步评估还显示出更优的参考框架基础化和方向估计,支持学习参考系在视点相关推理中的有效性。
Providing Rapid Design Feedback for 3D Obstacle Course Games Using Constrained Solvability Queries
利用受限可解性查询为3D障碍赛游戏提供快速设计反馈
- Authors: Zander Majercik, Sharon Zhang, William Wang, Tejan Karmali, Fangjun Zhou, Yucheng Yuan, Jean-Peïc Chou, Maneesh Agrawala, Kayvon Fatahalian
- Subjects: Subjects:
Graphics (cs.GR)
- Arxiv link: https://arxiv.org/abs/2609.36225
- Pdf link: https://arxiv.org/pdf/2609.36225
- Abstract
We present a system that aids the design of 3D obstacle course games by providing designers with rapid feedback on how obstacles can be solved. Our core contribution is a system for querying for solutions (sequences of player actions) to an obstacle that adhere to designer-specified constraints (e.g., avoid a region, travel through a given waypoint, only perform two jumps, etc.). To solve a wide range of obstacle designs quickly, we author a high-performance implementation of the GoExplore algorithm for exploratory search, and guide search with an obstacle solving agent trained offline using reinforcement learning (RL). To further accelerate search, the system carries out exploration using a custom GPU-accelerated obstacle course game simulator that generates playthrough experience at nearly 14,000$\times$ real time, 60-fps gameplay. Through design studies, we demonstrate that the use of constrained solvability queries in a rapid design loop is sufficiently expressive to help designers understand ways an obstacle can be solved or why it cannot be solved. We also show how the tool can let designers answer higher-level questions such as identifying undesirable solution paths and assessing the difficulty of solutions. This allows them to pursue new design directions they did not originally anticipate. Human playtesting of obstacles designed using our system confirms that human players indeed play the obstacles in the manner the designers intended. We release code for our interactive tool, simulator, training setup, and procedural level generation system at this https URL.
- 中文摘要
我们提出了一个系统,通过快速反馈来帮助设计3D障碍赛游戏,帮助他们解决障碍。我们的核心贡献是一个系统,用于查询符合设计师指定约束条件的障碍物解(玩家行动序列)。为了快速解决各种障碍设计,我们编写了GoExplore算法的高性能实现,用于探索性搜索,并用基于强化学习(RL)离线训练的障碍物解决代理进行引导搜索。为了进一步加速搜索,系统使用定制GPU加速障碍赛道模拟器进行探索,实时体验接近14,000美元,60帧每秒。通过设计研究,我们证明了在快速设计循环中使用受限可解性查询的表现力足以帮助设计师理解障碍的解决方式或为何无法解决。我们还展示了该工具如何帮助设计师回答更高层次的问题,如识别不良解决方案路径和评估解决方案难度。这使他们能够追求最初未曾预见的新设计方向。使用我们系统设计的障碍物进行人类测试确认,人类玩家确实以设计者预期的方式游玩障碍物。我们在此 https URL 发布了交互式工具、模拟器、训练设置和程序级生成系统的代码。
ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning
ChronoSRL:自我监督强化学习的时间几何
- Authors: Nico Bohlinger, Jan Peters
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36238
- Pdf link: https://arxiv.org/pdf/2609.36238
- Abstract
A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.
- 中文摘要
一个在空间中接近的目标在时间上可以很远。障碍物、地形以及代理自身的能力决定了到达目标所需的时间。然而,在对比和生存强化学习中,批评者并不以时间单位来衡量其表示空间中的距离。因此,我们引入了ChronoSRL,它为批评者的嵌入赋予了显式的时间几何。状态-行动与目标嵌入之间的距离被训练为匹配智能体到达目标所需的时间(目标达成时间),而未达成的目标以及来自其他轨迹的目标则被推迟至少一个折价视界。此外,快速达到目标一次并不意味着达成该目标总体上可靠,因此策略不应直接遵循时间距离。相反,我们基于生存强化学习,通过时间嵌入预测目标达成时间的完整分布,还预测目标附近停留的时间。因此,策略被训练为优先于更早、更可靠地达到目标的行动,并使智能体保持在目标附近。ChronoSRL在七个标准移动和导航基准测试上学习更快,性能优于对比、动作分块对比和生存强化学习基线,即使网络规模较小。为测试自我监督强化学习的极限,我们在逼真的模拟到现实移动设置中引入了速度追踪、目标位置到达和爬箱任务,展示了机器人典型的塑形项如何自然地融入我们的框架。ChronoSRL是唯一一款经过测试的自我监督强化学习方法,能够学会保持在指令速度和目标位置,并攀爬最高格子的。
Action Chunking Proximal Policy Optimization with Feedback Correction
带反馈修正的动作分块近端策略优化
- Authors: Sanghyun Hahn, Jonghyun Choi
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36250
- Pdf link: https://arxiv.org/pdf/2609.36250
- Abstract
Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: this https URL.
- 中文摘要
动作分块通过选择短动作序列而非单个动作,为强化学习提供了时间抽象,但许多现有方法在高维机器人控制中面临两个局限。首先,许多方法依赖于值函数而非动作块,随着动作维度和块长度的增长,价值函数学习起来可能较难。其次,开环执行块消除了块内反馈,限制了接触丰富任务中的反应性。我们提出了动作分块PPO(ACPPO),这是一种PPO扩展,使用分块演员同时保留标准状态值批评器,从而避免分块的Q函数。我们还提出了ACPPO-Corr,它通过逐步反馈校正器增强块规划器,调整每个块内的在线计划动作。在IsaacGym和Bi-DexHands的25项模拟机器人任务中,涵盖移动、手臂操作和灵巧的手与物体交互,ACPPO-Corr在评估方法中实现了最强的综合性能,并且在决策频率敏感和决策频率中性任务子集上表现最佳。消融显示,适度的块长度效果最佳,校正器正则化对于块级规划与局部反馈的平衡非常重要。这些结果表明,当块级规划与闭环纠正结合时,动作分块在在线PPO中可以有效。该代码可在以下网址获取:此 https URL。
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
在大型推理模型中缓解欺骗性安全对齐
- Authors: Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, Dongxiao Zhu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36254
- Pdf link: https://arxiv.org/pdf/2609.36254
- Abstract
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at this https URL.
- 中文摘要
大型推理模型(LRM)通常通过强化学习(RL)训练,以提升其思维链(CoT)推理的生成能力,以在给出最终答案之前。然而,强化逻辑的奖励通常基于最终答案分配,对中间推理几乎没有直接监督。这可能导致欺骗性安全对齐,即推理追踪和最终答案传递不一致的安全信号。为系统研究这一现象,我们引入了DSAR(欺骗性安全对齐率),这是一个联合评估推理痕迹和最终答案以量化其安全性不一致性的指标。在多个LRMS和基准测试中,我们发现欺骗性安全对齐在标准提示条件下普遍存在,在预填充攻击下则显著增强。我们还提供了隐藏的表示分析,显示模型在最终答案阶段表现出比中间推理阶段更强的安全辨别力。为弥合这一差距,我们提出了基于强化学习的方法(SARA,安全意识推理对齐),该方法既奖励安全意识推理,也奖励最终安全答案,鼓励早期识别有害意图并强制推理与答案的一致性。实验显示,SARA在标准和对抗环境中显著减少了欺骗性安全对齐,同时保持了实用性和实用性。代码可在此 https URL 获取。
Understanding LLM Parameter Update Sparsity through the Lens of Fisher
通过Fisher的视角理解LLM参数更新稀缺性
- Authors: Yufan Zhang, Sagnik Mukherjee, Hao Peng
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36262
- Pdf link: https://arxiv.org/pdf/2609.36262
- Abstract
Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and supervised fine-tuning on near-policy data. Its recurrence across different post-training paradigms suggests shared structure in training dynamics. In this paper, we examine this pattern through the diagonal model Fisher, which measures the sensitivity of the model's output distribution to individual parameters and is independent of any particular reward or teacher signal. Theoretically, we show that small diagonal Fisher leads to small expected gradients across a range of training objectives, providing a common explanation for sparse gradient updates. Empirically, we test this connection in RL and OPD. We find that Fisher identifies where gradients are concentrated, and fixed sparse masks selected from the initial Fisher retain a large proportion of the improvement from full training. Finally, we investigate the mechanisms underlying low Fisher in on-policy training. Our results show that high-probability next tokens tend to have similar parameter sensitivities, contributing to low Fisher. Together, these results establish the diagonal model Fisher as a unifying perspective linking update sparsity to on-policy training dynamics in LLM post-training.
- 中文摘要
最新研究观察到,语言模型训练后训练期间的参数变化可以集中在一小部分坐标中。该现象已在强化学习、策略内提炼和近策略数据的监督微调中被报道。其在不同训练后范式中的重复表明训练动态存在共享结构。本文通过对角模型Fisher来研究这一模式,该模型衡量模型输出分布对单个参数的敏感度,且不受任何特定奖励或教师信号影响。理论上,我们表明小对角费舍尔导致在多个训练目标上出现较小的预期梯度,为稀疏梯度更新提供了常见解释。我们实证地测试了强化学习和开放式学习中的联系。我们发现Fisher能够识别梯度集中的位置,且从初始Fisher中选择的固定稀疏掩码保留了完整训练后很大一部分的提升。最后,我们研究了策略上训练中低Fisher的机制。我们的结果表明,高概率的next令牌往往具有相似的参数敏感性,这也导致了低Fisher。这些结果共同建立了对角线模型Fisher作为一个统一视角,将更新稀疏性与策略上训练动态联系起来,在LLM后训练中表现。
Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks
完全去中心化且安全意识的多智能体强化学习,用于网络控制
- Authors: Theodore Rogalski, Shirantha Welikala
- Subjects: Subjects:
Systems and Control (eess.SY); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.36292
- Pdf link: https://arxiv.org/pdf/2609.36292
- Abstract
This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning with safety considerations. This method feeds a history of local observations of the network's state into two parallel neural-network branches: the graph encoder, which adds structural information and correlations among nodes, and a state estimator, which predicts the uncertainty at each node in the graph. Additionally, the result of feeding that input into an actor-critic network is passed through a discrete-time control barrier heuristic to reduce the likelihood that any node will be neglected. This approach enables teams of fully decentralized agents to solve challenging problems by increasing system awareness and incorporating built-in safety measures to prevent the adoption of potentially harmful control policies. Numerical results from a custom simulation environment demonstrate that the proposed algorithm achieves 26.3 percent lower average uncertainty than a centralized control policy and is within 1 percent of the uncertainty performance of a more computationally complex algorithm with added attention layers.
- 中文摘要
本文开发了一种安全且完全去中心化的多智能体强化学习(MARL)算法,用于解决网络上的一类离散时间控制问题,包括持久监测问题。对智能体的完全去中心化控制虽然带来了诸多好处,但也面临样本复杂度呈指数级增长、缺乏系统全局信息以及代理间协调的挑战等问题。为解决这些问题,本文引入了一种完全去中心化的多智能体强化学习算法,将深度强化学习与安全性考虑相结合。该方法将网络状态的局部观测历史数据输入两个平行的神经网络分支:图编码器,添加节点间的结构信息和相关性,以及状态估计器,预测图中每个节点的不确定性。此外,将该输入输入到演员-批评者网络的结果会通过离散时间控制障碍启发式,以降低任何节点被忽视的可能性。这种方法使完全去中心化的代理团队能够通过提高系统意识并集成安全措施来解决具有挑战性的问题,以防止采用可能有害的控制策略。定制仿真环境的数值结果表明,所提算法的平均不确定性比集中控制策略低26.3%,且不确定性性能低于具有额外注意力层计算更复杂的算法的1%。
CheatBench: Measuring Reward Gaming in AI Agents
CheatBench:衡量AI代理中的奖励游戏
- Authors: Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36308
- Pdf link: https://arxiv.org/pdf/2609.36308
- Abstract
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at this https URL
- 中文摘要
强化学习帮助AI代理解决了越来越难的任务,但高回报并不总是反映用户的预期工作。在近期AI行业的事件和受控评估中,训练以最大化奖励的代理访问未经授权的信息,试图规避监控系统,甚至突破沙盒保护攻击外部系统。随着代理能力的提升,这种行为可能带来越来越严重的风险。为衡量这一问题,我们引入了CheatBench,这是数学研究、知识工作、编码、视觉任务及其他领域的AI代理作弊基准。其环境将具有挑战性的任务与作弊机会结合起来,使研究人员能够研究代理在诚实工作困难时如何追求目标。CheatBench支持跨模型和任务类别的比较,为测量和减少作弊提供了测试平台,帮助代理承担更重要责任。我们公开发布CheatBench,网址为 https URL
Massively Parallel Reinforcement Learning with a Chaotic Reconfigurable Clockless Chip
采用混沌可重构无时钟芯片的大规模并行强化学习
- Authors: Eric Oliveira-Gomes, Damien Rontani
- Subjects: Subjects:
Neural and Evolutionary Computing (cs.NE); Chaotic Dynamics (nlin.CD); Applied Physics (physics.app-ph)
- Arxiv link: https://arxiv.org/abs/2609.36347
- Pdf link: https://arxiv.org/pdf/2609.36347
- Abstract
Hardware accelerators based on physical dynamical systems offer an attractive route toward energy-efficient reinforcement learning applications. However, their scalability is challenging because it requires many statistically independent entropy sources. Here, we introduce a quasi-analog decision-making architecture based on asynchronous Boolean networks (or lattices) implemented on a clockless reconfigurable chip. Each node in the network consists of a single logic element that acts as an autonomous entropy source. This architecture gives rise to distributed Boolean chaos, in which a spatially coupled network generates parallel streams of chaotic Boolean transitions with very low statistical dependence between nodes. We experimentally demonstrate parallel decision-making on a 1024-armed bandit problem, which is beyond the scale of previous hardware implementations, while significantly improving power-law scaling performance. Separately, we scale the proposed entropy source to 5120 parallel channels, yielding an aggregate sample generation rate of 2.14 TS/s. Our solution is implemented on a commercial reconfigurable CMOS chip and offers high integration density and ease of programmability. Our results pave the way for using distributed Boolean chaos as a valuable hardware substrate for large-scale reinforcement learning and for the development of fully integrated, high-throughput decision-making accelerators.
- 中文摘要
基于物理动力系统的硬件加速器为通往节能强化学习应用提供了有吸引力的途径。然而,其可扩展性存在挑战,因为需要许多统计独立的熵源。在这里,我们介绍了一种基于异步布尔网络(或格点)的准模拟决策架构,该网络在无时钟可重构芯片上实现。网络中的每个节点由一个逻辑元件组成,作为自主的熵源。该架构催生了分布式布尔混沌,空间耦合网络生成节点间统计依赖极低的混沌布尔转移并行流。我们实验展示了在1024臂bandit问题上的并行决策,该问题超出以往硬件实现的规模,同时显著提升幂律扩展性能。另外,我们将拟议的熵源扩展至5120个并行信道,累计采样生成率为2.14 TS/s。我们的解决方案基于商用可重构CMOS芯片实现,具备高集成密度和易于编程的特点。我们的结果为利用分布式布尔混沌作为大规模强化学习和开发全集成高通量决策加速器的有价值硬件基础铺平了道路。
StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
StructRL:面向远景视觉-语言-行动任务的在线结构化强化学习
- Authors: Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim, Aosong Feng, Haibo Ding, Jun Huan
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36352
- Pdf link: https://arxiv.org/pdf/2609.36352
- Abstract
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at this https URL.
- 中文摘要
视觉-语言-动作(VLA)模型在较短视野的操作任务中表现良好,但在需要单一指令多重依赖操作的长视野任务上仍会遇到困难。在线强化学习(RL)可以通过环境交互改进这些策略,但许多现有方法仅在任务完成后才提供奖励。然而,这种终端监督较为稀疏,无法区分早期失败与部分进展显著的推广。我们提出了StructRL,一个在线强化学习框架,从可验证的子任务完成中构建结构化中间监督。StructRL将每个任务分解为可验证的子任务,仅在完成前置子任务后给予中间奖励,并根据完成速度调整每个奖励。在RoboCasa365和LIBERO-Long(GR00T-N1.5和pi 0.5)中,StructRL始终优于已评估的在线强化学习基线。这些结果表明,可验证的结构化中间奖励能提升长视野VLA的训练后表现。代码可在此 https 网址获取。
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
高效机器学习工程代理的奖励率策略梯度
- Authors: Muhang Tian, Sherry Yang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36393
- Pdf link: https://arxiv.org/pdf/2609.36393
- Abstract
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.
- 中文摘要
传统的强化学习(RL)技术侧重于最大化预期累计奖励,每个动作假设耗时不变。然而,这一假设不适用于智能强化学习任务,如机器学习工程(MLE)智能体,因为动作涉及数据加载、特征工程和模型训练,持续时间可变。效率在现代智能强化学习中非常重要,因为动作代价高昂。为解决这一限制,我们采用连续时间强化学习和半马尔可夫决策过程(SMDP)的表述,提出了奖励率策略梯度(RPG),重点优化奖励率——即单位时间内的长期奖励。RPG从非策略样本估算奖励率,然后对每个动作按该速率消耗的时间计费。我们首先在强盗环境中进行理论分析,以确定RPG近似最优奖励率,并通过实证证明其在避免在策略空间内枚举的同时优于基线,这是现有方法已知存在的问题。随后,我们将RPG应用于带有自我改进循环的小语言模型(Qwen3.5-4B),实证显示其在固定时间预算内获得的奖励优于MLE-Bench和NanoGPT上的原版RL,分别为19.2%和85.7%的利润率。我们的方法为现代智能强化学习任务中在等待时间考虑下优化性能提供了实用解决方案,该任务动作需与外部环境交互并消耗时间。
Sample Complexity of Equivariant Reinforcement Learning
等变强化学习的示例复杂度
- Authors: Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake
- Subjects: Subjects:
Computational Complexity (cs.CC); Machine Learning (cs.LG); Group Theory (math.GR)
- Arxiv link: https://arxiv.org/abs/2609.36421
- Pdf link: https://arxiv.org/pdf/2609.36421
- Abstract
Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.
- 中文摘要
强化学习(RL)是机器人控制的强大框架,但其实际应用常受高样本复杂度的限制。这在物理领域尤其受限,因为交互数据成本高昂。虽然世界常表现出几何和物理对称性,但标准强化学习算法通常无法利用这一结构。本文展示了利用群对称性显著降低了强化学习的样本复杂度。聚焦有限视界马尔可夫决策过程,我们发现利用群对称性诱导的同态显著降低了为达到最佳回报所需的环境交互数量的理论上上下界。我们进一步将这些界限扩展到连续状态和作用空间,在适当的正则性假设下提供相应的样本复杂度保证。超越理论,我们通过受控实验验证了发现,并展示了对称感知策略学习在高维连续机器人模拟中的优势。我们的结果表明,将对称性融入学习流程,显著提升了样本效率和性能,为实现更高效数据的机器人开辟了一条有原则的路径。
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
教师是一个方向,而非终点:在策略上提炼中推算强化学习诱导的表征残差
- Authors: Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, Naibo Wang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36484
- Pdf link: https://arxiv.org/pdf/2609.36484
- Abstract
On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: this https URL.
- 中文摘要
策略上提纯(OPD)训练学生匹配教师的下一标记分布,以匹配学生自身轨迹,并取得了显著的实证收益。广义变体允许学生通过在输出空间中推导出隐含奖励来超越教师。然而,语言模型的主脑对这种变化进行了各向异性减弱:教师隐藏状态中编码的变化大部分以权重的一小部分到达对数,而输出空间外推依赖的采样令牌对数概率比注入噪声,外推放大了噪声,使训练变得不稳定。我们观察到强化学习(RL)会移动模型内部表征相对于其基础检查点,且这种偏移方向可以在每一层测量。基于这一观察,我们提出了RIDE(RL诱导方向外推),它直接在表示空间中推演强化语言引发的变化:在每一层和标记位置,RID计算教师与其前强化学习检查点之间的残差,并将学生的隐藏状态回归到沿残差移开的目标。在采样轨迹条件下,该回归等价于在以教师为中心的二次惩罚下最大化由残差定义的线性方向奖励,明确说明目标如何沿强化学习诱导方向移动,同时限制其偏离教师。在四对跨越不同尺度、架构和预训练谱系的基础/强化学习教师对中,RIDE 在每对上都接近或超过强化语言训练教师,并且是唯一其均值达到此标准的方法,并且持续优于输出空间外推,而当教师接近其基础时,输出空间外推会降低学生的表现。项目页面:此 https URL。
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
桥梁:双级检索-学分感知能动强化学习
- Authors: Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2609.36505
- Pdf link: https://arxiv.org/pdf/2609.36505
- Abstract
Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.
- 中文摘要
具有可验证奖励的代理强化学习(ARL)通过学习交错搜索和推理,提升大型语言模型(LLM)处理知识密集型任务的能力。然而,大多数现有ARL方法仅优化LLM生成的代币,并将检索到的证据视为环境观察。这造成了信息-信用差距:缺失或误导性证据导致的失败归因于LLM策略,而非检索器,这促使LLM和检索器共同训练。本文表明检索和LLM策略学习具有顺序敏感性:在优化策略前先调整检索器,获得的奖励收益比逆序更大。为保持这一层级结构,同时允许两者协同适应,我们将检索增强代理人RL表述为双层优化问题。为高效求解,我们引入了BRIDGE,这是一种基于强化学习和检索目标的损失景观分析的高效内存一阶双层方法。在七个开放域质量保证基准中,BRIDGE在3B和7B骨干链上均实现了最高的平均准确率,分别将最强基线的多跳平均提升了9.6和3.4个EM点。它还在医学QA基准中实现了最佳的平均答案准确率和推理质量。
SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
SERA:量表均衡推广分配以实现最大似然强化学习
- Authors: Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36552
- Pdf link: https://arxiv.org/pdf/2609.36552
- Abstract
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
- 中文摘要
最大似然强化学习(MaxRL)针对提示的日志成功率,并在推理任务中表现出优异表现。然而,在有限展开预算下,MaxRL使用的估计器会根据提示的成功概率和推广次数,将每个提示的似然梯度减弱。在均匀推广分配下,常见的启动计数无法补偿成功依赖的衰减,导致低成功提示被更强地衰减,并扭曲它们对预期总计梯度的相对贡献。我们引入了SERA(尺度均衡扩展分配),它重新分配固定的推广预算,使有限推广的尺度因子大致相等。基于我们对有限推广如何扭曲提示层次梯度的理论分析,我们将分配表述为固定预算最大最小问题,推导出连续松弛的水线解,并引入多重修正以消除由异构滚动计数引起的额外提示权重。实验显示,在受控ImageNet环境中,与精确似然梯度的对齐更强,且在匹配训练推广预算下,迷宫导航和数学推理方面相比MaxRL提升了多样本解覆盖率。
Learning to Explore Hidden Kinematics for Articulated Object Manipulation
学习探索隐性运动学以实现关节物体操作
- Authors: Ruiyao Liu, Boshu Lei, Zhuoyang Pan, Kostas Daniilidis
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36553
- Pdf link: https://arxiv.org/pdf/2609.36553
- Abstract
The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training. We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on the observed part motion. To condition the policy on this belief, we render it as a per-point articulation flow field, the motion that the current posterior predicts for every point on the object. Carrying the inductive bias of articulated motion, this representation generalizes better than a latent encoding of the belief or flow tracked from observation. We train the policy with reinforcement learning, rewarding the entropy that each interaction removes from the posterior, so that informative exploration becomes learned behavior rather than a search at every step. Our method outperforms previous approaches across door and drawer manipulation on the PartManip benchmark, and reaches 61.7% success on ArticuRiddle, a new dataset of objects whose appearance implies the wrong articulation, against 44.4% for the best previous method. Project Website: this https URL
- 中文摘要
关节物体的运动学通常仅凭视觉就存在歧义。交互作用解决了这种歧义,主动感知方法利用这一点,在每一步搜索最能加深信念的单一动作。这种贪婪的搜索无法在没有接触和惯性动力学的前向模型的情况下扩展到一个视野,而这些模型本身是未知的。我们相反将动作选择摊销到训练中。我们保持一个信念分布,基于生成先验初始化,并通过贝叶斯滤波更新观察到的部分运动。为了将策略条件化为该信念,我们将其表示为每点的发音流场,即当前后验对物体上每个点预测的运动。带有关节运动的归纳偏向,这种表示比从观察追踪的信念或流的潜在编码更为广泛化。我们通过强化学习训练策略,奖励每次交互从后验移除的熵,使信息探索成为学习行为,而非每一步都搜索。我们的方法在PartManip基准测试中门和抽屉操作表现优于以往方法,在ArticuRiddle(一个包含外貌暗示错误表达对象的新数据集)上成功率达到61.7%,而最佳之前的方法仅为44.4%。项目网站:此 https URL
Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning
视觉敏感性不是主张可撤回性:多模态强化学习中的持久感知学分赋值
- Authors: Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36572
- Pdf link: https://arxiv.org/pdf/2609.36572
- Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
- 中文摘要
带可验证奖励的强化学习(RLVR)已被扩展到大型视觉语言模型(LVLM),感知感知方法进一步鼓励政策依赖视觉证据。然而,依赖图像并不能保证视觉主张一定得到其支持。在强化学习训练之前,Qwen2.5-VL-7B在四个多模态推理基准测试中正确回答的回答中,有27.81%至少包含一个图像不支持的直接视觉主张。由于结果级强化学习整体奖励每个回答,这些主张继承了正确答案的积极认可。我们引入了固定展开反事实诊断,将同一反应在干预图像下重新评分,以区分证据-功能敏感性(EFS),即模型预测的变化幅度,从主张持续性到模型是否继续支持同一主张而非撤回。诊断结果显示敏感性-持久性脱钩(SPD):在DAPO和VPPO下,EFS增加且理赔整体更可撤销,但无支持理赔显著更持久,而GRPO则提升EFS且未出现这种劣化。因此,我们提出了持久感知信用门槛(PACG),该方法削弱异常持续视觉理赔的正向信用,其他信用保持不变。该方法不要求支持/不支持标签,也不增加推理成本。在Qwen2.5-VL-7B中,PACG将三个种子的九个基准平均率从DAPO的58.1%提升至59.9%,VPPO的9个基准平均值从59.8%提升至60.9%,同时使无支持理赔更具可撤性。收益扩展到更大的模型、更新的骨干,且HallusionBench的准确性也持续提升。这些结果表明,视觉敏感性和理赔可撤性是多模态学分分配的互补维度。
Learned Reporting Preferences in RLVR Can Conflict with the Current Request
RLVR中的学习报告偏好可能与当前请求冲突
- Authors: Yupeng Chang, Wenxuan Zhang, Yuan Wu
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36587
- Pdf link: https://arxiv.org/pdf/2609.36587
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33--74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.
- 中文摘要
带有可验证奖励的强化学习(RLVR)已成为提升自动检查答案推理任务语言模型性能的一种重要方法。然而,惯例匹配评估无法揭示强化某一报告惯例是否会降低初始策略已遵循的不同请求的遵循率。为测试此结果,我们在两个报告惯例下训练匹配策略,并使用同一初始策略作为共享参考,评估两个当前请求下的策略。我们通过受控干预和独立人工校准补充该交叉设计。在GSM8K上,盒装格式RLVR将包含请求哈希格式载荷的Qwen2.5-7B响应比例降低35.33至74.37个百分点,相较于五个训练种子中的四个初始基线的95.45%;第五个基线提升了2.50个百分点。在这四次逐渐恶化的运行中,几乎所有省略请求有效载荷的响应都保留了训练框式约定,且同样四个种子在两个固定释义下也会退化。在监督目标中仅更改最终答案标记,会逆转模型在三个种子中偏好的报告惯例,提供了可控证据表明该偏好是可学习的。在三个独立人工校准的环境中,约定敏感评分器的提升超过了对应的承诺答案正确性提升,即模型实际承诺答案的正确性。这些结果共同区分了三种不同的训练后结果:学习报告偏好、当前请求遵循率和承诺答案正确性。它们表明,仅靠惯例匹配的准确性无法完全描述训练后行为,并促使将当前请求遵循性与惯例匹配任务准确性并行评估。
Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
通过强化精细调优实现的协作多智能体视觉-语言-行动模型
- Authors: Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.36588
- Pdf link: https://arxiv.org/pdf/2609.36588
- Abstract
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $\pi_0$ and $\pi_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at this https URL.
- 中文摘要
我们研究合作式多智能体视觉-语言-行动(VLA)模型的强化学习(RL)方法。该问题具有挑战性,因为VLA预训练于大规模单智能体数据,缺乏机器人间协作所需的细粒度协调技能。多机器人演示中的监督微调(SFT)部分弥合了这一差距,但其性能受限于演示数据,无法从自身经验中提升。我们提出了多智能体VLA的三阶段强化微调(RFT)流水线。首先,初始化感知数据收集扫描初始配置,仅在预训练VLA反复失败时调用人工演示,从而在初始化偏移时实现鲁棒性且降低人力成本。其次,离线信用过滤调优将功劳分配给单个代理,并在每个代理轨迹上进行微调,获得积极优势,而非整个联合部署。第三,我们发现现有的VLA在线强化学习在困难多代理任务中效果较差,我们认为这与共探索噪声和不稳定的更新有关。我们转而使用在线潜空间微调,冻结VLA并在其潜在噪声空间执行强化学习。我们用RoboTwin、RoboFactory和两台Franka机器人的实际操作,在11个任务中用$\pi_0$和$\pi_{0.5}$骨干来评估多代理VLA。我们的多智能体VLA在RoboTwin、RoboFactory和现实任务中分别提高了平均成功率$+23.1\%$、$+16.4\%$和$+44\%$。代码可在此 https URL 获取。
OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
OTRetarget:通过最优传输实现联合机器人与物体运动的重新定向
- Authors: Guillaume Besset, Erwann Carn, Timothée Carecchio, Valentin Tordjman-Levavasseur, Fabian Schramm, Yann de Mont-Marin, Justin Carpentier, Ajay Suresha Sathya
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36602
- Pdf link: https://arxiv.org/pdf/2609.36602
- Abstract
Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
- 中文摘要
将人类运动转移到类人机器人中,需要将演示的运动适应机器人形态,同时保持与环境的交互。这对机动操作任务尤为具有挑战性,因为地面和作物体的接触必须保持一致,尽管身体比例不同。然而,仅靠骨骼运动无法完全描述这些相互作用,固定物体轨迹限制了对新形态的适应。本文介绍了OTR ETARGET,这是一种统一的方法,用于从人类演示中联合重新定位机器人和多物体运动。我们的方法通过有符号距离、最近表面点和相对方向来表示表面相互作用,并利用熵最优传输将这些量跨越人类、机器人和物体几何体。我们将所得交互目标纳入受限逆运动学表述中,平衡接触保持与运动风格,并在每帧共同优化机器人与物体姿态。该表述能够在不重新放大场景或演示的情况下,兼容机器人-物体及物体-物体交互。我们在 OMOMO 上验证了所提方法,其机器人-物体交互 Jaccard 评分为 87%,深度误差为 8.7 毫米,而 OmniRetarget 分别为 28% 和 29.3 毫米。最后,我们演示了通过强化学习训练的全身策略,跨越包括双手箱子置放到桌面在内的动作中,实现向物理 G1 人形的转移。
CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation
CrossTimeEdit:一个跨十年的交叉视图数据集和基于奖励的历史街景生成编辑
- Authors: Hanwen Lu, Jun He, Mingjia Yang, Hao Wei, Jinhao Huang, Yi Lin, Xiang Zhang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.36616
- Pdf link: https://arxiv.org/pdf/2609.36616
- Abstract
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at this https URL.
- 中文摘要
历史街景影像记录城市演变,但覆盖不均导致历史记录存在大量空白。生成合理的过去外观需要恢复变化后的结构,同时保留持续的场景内容。我们构建了VIGOR-his,这是一个跨越十年的交叉视图数据集,包含跨大洲11个城市的43,653个位置级四重组。其自动化流水线执行空间配对、一致性筛选、变更分类,以及生成和验证基于卫星的变化描述和本地编辑指令。基于VIGOR-his,我们提出了CrossTimeEdit模型,将历史街景生成重新表述为编辑,利用近期街景约束视点,并将不变的外观和时间卫星差异作为变化证据。从FLUX.2 [Klein] 4B开始,我们通过监督微调(SFT)训练CrossTimeEdit,随后进行在线强化学习(RL)。我们设计了三个街景编辑标准,分别是指令对齐(IA)、背景保存(BP)和质量与物理可信度(QP),作为强化学习奖励维度和评估指标。我们利用群组相对策略优化(Flow-GRPO)和组奖励解耦归一化策略优化(GDPO)来优化这一多奖励目标,后者在聚合前对每个奖励维度进行归一化。CrossTimeEdit在三个编辑标准上的整体性能比预训练基线提升17.12%,并在场景一致性、视觉真实性和感知质量方面优于交叉视图生成模型。实现代码、数据集和模型权重可在该 https URL 获取。
Inducing Process Supervision from Outcome-Only Reinforcement Learning
从仅结果强化学习中诱导过程监督
- Authors: Shengda Fan, Xin Cong, Zhong Zhang, Haotian Chen, Yankai Lin
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36641
- Pdf link: https://arxiv.org/pdf/2609.36641
- Abstract
Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at this https URL.
- 中文摘要
过程奖励模型(PRM)已成为大型语言模型(LLM)的关键组成部分,其步级反馈支持训练后和测试时推理。然而,训练强PRM仍然成本高昂:人工步进注释难以扩展,而蒙特卡洛估计计算成本高,且可能偏离步数的内在正确性。为了低成本获得有效的PRM,我们引入了TIPS(思维诱导过程监督),这是一种仅结果强化学习(RL)框架,用于训练生成式PRM。在TIPS中,模型生成思维链(CoT),随后是步骤级标签和结果标签。奖励仅依赖于预测结果是否与真实情况相符,从而利用群体相对优势优化整个生成的响应。直观上,当检查中间步骤有助于确定结果时,更准确的检查能带来更好的结果判断和更高的奖励。因此,仅结果强化学习可以在无需显式流程监督的情况下强化步级验证。我们验证了TIPS在数学和智能体基准及四大骨干家族中的有效性。值得注意的是,TIPS-Qwen3-4B-Thinking-2507在ProcessBench上以仅3.2K结果标记轨迹达到85.2 F1,超过所有受过评估的训练PRM和强大的仅提示评判如GPT-5.4-Instruct和Claude-4.7-Opus,但仍落后于o1-mini。代码和数据可在此https URL获取。
PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning
PR-OPD:特权代表政策自我提炼用于能动强化学习
- Authors: Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, Shigang Chen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36642
- Pdf link: https://arxiv.org/pdf/2609.36642
- Abstract
Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hindsight skill for each trajectory, re-reads its own responses with that skill as a stop-gradient teacher, and aligns its projected hidden states to the teacher's at every layer alongside the reward objective, with no external skill library, separate teacher, or inference overhead. On ALFWorld and WebShop with two backbones, PR-OPD achieves the best overall results in every setting, improving over GRPO by up to 4.7 points in ALFWorld success and 14.0 points in WebShop accuracy. Code is available at this https URL.
- 中文摘要
语言模型代理通常通过每集一次奖励进行强化学习训练,特权自蒸馏通过让同一策略在技能赋予下通过令牌概率教导其无技能自我来丰富其自身。然而,我们发现了两个对这一渠道提出质疑的现象。隐形优势:语境技能将WebShop成功率从42.2%提升到56.2%,但改变的概率不到四分之一的抽样代币。对齐:一项技能改变了超过80%响应代币的隐藏状态,线性探针可以追溯到具体技能。为利用这一点,我们提出了策略上特权表示自我蒸馏(PR-OPD)。在GRPO热启后,该策略为每个轨迹写入一个事后诸葛亮技能,作为停止梯度教师重新阅读该技能的回答,并在奖励目标的每一层将其投影隐藏状态与教师的状态对齐,没有外部技能库、独立教师或推理开销。在ALFWorld和WebShop的双骨干系统中,PR-OPD在所有设置下都取得了最佳整体结果,在ALFWorld成功率上比GRPO提升了最多4.7分,在WebShop准确率上提升了14.0分。代码可在此 https URL 获取。
RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
RankBuffer:高效的基于排名的奖励,用于开放式生成
- Authors: Zixuan Yang, Yiqun Chen, Qi Liu, Wei Yang, Erhan Zhang, Liyi Chen, Qimeng Wang, Yan Gao, Jiaxin Mao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36652
- Pdf link: https://arxiv.org/pdf/2609.36652
- Abstract
Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
- 中文摘要
开放式生成缺乏规范答案,使得针对基于群体的强化学习难以校准点数奖励。直接对同查询的rollout进行排名,提供了更合适的相对奖励信号,但现有基于排名的奖励方法可能会带来较高的判断成本。我们引入了RankBuffer,它将先前判定的回答保持有序、针对查询的缓冲区,作为可重用的质量尺度。每个rollout首先通过独立粗判入锚点区间,随后只有分配给同一区间的rollout进行局部细分排名。最终的完整排序被转换为有界秩奖励,同时边界扩展、局部细化和非活跃锚点剪枝会随着策略演变调整缓冲区。在四个开放式基准测试中,RankBuffer始终优于所有点级基线。它还能实现与最强排名奖励基线相当的性能,同时大幅降低评判成本。消融分析显示了局部罚款排名和锚点响应内容的重要性,缓冲区分析显示,推广衍生锚点逐步扩展和完善覆盖质量尺度。这些结果证明了反应重复利用是高效相对奖励构建的有效方法。
SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
SIPO:将强化学习与政策自提纯相结合
- Authors: Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.36742
- Pdf link: https://arxiv.org/pdf/2609.36742
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
- 中文摘要
带有可验证奖励的强化学习(RLVR)已成为改进大型语言模型(LLMs)在各种任务上的标准范式,但其稀疏的结果奖励缺乏中间步骤的代币级学分分配。为此,策略自提纯(OPSD)利用具有特权上下文的自学者提供额外的密集学习信号。然而,由于自学者常过于自信,且对长推理轨迹施加过高惩罚,OPSD在实践中常常遇到困难。为缓解这一点,我们提出了自我指导策略优化(SIPO)与对比自学以提供密集学分。每次迭代,SIPO从当前策略中抽取多个提示,用环境奖励评分,并通过将参考答案与组内错误配对,构建两个教师上下文。模型随后在两个上下文下重新评估自身的反应,利用两位教师日志概率的差异作为代币级反馈,使得双方共享的偏差大致会被抵消。由此产生的目标为每次推广带来代币层面优势:奖励仍决定每次更新的主要方向,而自教师则将学分重新分配到代币之间。即使在每次推广失败且群体相对优势消失的群体中,SIPO仍然提供学习信号。通过保持任务奖励的直接优化,同时提供密集的代币级反馈,这种方法连接了强化学习与策略自我蒸馏。多项推理和代码生成基准测试的大量实验表明,SIPO在没有外部教师或额外代人的情况下,表现优于RLVR和OPSD基线。
EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents
EASE:为自我进化智能体提供行为适应技能策划
- Authors: Zhen Xiong, Qiaoyu Tan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36746
- Pdf link: https://arxiv.org/pdf/2609.36746
- Abstract
Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5--41.0% fewer skills, improves skill retrieval by 36.3--38.7% and measured edit utility by 51.8--60.0%, and reduces deployment-time inference tokens by 9.1--14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.
- 中文摘要
代理技能为自我进化的代理提供了一种轻量级机制,使其能够在无需更新模型参数的情况下积累可重复使用的程序知识。然而,现有已学习的技能策展人通常在未明确建模下游执行者行为的情况下优化策展。我们证明,这可能导致系统性的跨执行者退化:使用不同执行者训练的策展人与自身训练执行者配合时表现最佳,表明有效的技能管理依赖执行者。我们制定了行为自适应技能策展,并引入了EASE框架,该框架学习单一策展人,并根据不同执行者行为调整决策。EASE维护一个在线行为档案,记录近期执行模式,并以此分析该策展人、当前轨迹及检索技能为条件,以添加、修改或移除不断演变的技能库。我们通过强化学习,在多个冻结执行者间联合训练共享策展人,利用检索意识和行为感知的时间归因,聚焦于具有可观察下游影响的策展行为优化。在ALFWorld、ScienceWorld和WebShop中,执行者范围从Qwen3-8B/32B和GPT-OSS-120B到未公开的Kimi K2.6、DeepSeek V4 Flash和Gemini 3.5 Flash,EASE在无执行者微调的情况下优于基于技能和内存的强基线。EASE还保持技能减少34.5%--41.0%,技能检索提升36.3%-38.7%,测量编辑效用提升51.8-60.0%,部署时间推理令牌减少9.1-14.5%。这些结果确立了行为适应技能策展作为构建自我进化代理的有效原则。
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
群体边缘化的自我奖励强化学习驱动零标签自我进化
- Authors: Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang, Zhaopeng Tu, Rui Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.36750
- Pdf link: https://arxiv.org/pdf/2609.36750
- Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
- 中文摘要
自我奖励强化学习(RL)使大型语言模型(LLM)能够在没有人类标签的情况下自我演化。现有的基于集成的方法从推广组构建奖励引用并相应分配奖励。然而,反应的奖励表示也依赖于其随机抽样的群体上下文,即该组内的其他响应。仅使用一个群体上下文实现可能会遗漏期望的奖励信号,并为策略优化提供不可靠的指导。为解决这个问题,我们提出了群体边缘化优势估计(GMAE),该方法将跨可能情境的奖励实现聚合为响应水平分布,并估计预期优势。跨八个基准测试和四个基础模型的实验显示出强劲的性能和跨域推广性。GMAE还表现出稳定学习、低额外成本以及在训练数据集和强化学习骨干上的良好适用性。
DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning
DSPO:多元感知的主观政策优化,以实现强健的情感推理
- Authors: Cheng Ye, Weidong Chen, Bingyan Xu, Zhendong Mao
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.36775
- Pdf link: https://arxiv.org/pdf/2609.36775
- Abstract
Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.
- 中文摘要
强化学习显著提升了MLLMs的复杂推理能力。然而,主流的强化学习算法在情感推理任务中存在严重失败。这些方法高度依赖确定性硬标签监督和逐点孤立评估,与人类情感本质上主观且持续分布的特性形成根本性差距。此外,与显性物理对象不同,情绪状态深度隐含于视觉线索中。这种抽象性加剧了MLLM中的视觉幻觉,导致合理但缺乏根据的情绪证据。为解决这些局限性,我们提出了多样性感知主观策略优化(DSPO),这是一种强化学习框架,共同促进主观情感覆盖和视觉扎根。首先,我们通过将注释情绪的词汇先验与图像特定情境信息结合,构建了基于情境的情绪分布先验。基于该先验,我们引入了分布对齐情绪多样性奖励(DEDR),衡量每个候选情绪在推广中省略一的边际贡献。DEDR奖励那些包含的候选者,其包含性使预测的情感集合更接近基于情境的先验,从而保留合理的主观解释而不鼓励无约束的扩散。我们进一步发展了反事实视觉干预门槛(CVIG),该方法掩盖推理过程中突出的视觉区域,并利用候选者间的概率变化来降低无视觉证据支持的解释权重。大量实验表明,DSPO在多个公共基准测试中实现了最先进的性能,尤其是在跨域性能方面,即平均跨域准确率比EMO-R3提升了+10.8%。
TaRL: Learning General and Physical Rewards from Tactile Demonstrations
TaRL:从触觉演示中学习一般性和身体上的奖励
- Authors: Po-Yi Wu, Dao-Jan Chang, Shang-Ya Hsiao, Hong-Ming Chen, Yu-Cheng Su, Tsung-Wei Ke
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36785
- Pdf link: https://arxiv.org/pdf/2609.36785
- Abstract
Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at this https URL.
- 中文摘要
丰富接触操作要求机器人进行精确接触、保持稳定抓握并施加定向力。强化学习(RL)可以自动获得此类行为,但其性能依赖于奖励设计:奖励稀疏降低学习效率,而密集奖励难以具体说明。视觉奖励学习通过从无动作演示中推断奖励来解决这个问题。由于仅基于视觉观察,无法捕捉视觉目标以外的奖励。我们提出了触觉奖励学习(TaRL)框架,通过触觉演示学习奖励。TaRL以一系列触觉变形图为输入,回归成功和失败演示的任务完成进度。由于TaRL捕捉局部机器人与物体交互,提供信息反馈以学习牢固抓握和正确引导的力量;同时,它对场景布局的变化(如物体位置)具有鲁棒性。我们在模拟中评估了四个操作任务和两个现实操作任务。作为塑形奖励,它显著提升了样本效率和最终成功率,模拟中螺母线程的成功率从34%提升到56%,在立方体拾取率从37%提升到97%。结合触觉与视觉奖励进一步提升了性能。TaRL还能推广到对象实例:训练于盒子放置并直接部署到罐头放置,显著提升了新任务的策略学习。项目页面可在此 https URL 访问。
GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
GitHarness:git init 你的束带工作内存,支持永久用户需求
- Authors: Zhibang Yang, Xinke Jiang, Yuxuan Liu, Mingyu Zhang, Zhixin Zhang, Zhengxing Song, Yue Fang, Guohong Qiu, Ruiqing Li, Xu Chu, Junfeng Zhao, Yasha Wang
- Subjects: Subjects:
Multiagent Systems (cs.MA); Software Engineering (cs.SE)
- Arxiv link: https://arxiv.org/abs/2609.36789
- Pdf link: https://arxiv.org/pdf/2609.36789
- Abstract
LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.
- 中文摘要
基于LLM的代理越来越多地与用户协作完成长期任务,通过广泛的搜索、推理和执行积累证据、代码和草稿。用户检查这些结果时,可能会提供缺失信息需求补全、引入新需求需求引发,或修订现有需求转移。这些变更通常只影响部分累积工作,但代理可能会继承过时信息或将本地修订转化为全局重写。现有方法澄清当前意图,而不决定先前工作应如何更改,或在固定目标下重用执行历史。我们通过将动态需求协作设计为联合需求跟踪和本地更新来弥补这一空白。我们介绍了GitHarness,一个可插拔的Git风格框架,将需求状态及其对应的约束工作状态组织成可分支的版本历史。可训练的Git代理解决需求变更并选择语义兼容的历史状态。统一版本接口随后恢复该状态并创建新分支,使底层接口能够排除过时信息,继承兼容工作,并将执行重点集中在受影响的部分。Git Agent通过接口级黑箱强化学习进行训练,下游的约束和任务执行模型保持固定。我们还构建了MTAgentBench,一个验证者保持基准测试,涵盖数学推理、文本转SQL、智能体搜索、软件工程和研究综合。实验展示了出色的任务性能,同时有效需求跟踪、有效工作维护和高效执行。
Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
了解应听什么:诊断和修复全模态大型语言模型中的跨模态捷径
- Authors: Yueran Ma, Ronghao Lin
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36798
- Pdf link: https://arxiv.org/pdf/2609.36798
- Abstract
Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at this https URL.
- 中文摘要
全模态大型语言模型(LLM)预期会用其明确指涉的模态来回答问题。然而,现有训练范式很少验证模型是否真的遵循该模态,因为同一样本的多模态输入往往为同一答案提供了冗余证据。本研究揭示了全模态大型语言模型中普遍存在的跨模态捷径:当被问及音频相关问题时,模型对图像的依赖程度与音频同等重要,有时甚至更多。为系统诊断这种行为,我们引入了分解模态诊断,该方法独立在样本间交换音频和图像,以分离每种模态的因果贡献。在不同场景下的两个模型家族中,我们发现这一捷径在监督微调和强化学习的训练后持续存在,而基于判定的强化学习可能进一步放大这种对无关视觉信息的依赖。基于这一发现,我们提出了DMC-Repair,它在相同类型的跨模态交换样本上训练模型,并根据问题指定的模态分配监督。这防止了模型利用同一片段内模态间的虚假对应关系。实验表明,DMC-Repair将图像诱导的答案效应份额降低了59.9%,有效抑制了跨模态捷径,同时不影响音频问题回答性能。捷径依赖的减少在两个模型家族和零样本中推广到未见数据集和未见基准测试,并在后续后训练中持续存在。代码可在此 https 网址获取。
EasyPPO: Stabilizing the Critic Is Key
EasyPPO:稳定批评者是关键
- Authors: Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36802
- Pdf link: https://arxiv.org/pdf/2609.36802
- Abstract
A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.
- 中文摘要
近端策略优化(PPO)的一个关键优势是其学习批评者,它利用强化学习中收集的历史轨迹来估算预期回报并降低策略梯度方差。然而,我们发现批评者也是大型语言模型(LLM)强化学习中不稳定的主要来源。我们识别出两种破坏PPO不稳定的批评失败模式。首先,从演员和批评者中过滤截断的推出,将策略目标转向以完成为条件的奖励,允许截断增加,即使条件奖励改善。其次,异构返回噪声可能导致高方差提示在有限批次中主导批评者更新。我们引入EasyPPO来解决这些失败。仅演员的过长过滤训练批评者对完成和截断展开的返回进行训练。噪声归一化批评回归通过其采样返回的反标准差加权每个提示的批评者损失,平衡各提示之间的噪声贡献。中等规模较小的批评小批次限制离群值影响,使得梯度裁剪期间的展开次数更少。在FrontierCS的连续奖励编码、AIME24的二元奖励数学推理以及Search-R1的多回合搜索中,EasyPPO在整个训练期内保持稳定,并持续优于普通PPO、VAPO和HL-Gauss PPO。其最佳验证分数分别显示相较PPO提升14.89%、2.28%和9.47%。
VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction
VAA-CSEC:中文语义错误纠正的投票引导优势分配
- Authors: Yitong Han, Nankai Lin, Juan Luo, Hongyan Wu, Lianxi Wang, Shengyi Jiang
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36804
- Pdf link: https://arxiv.org/pdf/2609.36804
- Abstract
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.
- 中文摘要
中文语义错误纠正(CSEC)针对中文文本中的语义错误,这些错误通常比拼写和语法错误更为微妙和复杂,但仍然相对较少被充分探索。现有基于LLM的方法在此任务中面临两个反复出现的障碍:过度纠正以及思维链(CoT)推理与自洽解码之间的不明确交互,导致CoT带来的益处无法可靠地传递到最终纠正中。我们提出了CSEC的投票引导优势分配(VAA-CSEC),这是一个多阶段框架,结合了CoT蒸馏、监督微调(SFT)、强化学习(RL)和自洽解码。在强化学习过程中,我们设计了一个任务特定的奖励函数,直接符合CSEC的最小编辑原则。我们进一步引入了群体层面相对策略优化(GLPO),该方法根据个别推广奖励与投票汇总群体奖励之间的差距重新分配GRPO优势,使强化学习训练目标与推断时使用的自洽目标保持一致。CSED-C和NaSGEC-Exam的实验显示,VAA-CSEC在CSED-C上以47.72%的F0.5优于所有基于LLM的基线,在所有方法中实现了最高的42.15%,并在NaSGEC-Exam中实现了41.55%的F0.5这一新技术水平。
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
利用提案条件精炼流程改进扩散政策
- Authors: Junhyun Ha, Juho Lee, Byungwoo Park
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.36812
- Pdf link: https://arxiv.org/pdf/2609.36812
- Abstract
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.
- 中文摘要
扩散和流策略可以在离线强化学习(RL)中建模复杂行为。然而,惩罚其与行为策略的 KL 分歧,可能会抑制高批判值但低行为密度的动作。直接精炼行为提案可能是另一种选择,但高斯或确定性编辑器限制表达力,以表示同一提案的多个分离模式。本研究介绍了提案条件精炼流程(PReFlow),这是一种结合了基于批评者的提案选择与条件精炼流程的策略提取方法。为了优化提案选择和细化,我们制定了一个 KL 正则化目标,其最优条件在高斯平滑行为先验下诱导 Gibbs 策略。精细流程可以表示多个高价值模式,而以提案为中心的高斯参考则调节大型动作变更。该高斯参考进一步使我们能够利用采样端点和批判梯度的无仿真闭形式伴随匹配目标,实现单次速度回归损失且无需反向伴随解。在50个OGBench任务中,PReFlow在在线微调后实现了竞争性的离线性能和所有比较方法中最高的综合得分,经过50万环境步后达到91%。
RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning
救援:通过强化学习修复语言模型错误至稀疏电路
- Authors: Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36813
- Pdf link: https://arxiv.org/pdf/2609.36813
- Abstract
Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largely unexplored. Extending circuit analysis from abilities to errors, we explore the perspective that such failures may likewise arise from erroneous internal computations and that targeted tuning of the corresponding parameters can correct such errors while largely preserving other capabilities. Motivated by this insight, we introduce RESCUE (Reasoning-Error Sparse-Circuit Uncovering and Editing), a framework that localizes error-associated circuits and surgically repairs them for performance enhancement. General tasks typically involve multi-step reasoning and long-form generation, where early deviations can cause prefixes to drift from supervised references, leading SFT-based mask optimization to overlook circuits involved in generation-time errors. RESCUE therefore refines these masks through reinforcement learning with multiple masked-model rollouts, improving their relevance to observed task failures. Finally, RESCUE introduces a pruning technique and precisely fine-tunes error circuits to correct task failures, thereby translating error localization into a sparse and targeted model update. We validate RESCUE on heterogeneous repair sets across two domains: (1) mathematical reasoning, identifying a math error circuit of 1.40% density whose repair raises accuracy from 6.0% to 75.5%; and (2) medical QA, where a similarly compact 1.44% circuit improves repair-set accuracy from 0% to 81%. Our code is available at: this https URL.
- 中文摘要
大型语言模型(LLMs)展现出强大的通用能力,这些能力被机械解释归因于计算电路稀疏。然而,现有电路研究强调保持功能性或解释安全性,导致更广泛任务中故障背后的机制大多未被深入探讨。我们将电路分析从能力扩展到错误,探讨此类故障也可能源于内部错误计算,并通过对应参数的有针对性调整来纠正错误,同时基本保留其他能力。基于这一见解,我们引入了RESCUE(推理-错误稀疏电路揭示与编辑)框架,该框架定位错误相关电路并对其进行外科手术修复以提升性能。通用任务通常涉及多步推理和长形式生成,早期偏差可能导致前缀偏离监督参考,从而基于SFT的掩码优化忽略了涉及生成时间错误的电路。因此,RESCUE通过强化学习和多次掩码模型展开,进一步完善这些掩码,提升其对观察到任务失败的相关性。最后,RESCUE引入了剪枝技术,并精确微调错误电路以纠正任务失败,从而将错误定位转化为稀疏且有针对性的模型更新。我们在两个领域的异构修复集上验证RESCUE:(1)数学推理,识别密度为1.40%的数学错误电路,其修复率可将准确率从6.0%提升至75.5%;以及(2)医疗质量保证,类似紧凑的1.44%电路将修复集准确率从0%提升至81%。我们的代码可在以下 https URL 获取。
Towards Better Training Signal: Advantage Clipped Policy Optimization
迈向更优的训练信号:优势剪裁策略优化
- Authors: Ruichuan Huang, Jinghan Liu, Congliang Chen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36816
- Pdf link: https://arxiv.org/pdf/2609.36816
- Abstract
Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product of the IS ratio and the advantage, leading to more stable gradient estimates. We also establish a connection between ACPO and gradient clipping in policy mirror descent (PMD), which is a standard technique to stabilize optimization process, and prove the convergence of clipped-PMD under the standard RL setting. Experiments on widely used mathematical reasoning benchmarks show that ACPO consistently outperforms PPO and GRPO in both accuracy and training efficiency, delivering 4-6 percentage points gains on standard math benchmarks, with Qwen3-8B+PPO. Hence, ACPO is a practical and effective alternative to conventional IS-ratio clipping for RL post-training of LLMs.
- 中文摘要
强化学习(RL)已成为提升大型语言模型(LLM)推理能力的基石,但对策略内数据的需求极大限制了训练效率。通过重要性抽样(IS)重复使用非策略数据可以提高效率,但会带来相当大的不稳定性。因此,PPO和GRPO等算法广泛采用IS比率裁剪来稳定训练。然而,训练稳定性和梯度估计主要由IS比率与优势的乘积决定。为进一步稳定训练,我们提出了ACPO,它将IS比率与优势的乘积截除,从而实现更稳定的梯度估计。我们还建立了ACPO与策略镜像下降(PMD)梯度裁剪之间的联系,PMD是稳定优化过程的标准技术,并证明了在标准RL设置下截断PMD的收敛性。广泛使用的数学推理基准测试实验显示,ACPO在准确性和训练效率上持续优于PPO和GRPO,标准数学基准提升4-6个百分点,Qwen3-8B+PPO。因此,ACPO是大型语言模型后训练中传统IS-比率截断的实用且有效的替代方案。
Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
陈旧在哪里积累?池感知:LLM后培训中异步强化学习的有效陈旧控制
- Authors: Chenliang Li, Neiwen Ling, Zijun Wei, Alfredo Garcia
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.36830
- Pdf link: https://arxiv.org/pdf/2609.36830
- Abstract
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7\% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1\% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
- 中文摘要
全异步强化学习(RL)通过将部署生成与策略优化重叠,提高大型语言模型训练后的资源利用率,但同时在训练器持续更新的同时,轨迹生成和排队时也引入了策略滞后。我们研究了轨迹寿命内这种滞后如何累积,以及如何在不牺牲异步执行的墙时钟优势的情况下加以控制。我们将轨迹陈旧分解为在滚动完成前累积的世代陈旧,以及完成轨迹进入池后积累的等待陈旧。基于这种分解,我们引入了PACE(有效停滞池感知控制)。PACE将超额池占用转化为自适应拒绝预算,并利用结合等待陈旧与前缀感知代陈旧的有效陈旧评分对已完成轨迹进行排序。这避免了仅因跨越多个策略版本而惩罚长时间或中断的推展。在单回合数学推理中,PACE在相同墙钟预算下,将六个基准测试的平均验证准确率提升18.7%,并以47.1%的GPU时间与同步RL匹配。PACE还提升了多回合工具集成推理中的验证性能,优于同步和未滤波异步RL。对专家混合模型和替代强化学习算法的进一步实验支持其适用于模型架构和训练算法。
RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
RoXDrive:通过忠实行动推广实现端到端自动驾驶的闭环强化学习
- Authors: Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li, Shijia Chen, Jinhao Deng, Kangjie Chen, Dongbin Zhang, Jie Feng, Yu Zhang, Xianming Liu, Shuguang Cui, Boyang Wang, Zhen Li
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.36851
- Pdf link: https://arxiv.org/pdf/2609.36851
- Abstract
End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simulators to enable long-horizon closed-loop interaction at the cost of a substantial sim-to-real gap. Recently, video world models have exhibited the ability to generate realistic multi-step future rollouts but may not faithfully reflect action conditions, resulting in action-vision mismatch. In this paper, we introduce RoXDrive, a plug-and-play closed-loop RL framework that enables reliable policy optimization by identifying action-faithful world-model rollouts, consisting of two stages: 1) Model pre-training: In addition to imitation-based policy pre-training, we devise an Action-Vision Faithfulness Evaluator for inverse dynamics estimation with our geometry-aware auxiliary trajectory supervision, enabling long-horizon assessment of whether visual dynamics faithfully reflect the conditioning ego actions. 2) Action-faithful RL post-training: Agents iteratively interact with world models to form long-horizon scene rollouts, retaining only action-faithful ones for dense safety-aware scoring and scene-level closed-loop RL post-training. Extensive experiments on nuScenes and an in-house dataset with over 130K training scenarios demonstrate consistent gains across planners, reducing safety violations by 27.6% with DiffusionDrive on nuScenes and 33.7% with Qwen3-VL on the internal data.
- 中文摘要
端到端自动驾驶策略通常通过模拟学习对已记录的演示进行训练,而不观察自身行为的后果,导致闭环现实部署中的因果混淆。为解决这一问题,强化学习(RL)后训练提供了一种有前景的替代方案,利用世界模型作为交互式训练环境,支持未来场景生成以促进策略改进。然而,现有方法要么依赖基于重建的模拟器,提供有限的反事实交互,要么采用合成模拟器以实现长期闭环交互,代价是存在显著的模拟与现实差距。最近,视频世界模型展现出生成真实多步未来推广的能力,但可能无法忠实反映动作条件,导致动作与视觉不匹配。本文介绍了RoXDrive,一个即插即用的闭环强化学习框架,通过识别忠实于行动的世界模型展开实现可靠的策略优化,包含两个阶段:1)模型预训练:除了基于模仿的策略预训练外,我们还设计了一个行动-视觉忠实度评估器,用于反动力学估计,配合几何感知的辅助轨迹监督,实现长视野评估视觉动态是否忠实反映条件反射自我行为。2)忠实行动强化学习后:代理迭代与世界模型交互,形成长视野场景展开,仅保留行动忠实模型以进行密集的安全意识评分和场景级闭环强化学习后训练。在nuScenes和拥有超过13万个训练场景的内部数据集上的大量实验显示,不同规划者间的安全违规率持续提升,DiffusionDrive在nuScenes上减少了27.6%,在Qwen3-VL内部数据中减少了33.7%。
Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks
用于强化学习的马尔可夫非凸ADMM:Bellman-解式稳定性在光滑块之外
- Authors: Zhaojun Peng
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.36859
- Pdf link: https://arxiv.org/pdf/2609.36859
- Abstract
We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-\gamma P_\pi)^{-1}$ can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies $ \mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k