生成时间: 2026-09-29 22:49:42 (UTC+8); Arxiv 发布时间: 2026-09-29 20:00 EDT (2026-09-30 08:00 UTC+8)
今天共有 144 篇相关文章
Keyword: reinforcement learning
MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning
MaD-RL:校准LLM与强化学习的匹配分布
- Authors: Sourabh Kulkarni, Ksheeraj Sai Vepuri, Basar Demir, Jason Bohrer, Emily Shen, Jianfa Chen, Nan Jiang, Ankit Jain, Harihar Subramanyam, Mannat Singh, Chirag Nagpal
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31644
- Pdf link: https://arxiv.org/pdf/2609.31644
- Abstract
Reinforcement learning (RL) is widely used in language-model post-training to maximize rewards assigned to individual model outputs, such as scores from binary verifiers or reward models trained on human feedback. However, applications such as synthetic-data generation, fairness-related constraint satisfaction, and policy exploration require controlling the distribution of outputs across model generations rather than only maximizing expected reward. We propose a general RL-based framework for \textit{Distribution Matching} allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution. Empirically, we demonstrate that dominant post-training recipes such as Group Relative Policy Optimization (GRPO) reduce output diversity by concentrating policy probability towards a single mode. Entropy regularization and sampling temperature can improve the spread of the distribution but have constrained effectiveness, limited to apply only in token space and toward uniform distributions. We show that prior work in this area is a specific case of Distribution Matching involving the $L_2$ divergence. We then propose reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justification. Finally, we demonstrate the effectiveness of our approach on a set of experiments involving mathematical reasoning and programming.
- 中文摘要
强化学习(RL)广泛应用于语言模型后训练,以最大化分配给单个模型输出的奖励,例如二元验证器的得分或基于人类反馈训练的奖励模型。然而,合成数据生成、公平性相关约束满足和策略探索等应用需要控制模型世代间输出的分布,而不仅仅是最大化期望奖励。我们提出了一个基于 \textit{分布匹配}的通用 RL 框架,允许将模型输出的潜在类别属性分布匹配到指定目标分布。通过实证,我们证明了主导的后训练方案,如群体相对策略优化(GRPO),通过将策略概率集中于单一模式来减少输出多样性。熵正则化和抽样温度可以改善分布的扩散,但其有效性受限,仅限于符号空间和均匀分布。我们证明该领域的先前工作是涉及$L_2$散度的分布匹配特例。随后,我们提出了其他发散如基尔摩尔·拉姆和詹森-香农的奖励函数,并以理论理由进行激励。最后,我们展示了我们方法在涉及数学推理和编程的一组实验中的有效性。
Humanoid Badminton: Learning Dynamic Racket Skills from Limited Human Motion Data
类人羽毛球:从有限的人体运动数据中学习动态球拍技能
- Authors: Jingzhi Cui, Zhexiong Wang, Bangjie Xu, Pengyu Zhao, Youyuan Li, Zhi Su, Peng Ren, Mengdi Xu, Chao Yu, Yi Wu, Luyang Wang, Zhongyu Li
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.31840
- Pdf link: https://arxiv.org/pdf/2609.31840
- Abstract
High-speed racket sports provide a demanding testbed for humanoid robots, requiring time-critical decisions, precise striking, and dynamic whole-body coordination. In badminton, fast-changing shuttle trajectories require timely contact decisions, while successful returns demand precise racket pose and velocity within a brief contact window and across a broad three-dimensional striking workspace. Human motion data provide valuable priors for such athletic skills, but usable badminton references are limited and imperfect. Direct tracking provides insufficient executable variation for diverse shuttle conditions, while purely task-driven optimization may produce unnatural motion. To address these challenges, we present a three-stage hierarchical reinforcement learning framework for dynamic humanoid badminton. First, task-randomized motion augmentation expands sparse annotated hitting events into executable target-conditioned stroke variations, forming a continuous latent skill space. Second, a high-level planner outputs continuous latent skill codes to compose these skills online according to the observed shuttle state. Third, a context-conditioned adversarial regularizer encourages more natural planner-level skill usage while preserving return performance. When deployed on a real humanoid robot, our system achieves sustained multi-skill rallies with human players, including forehand, backhand, and highly dynamic jump returns. This is the first real-world humanoid racket-sport system to demonstrate multi-skill human--robot rallies including highly dynamic jump returns.
- 中文摘要
高速球拍运动为类人机器人提供了严苛的测试平台,要求时间关键的决策、精准的击球和动态的全身协调。在羽毛球中,快速变化的羽毛球轨迹需要及时的接触决策,而成功的回球则需要在短暂接触窗口内、跨越宽广的三维击球工作空间内,保持精准的球拍姿势和速度。人体运动数据为此类运动技能提供了宝贵的先验,但可用的羽毛球参考有限且不完美。直接跟踪在不同折返条件下的可执行变异性不足,而纯任务驱动的优化则可能产生不自然的动作。为应对这些挑战,我们提出了一个三阶段的层级强化学习框架,用于动态类人羽毛球。首先,任务随机动作增强将稀疏注释的击球事件扩展为可执行的目标条件划球变化,形成连续的潜在技能空间。其次,高级规划器输出连续潜在技能代码,根据观察到的穿梭状态在线组合这些技能。第三,情境条件的对抗规则器鼓励更自然的规划者级技能使用,同时保持回球表现。当部署于真实的人形机器人上时,我们的系统实现了持续的多技能回合,包括正手、反手和高度动态的跳回球。这是首个展示多技能人机拉力(包括高度动态跳回球)的真实人形球拍运动系统。
PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning
PHIRL:将学习奖励与任务进展对齐,实现逆向强化学习
- Authors: Hang Yu, James Staley, Cheng Xi Tsou, Xiujin Liu, Wenchang Gao, Jindan Huang, Shijie Fang, Zhegong Shangguan, Angelo Cangelosi, Reuben Aronson, Elaine Short
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31855
- Pdf link: https://arxiv.org/pdf/2609.31855
- Abstract
Human demonstrations provide dense policy-level information but sometimes lack local precision. Human feedback presents accurate local critiques, but offers sparse evaluations rather than direct policy guidance. We propose Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that learns robust reward functions by jointly leveraging demonstrations and feedback. Specifically, we use progress, a feedback modality that describes cumulative task completion. PHIRL iteratively infers a reward function from demonstrations via inverse reinforcement learning, calculates the learned rewards over the progress-annotated demonstrations, and aligns the rewards with progress annotations over four dimensions. We evaluate PHIRL on real and simulated robot tasks, with additional exploration using a fine-tuned vision-language model to provide progress feedback. Results demonstrate that PHIRL significantly outperforms the baselines, achieving substantially higher environmental return rewards and task success with only twenty percent of demonstrations annotated. Analysis of reward-hacking scenarios demonstrates that PHIRL learned reward functions are reliable against exploitation.
- 中文摘要
人工演示提供密集的政策层级信息,但有时缺乏局部精确度。人工反馈提供准确的局部批评,但评估稀疏,而非直接的政策指导。我们提出了进步启发式逆强化学习(PHIRL),这是一个数据高效的框架,通过结合演示和反馈学习稳健的奖励函数。具体来说,我们使用进度,这是一种描述累积任务完成的反馈模式。PHIRL通过逆强化学习迭代从演示中推断奖励函数,计算进度注释演示上的学习奖励,并将奖励与四个维度的进展注释对齐。我们在真实和模拟机器人任务上评估PHIRL,并通过微调的视觉语言模型进一步探索以提供进展反馈。结果表明,PHIRL显著优于基线,仅注释20%的演示,实现了显著更高的环境回报奖励和任务成功率。对奖励黑客场景的分析表明,PHIRL学习到的奖励函数对被利用具有可靠性。
Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining
股票交易的深度强化学习:基于前向再训练的基准分析者-批评者方法
- Authors: Bicheng Wang, Xinyi Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.31870
- Pdf link: https://arxiv.org/pdf/2609.31870
- Abstract
Consistently profitable trading is difficult because equity markets are noisy, non-stationary, and only partially predictable from historical data. We benchmark five deep reinforcement learning (DRL) actor-critic methods: A2C, PPO, DDPG, TD3, and SAC, that learn trading actions end-to-end from market states, and compare them with a supervised price-forecasting baseline. Using daily data for 20 large-capitalization S&P 500 stocks from 2000 to 2020, enriched with trend-following technical indicators and log min-max scaling, we train on 2000-2018 and backtest on 2019-2020. Each agent is evaluated both when trained once and under forward retraining, in which it is retrained on all data available before each successive test window. DDPG achieves the highest annual return (55.5%), Sharpe ratio (1.38), and alpha (0.22), but also the highest market beta (1.24). TD3 and SAC offer a better risk-return balance, with Sharpe ratios of 1.37 and 1.33 and maximum drawdowns of about 25%. Forward retraining improves A2C, PPO, and SAC, leaves TD3 essentially unchanged, and reduces DDPG's annual return from 55.5% to 29.8%, consistent with TD3's greater robustness to hyperparameters. The forecasting baseline has the smallest maximum drawdown (9.6%) and the lowest beta (0.31), underscoring a trade-off between the higher returns of end-to-end DRL and the lower risk of forecast-driven strategies.
- 中文摘要
持续盈利交易很难,因为股市噪声大、非平稳,且仅能部分从历史数据预测。我们以五种深度强化学习(DRL)演员-批判方法为基准:A2C、PPO、DDPG、TD3和SAC,这些方法从市场状态端到端学习交易行为,并与监督价格预测基线进行比较。我们利用2000年至2020年间20只大市值标普500股票的每日数据,结合趋势跟踪技术指标和对数最小最大尺度,2000-2018年训练,2019-2020年回测。每个代理既在一次训练中评估,也在前向再训练中对每个后续测试窗口前的所有可用数据进行再训练。DDPG实现了最高的年回报(55.5%)、Sharpe比率(1.38)和α(0.22),同时也是最高的市场贝塔(1.24)。TD3和SAC提供了更好的风险-回报平衡,Sharpe比率分别为1.37和1.33,最大回撤率约为25%。前期再训练提升了A2C、PPO和SAC,TD3基本保持不变,并将DDPG的年回报从55.5%降至29.8%,这与TD3对超参数更强的鲁棒性相符。预测基线的最大回撤最小(9.6%)和最低的贝塔值(0.31),强调了端到端DRL高回报与预测驱动策略风险较低之间的权衡。
FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization
FARE:深度强化学习,实现公平暴露、受限不确定性意识的金融内容个性化
- Authors: Arundeep Chinta, Lucas Vinh Tran, Jay Katukuri
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2609.31890
- Pdf link: https://arxiv.org/pdf/2609.31890
- Abstract
Content personalization systems in financial services must ensure fair exposure across diverse offerings-a requirement driven by contractual obligations and the need to prevent "rich-get-richer" dynamics where content with high click-through rate (CTR) dominates while other relevant products receive minimal visibility. Share of Voice (SOV) constraints, which guarantee each content category a target fraction of top-position exposure, address this by promoting product diversity and balanced user discovery. While re-ranking layers atop CTR models are common in practice, we propose two key novelties: (1) framing SOV-constrained ranking as a deep reinforcement learning problem analogous to constrained trade execution in algorithmic finance, and (2) explicitly incorporating CTR prediction uncertainty into the agent's state space and policy design-enabling larger ranking adjustments for high-uncertainty predictions where deviation from CTR-optimal ordering is less costly. We introduce FARE (Fair Ranking Executor), a modular uncertainty-aware execution layer that translates any black-box CTR model's predictions into SOV-fair rankings without retraining the underlying model. Our uncertainty-weighted proportional control policy (FARE-PC) and learned neural policies (FARE-ES, FARE-PPO) demonstrate that uncertainty-aware approaches can substantially reduce SOV deviation from fairness targets while minimizing engagement loss, with gradient-free evolution strategies outperforming policy gradient methods on synthetic data and the ordering reversing on KuaiRand-Pure.
- 中文摘要
金融服务中的内容个性化系统必须确保多样化产品间的公平曝光——这一要求源于合同义务以及防止“富人越富”的动态,即高点击率(CTR)内容占主导地位,而其他相关产品曝光度极低。语音份额(SOV)约束,保证每个内容类别在顶级曝光中占目标比例,通过促进产品多样性和平衡用户发现来解决这一问题。虽然在CTR模型上重新排序层级在实践中很常见,但我们提出了两个关键创新:(1)将SOV约束排名框架为类似算法金融中受限交易执行的深度强化学习问题;(2)明确将CTR预测不确定性纳入代理的状态空间和策略设计,从而实现对高不确定性预测的更大排名调整,尤其是在偏离CTR最优排序成本较低的情况下。我们引入了FARE(公平排名执行者),这是一个模块化的不确定性感知执行层,能够将任何黑箱CTR模型的预测转化为SOV公平排名,而无需重新训练底层模型。我们的不确定性加权比例控制策略(FARE-PC)和学习神经策略(FARE-ES、FARE-PPO)证明,不确定性感知方法能够显著减少SOV偏离公平目标,同时最小化参与度损失,无梯度演化策略在合成数据上优于策略梯度方法,KuaiRand-Pure的排序反转表现更佳。
CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense
CyberWorld:样本高效自主网络防御的世界模型
- Authors: Ryozo Masukawa, Sanggeon Yun, Raheeb Hassan, Hyunwoo Oh, SungHeon Jeong, Mohsen Imani
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2609.31893
- Pdf link: https://arxiv.org/pdf/2609.31893
- Abstract
Deep reinforcement learning has become a prominent approach to autonomous cyber defense. Existing methods are predominantly model-free and consequently require extensive environment interaction. World models provide an alternative by learning predictive dynamics and optimizing policies through imagined trajectories, yielding substantial gains in sample efficiency in robotics and embodied control. Extending this paradigm to cybersecurity raises a fundamental question: what should constitute the "world" in a cyber world model? We introduce CyberWorld, a Dreamer-style world modeling framework that learns latent cyber dynamics from vector, graph, textual, and multimodal representations of the defended network. Across all four scoreable CyberWheel attack strategies, the graph-based CyberWorld variant exceeds a strategy-agnostic control after 3.6k-15.8k environment steps, compared with millions of steps required by model-free PPO. Across representation choices, graph structure provides greater robustness under topology-dependent attacks, while simpler representations remain competitive in overall performance. Among successful runs, the number of episodes required to reach the control remains approximately constant as network size increases from 15 to 100 hosts. These results establish learned cyber dynamics as a sample-efficient and scalable basis for autonomous defense, and identify world representation as a central design axis for robustness and scalability.
- 中文摘要
深度强化学习已成为自主网络防御的重要方法。现有方法大多无模型,因此需要大量环境交互。世界模型通过学习预测动力学和通过想象轨迹优化策略,提供了一种替代方案,从而显著提升机器人和具象控制的样本效率。将这一范式推广到网络安全引发了一个根本问题:在网络世界模型中,什么应构成“世界”?我们介绍了CyberWorld,一种Dreamer风格的世界建模框架,通过防御网络的向量、图、文本和多模态表示学习潜在网络动态。在所有四种可评分的CyberWheel攻击策略中,基于图的CyberWorld变体在36k至15800个环境步后超过了策略无关控制,而无模型PPO则需数百万步。在各种表示选择中,图结构在拓扑依赖攻击下提供了更高的鲁棒性,而更简单的表示则在整体性能上保持竞争力。在成功运行中,随着网络规模从15主机增加到100主机,达到控制所需的集数大致保持不变。这些结果确立了学习到的网络动态作为自主防御的样本高效和可扩展基础,并将世界表示视为鲁棒性和可扩展性的核心设计轴。
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
理解 LLM 后期培训中 SFT、RLVR 和 OPD 之间的协同效应
- Authors: Emre Can Acikgoz, Yang Li, Zeyu Leo Liu, Srijan Bansal, Dilek Hakkani-Tür, Shafiq Joty, Semih Yavuz
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31900
- Pdf link: https://arxiv.org/pdf/2609.31900
- Abstract
Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves the current model can make the next stage less effective. Through controlled experiments with Qwen3 models on math and science reasoning, we first characterize OPD across nine student-teacher pairs spanning 2x to 53x parameter ratios and show that OPD effectiveness depends on student-teacher compatibility rather than teacher scale alone. The surrounding stages of OPD reshape this compatibility in three ways: (1) A brief SFT warm-up improves subsequent OPD, while an RLVR-strengthened student regresses under distillation from the same teacher. (2) Adapting the teacher with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Following these two interventions, we find that combining teacher adaptation and student warm-up alone raise average OPD accuracy from 29.2\% to 43.8\% (50\% relative improvement) after the same number of distillation steps, with additional preparatory training. (3) At comparable accuracy, OPD leaves a stronger initialization for downstream RLVR than SFT, with a gap that widens as RL compute scales. Our results suggest that each post-training stage should be chosen not only for the capability it adds, but for the learning interface it creates for the next stage.
- 中文摘要
现代LLM后期培训将监督微调(SFT)、可验证奖励强化学习(RLVR)和策略提炼(OPD)组成多阶段流程,但这些阶段通常单独设计和评估。我们表明这一组合具有重要性:改进当前模型的阶段可能使下一阶段效果降低。通过对数学和科学推理的Qwen3模型进行受控实验,我们首先在9对师生对中表征OPD,这些对比参数为2x至53倍,并表明OPD的有效性依赖于师生兼容性,而非仅教师尺度。OPD的周边阶段以三种方式重塑了这种兼容性:(1)短暂的SFT热身提升后续OPD,而RLVR强化的学生在同一教师的提炼下退步。(2)用RLVR调整教师,使其提升下游OPD的准确性,与其新增能力成比例。在这两项干预措施之后,我们发现,结合教师适应和学生热身,在相同提炼步骤并加额外准备训练后,平均OPD准确率从29.2%提升至43.8%(相对提升50%)。(3)在相当准确性下,OPD对下游RLVR的初始化比SFT更强,且随着强化学习计算规模的扩大,差距会扩大。我们的结果表明,每个培训后阶段不仅应因其新增的能力而选择,还应根据其为下一阶段创建的学习界面而选择。
Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning
利用强化学习增强胸部X光报告生成中的视觉推理能力
- Authors: Denis Musinguzi, Andrew Katumba, Prasenjit Mitra
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.31911
- Pdf link: https://arxiv.org/pdf/2609.31911
- Abstract
Medical report generation has made significant progress with the rise of modern vision-language models and the growing availability of large-scale medical datasets. However, hallucinations remain a major challenge, largely due to the limitations of supervised fine-tuning (SFT), which prioritizes lexical similarity to reference reports rather than clinical correctness. While reinforcement learning has shown strong performance in domains with verifiable rewards such as mathematics and code generation, its application to open-ended medical tasks remains limited. Existing work focuses on evaluating final answers, overlooking the model's reasoning, despite evidence that flawed reasoning can degrade overall performance. In this study, we propose a framework that verifies the model's reasoning process by integrating anatomical regions, bounding boxes, and region-level textual descriptions. We design spatial and factual reward mechanisms to ensure that the model's reasoning is both visually grounded and factually accurate. Starting from Qwen3-VL-8B-Instruct as our base model, we adapt it to the medical domain using supervised fine-tuning, introduce reasoning capability through a cold-start SFT stage, and refine it with reinforcement learning. We find that RL provides performance gains beyond those achievable through SFT alone, and that jointly verifying both reasoning steps and final outputs yields larger improvements than verifying either in isolation. We further identify multiple modes of reward hacking in the RL stage. Finally, the model's structured think traces enhance interpretability, making its outputs easier to audit for clinical use.
- 中文摘要
随着现代视觉语言模型的兴起和大规模医疗数据集的日益普及,医学报告生成取得了显著进展。然而,幻觉仍是一个重大挑战,主要源于监督微调(SFT)的局限性,该技术更注重词汇与参考报告的相似性,而非临床正确性。虽然强化学习在具有可验证奖励的领域(如数学和代码生成)表现出良好表现,但其在开放式医疗任务中的应用仍然有限。现有研究侧重于评估最终答案,忽视模型的推理,尽管有证据表明推理错误会降低整体表现。本研究提出一个框架,通过整合解剖区域、边界框和区域级文本描述来验证模型的推理过程。我们设计空间和事实奖励机制,确保模型推理既具视觉基础又事实准确。以Qwen3-VL-8B-Ininstruction为基础模型,我们通过监督微调将其适应医学领域,通过冷启动SFT阶段引入推理能力,并通过强化学习进行细化。我们发现,强化学习提供了超越单靠SFT可实现的性能提升,且联合验证推理步骤和最终输出带来的改进比单独验证任何一方更为显著。我们还进一步识别了强化学习阶段的多种奖励黑客模式。最后,模型的结构化思维痕迹增强了可解释性,使其输出更易于临床审计。
On-Policy Attention Linearization
政策上的注意力线性化
- Authors: Arian Raje, Anupam Nayak, Anthony Fei, Akaash Parthasarathy, Mohamed Abdelfattah, Gauri Joshi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.31947
- Pdf link: https://arxiv.org/pdf/2609.31947
- Abstract
Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.
- 中文摘要
用线性注意力取代大多数软极大注意力层的混合变换器架构,以极低的内存成本提供变换器级别的质量。越来越多的研究不是预训练此类模型,而是从已训练好的全注意力变换器中提炼出来。然而,这些提炼模型在长上下文检索和推理任务中常常崩溃,尤其是在思考模式下,混合架构的效率提升最为关键。由于线性注意力层必须将上下文压缩为固定大小状态,其错误在长序列中会累积。由于非策略提炼从未教会学生模型从这种漂移中恢复,需要更长序列长度的任务变得尤为具有挑战性。我们引入了策略上注意力线性化(OPAL),其中混合注意力学生采样自身的长上下文轨迹,并接受冻结全注意力教师的密集监督。将OPAL应用于Qwen3-4B和MiMo-7B-RL-0530,我们在常识推理中恢复了87美元至94%美元的全注意力表现,NIAH检索中100美元,仅用3B训练代币实现了83美元至93%美元的数学推理。我们无需监督微调(SFT)或可验证奖励的强化学习(RLVR)即可实现这些结果。与最强的先前线性化方法相比,该方法恢复教师检索表现约68%美元和绝对平均数学推理准确率21.6美元,OPAL则完全恢复了67.6美元——72.2美元。
CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
CaptchaArena:一个大规模、细粒度的数据集,用于在交互式验证码上训练计算机代理
- Authors: Zhenhao Zhang, Zhaoyu Fan, Haohan Ying, Jingwen Hu, Hancen Fan, Junhao Zhou, Zitian Chen, Linchao Zhu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.31957
- Pdf link: https://arxiv.org/pdf/2609.31957
- Abstract
Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at this https URL.
- 中文摘要
⚠️ 翻译失败
Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation
交互式分布稳健多智能体学习,采用一般函数近似
- Authors: Debamita Ghosh, George K. Atia, Yue Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32048
- Pdf link: https://arxiv.org/pdf/2609.32048
- Abstract
Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $\phi$-divergence uncertainty sets. We propose RoMEX-$\phi$, a model-free framework that integrates equilibrium-based exploration with dual fitted learning. Through a functional dual representation of the robust multi-agent Bellman operator, RoMEX-$\phi$ enables tractable worst-case value estimation from nominal interaction data using a centered empirical robust discrepancy. We introduce the robust Multi-Agent Decoupling Coefficient (robust MADC) to characterize the intrinsic exploration complexity arising from strategic interactions and adversarial transition uncertainty. We establish sublinear robust regret guarantees governed by the robust MADC rather than explicitly by the state and joint action space sizes, replacing tabular dependence with intrinsic function-class complexity. Numerical experiments on a scalable general-sum DRMG under total variation uncertainty show that RoMEX-$\phi$ is substantially more resilient to transition shifts than its non-robust counterpart while remaining competitive with an exact tabular robust baseline. Our results provide a scalable framework for distributionally robust multi-agent reinforcement learning with general function approximation.
- 中文摘要
模型错误指定在多智能体强化学习中构成根本性挑战,因为转移不确定性可能通过智能体间的战略互动被放大。分布稳健马尔可夫博弈(DRMG)为应对此类不确定性提供了原则性框架,但现有方法常依赖限制性假设,或在大状态和联合动作空间中扩展性较差。我们在一般函数近似和$\phi$-散度不确定性集的一般和DRMG中研究在线学习。我们提出了RoMEX-$\phi$,一种无模型框架,将基于均衡的探索与对偶拟合学习集成。通过对稳健多智能体Bellman算子的函数对偶表示,RoMEX-$\phi$使得利用中心经验稳健差异从名义交互数据中进行可处理的最坏情况值估计成为可能。我们引入了稳健的多智能体解耦系数(MADC),以表征战略交互和对抗性过渡不确定性带来的内在探索复杂性。我们建立了由稳健MADC而非状态和联合动作空间大小明确控制的亚线性鲁棒遗憾保证,用内在功能类复杂度取代了表式依赖。在全变异不确定性条件下,对可扩展广和DRMG的数值实验显示,RoMEX-$\phi$对过渡位移的韧性远高于非稳健对应,同时保持与精确表格稳健基线的竞争力。我们的结果为分布稳健的多智能体强化学习提供了可扩展的框架,并采用通用函数近似。
Graph Forward Distribution Matching for Molecular Inverse Design
分子逆设计中的图前向分布匹配
- Authors: Yihan Zhu, Yuhan Liu, Brett Savoie, Tengfei Luo, Meng Jiang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32056
- Pdf link: https://arxiv.org/pdf/2609.32056
- Abstract
Achieving precise control over multiple properties without sacrificing chemical validity remains a central challenge in molecular inverse design. Existing reinforcement learning (RL) methods fine-tune graph diffusion models by treating reverse sampling as a sequential policy, using a single terminal reward to optimize hundreds of coupled decisions. They often suffer from instability, validity collapse, and limited property gains. We introduce GraphFDM (Graph Forward Distribution Matching), a new online RL paradigm for graph diffusion that performs optimization through the forward process. GraphFDM uses valid generations to define a reward-tilted target distribution jointly optimized over graph size and molecular structure for each property condition, incorporating reinforcement signals into supervised learning without storing reverse trajectories. We derive the unique optimal target, prove a condition-wise improvement guarantee, and show that the fixed graph-size prior of standard graph diffusion leaves an irreducible matching gap. In multi-conditional polymer and small-molecule generation, GraphFDM achieves the lowest MAE on every target property, with reductions of up to 53.0\% relative to the strongest baselines and chemical validity above 0.99. It further generalizes to out-of-distribution property combinations.
- 中文摘要
在不牺牲化学效度的情况下,实现对多重性质的精确控制仍是分子逆设计的核心挑战。现有强化学习(RL)方法通过将逆向抽样视为顺序策略,利用单一终端奖励优化数百个耦合决策,微调图扩散模型。这些方法常常存在不稳定性、效度崩溃和有限的性质收益。我们引入了GraphFDM(图前向分布匹配),这是一种新的在线强化学习范式,通过前向过程进行优化。GraphFDM利用有效生成定义一个奖励倾斜的目标分布,针对每个属性条件在图大小和分子结构上联合优化,将强化信号纳入监督学习,且不存储反向轨迹。我们推导出唯一的最优靶标,证明了条件层面的改进保证,并证明标准图扩散的固定图大小先验留下不可约匹配缺口。在多条件聚合物和小分子生成中,GraphFDM在每个靶点性质下均可实现最低MAE,相较最强基线降低最多53.0%,化学效度高于0.99。它进一步推广到分布外性质组合。
Grasp2Twist: Learning Bimanual Dexterous Jar Opening by Reinforcement Learning
Grasp2Twist:通过强化学习学习双手灵巧罐开
- Authors: Mo Xu (1), Yunfu Deng (1), Jianuo Wang (2), Josiah Hanna (1), Bilge Mutlu (1) ((1) University of Wisconsin-Madison, (2) Independent Researcher)
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32064
- Pdf link: https://arxiv.org/pdf/2609.32064
- Abstract
This paper presents Grasp2Twist, a bimanual dexterous manipulation system that learns to grasp and twist open jar lids using reinforcement learning. Learning this task raises three challenges: learning a unified policy for a multi-stage task, sustaining lid twisting, and sim-to-real transfer. To address the first challenge, we introduce a continuous enclosure measure to guide grasp formation and a binary enclosure indicator to guide the grasp-to-twist transition for unified policy learning. We derive both from the geometric relationship between the object center and the convex hull formed by the hand's palm and fingertips. Kinematic constraints limit how far the hand can rotate the lid with fixed contacts, so sustained twisting requires finger contact reconfiguration. We use a three-stage curriculum to facilitate exploration of these contact changes and also improve robustness for sim-to-real transfer. With our approach, the learned policy demonstrates finger gaiting, reconfiguring hand-object contacts to sustain lid rotation. It transfers zero-shot to the physical system and achieves an 88% task success rate across six household containers, including peanut-butter, vitamin, and instant-coffee jars. Ablations further validate the roles of the geometric enclosure in grasp formation and the curriculum in contact-reconfiguration exploration.
- 中文摘要
本文介绍了 Grasp2Twist,一种双手灵巧操作系统,通过强化学习学习抓住和扭转打开的罐盖。学习该任务带来了三个挑战:为多阶段任务学习统一策略、维持盖子扭转以及模拟到现实的转移。为应对第一个挑战,我们引入了连续包围度量来指导抓握形成,以及二元围栏指示器用于指导抓握到扭转的过渡,实现统一策略学习。我们两者都源自物体中心与由手掌和指尖形成的凸壳之间的几何关系。运动学约束限制了手在固定接触下旋转盖子的距离,因此持续扭转需要手指接触的重新配置。我们采用三阶段课程,促进对接触变化的探索,并提升模拟到现实转移的鲁棒性。通过我们的方法,所学策略展示了手指步态,重新配置手与物体接触以维持盖子旋转。它将零射程转移到物理系统,并在六个家用容器(包括花生酱罐、维生素罐和速溶咖啡罐)中实现了88%的任务成功率。消融进一步验证了几何围护在握持形成中的作用以及课程在接触重构探索中的作用。
Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models
找到你做不到的事:针对自我改进VLA模型的代理现实强化学习
- Authors: Yuan Fang, Zechu Li, Haolei Tong, Puze Liu, Georgia Chalvatzaki
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32069
- Pdf link: https://arxiv.org/pdf/2609.32069
- Abstract
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $\pi_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: this http URL.
- 中文摘要
视觉-语言-行动(VLA)模型为机器人操作提供了强先验,但通常以冻结策略形式部署,无法从自身失败中改进。现实世界强化学习(RL)提供了持续改进的路径,但手动环境重置和任务成功监督阻碍自主学习。我们介绍 \textbf{FIND},一种能动的现实现实强化学习框架,闭合场景理解、弱点意识练习、自我评估和策略改进之间的循环,持续工作空间中实现。FIND 将自主练习重新框架为场景条件、性能感知的任务选择问题:每次部署后不再恢复预定义场景,而是利用生成场景决定下一步练习内容。视觉语言代理从预定义库中识别可行任务,优先排序近期成功率较低的任务,并通过配对的执行前后观察评估结果。我们以冻结的$\pi_{0.5}$ VLA和残余的非策略强化学习实现FIND。在八个真实世界操作任务中,独立人类评估成功率从$55提升至$71.9%$。代表性运行可在互动后6小时内完成456次自主集,需要30次场景恢复干预,在线学习期间无需人工提供奖励标签。消融和系统评估进一步审视关键设计选择、代理评估准确性和人工干预需求。我们的网站公开访问地址为:此http URL。
When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk
何时人类应重新掌控?在动荡的人工智能风险下的最佳委派
- Authors: Haoze Yan, Julien Roze, Ved Upadhyay, Unal Tatar, Thibaut Mastrolia
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2609.32083
- Pdf link: https://arxiv.org/pdf/2609.32083
- Abstract
Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating periods of elevated risk. We introduce a continuous-time framework for learning adaptive human oversight under such turbulent AI risk. Existing oversight and delegation formulations condition on history but do not model incident clustering, or its suppression by supervision effort, jointly with the delegation decision and this study fixes this gap. The self-exciting dynamics capture how risk events increase the likelihood of subsequent events, making their timing and history central to decision-making. We formulate a stochastic control problem combining human actions, monitoring effort, and switching between human-AI-assisted operation and full AI delegation, balancing operational rewards against oversight costs, and cascading AI-failures and induced uncertainty. Human participation is an endogenous component of risk management: the policy determines both when oversight is needed and how much effort to allocate. We study a relaxed switching formulation and propose Hawkes-PPO, a policy-gradient method that uses a bank of exponential filters of observed incident times. In a synthetic environment it attains a higher risk-adjusted objective than either fixed regime and approaches an approximate full-information oracle. We illustrate our results with numerical simulations by examining how cascade risks influence intervention and delegation, connecting reinforcement learning with adaptive human oversight of AI systems. In particular, we illustrate the benefit of our switching strategy and Hawkes-PPO algorithm to monitor the project efficiently along time, reducing turbulent risks occurrences and costs.
- 中文摘要
部署人工智能系统需要决定何时委派任务,何时人工介入监控和降低人工智能操作引发的风险。当故障聚集时,这些决策具有挑战性:幻觉或有害输出可能触发更多错误,造成风险升高期。我们引入了一个连续时间框架,用于在如此动荡的人工智能风险下学习自适应人类监督。现有的监督和委派表述以历史为基础,但未与委派决策结合对事件聚类或监督抑制进行建模,本研究弥补了这一空白。自我激励动态体现了风险事件如何增加后续事件的可能性,使其时机和历史成为决策的核心。我们提出了一个随机控制问题,结合了人类行动、监控努力以及在人机-AI辅助操作与完全AI委托之间切换,平衡运营回报与监督成本,以及AI失败的连锁反应和诱发的不确定性。人类参与是风险管理的内生组成部分:政策决定何时需要监督以及分配多少精力。我们研究了一种宽松的切换表述,提出了Hawkes-PPO策略梯度方法,该方法使用一组指数级的观察事件时间滤波器。在合成环境中,它获得比任何固定模式更高的风险调整目标,并接近一个近似的全信息预言机。我们通过数值模拟展示结果,考察级联风险如何影响干预和委派,将强化学习与自适应人类对AI系统的监督相结合。特别是,我们展示了切换策略和Hawkes-PPO算法在高效监控项目、降低动荡风险和成本方面的优势。
Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
发挥到标准:可证明的最优四边形块分解的强化学习
- Authors: Arjun Narayanan, Per-Olof Persson
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32146
- Pdf link: https://arxiv.org/pdf/2609.32146
- Abstract
A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of any all-quadrilateral mesh of a given domain purely based on its topology and corner angles. We train a reinforcement learning agent to build decompositions that reach this bound, which we call par. It acts directly on the mesh's half-edge data structure through local edits, with a policy network whose convolutions follow the mesh's own connectivity, so it applies unchanged to domains larger than any seen in training. The reward targets the floor directly, and it is sparse: random play reaches it on no domain with more than eight sides. We overcome this exploration barrier via behaviour cloning on optimal meshes that are trivial to construct, walked backward into demonstrations, before training it with PPO. On 96 held-out domains the agent produces an all-quadrilateral mesh on every one, a usable one on 95.7 on average, and a provably optimal one on 90; Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none, and even at three to fourteen times the elements never produces a more regular mesh. On 64 domains twice the training size the agent completes all, is usable on 62, and keeps a median excess over par below one against Gmsh's 39 at the same element count.
- 中文摘要
平面域的四边形块分解通过其是否完备、元素是否良好以及顶点中有多少不规则来判断。最后一个有可证明的底层:离散高斯-邦内恒等式仅基于其拓扑和角角,强制对给定域中任意全四边形网格的总顶点不规则性设定下界。我们训练强化学习代理构建达到该界限的分解,称之为par。它通过局部编辑直接作用于网格的半边数据结构,策略网络的卷积遵循网格自身的连通性,因此对训练中见到的任何域都不变。奖励直接针对底层,且稀疏:随机对应的域中不超过八边的域。我们通过在最优网格上进行行为克隆克服了这一探索障碍,这些网格易于构建,先回溯演示,然后用PPO训练。在96个保留域上,代理对每个域生成全四边形网格,平均在95.7域生成一个可用网格,在90个域生成一个可证明的最优网格;Gmsh在相同元素数下最强配置完成51个,在38个可用,在零单元上最优,即使元素数达到3到14倍,也从未产生更规律的网格。在64个训练规模的两倍域中,智能体完成所有,62个可用,且中位过剩低于1,而Gmsh在相同元素数下为39个。
Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles
不确定性感知在线算法的模拟集群选择
- Authors: Yongyi Guo, Zifan Xu, Ziping Xu, Kelly W. Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32170
- Pdf link: https://arxiv.org/pdf/2609.32170
- Abstract
The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bootstrap resampling---and selects the online algorithm with the best average performance across the ensemble. While ensemble-based approaches have been used to mitigate distribution shift and facilitate sim-to-real transfer, we formally show that this approach can mitigate the effects of limited data when fitting the simulator and theoretically has significant regret gains compared to Plug-In selection in multi-armed bandits. We also empirically investigate the Uncertainty-Aware selection approach in deep RL experiments on robotic control tasks that involve selecting reward-shaping hyperparameters, and show that it leads to more reliable selection and improved online performance.
- 中文摘要
在线强化学习的性能关键依赖于设计选择,尤其是影响探索的选择。这些选择通常通过拟合模拟器到离线数据、评估模拟器中的候选算法,并部署表现最佳的算法来选择。最简单的插件选择规则仅仅选择拟合模拟器上表现最佳的算法,这使得拟合模拟器的离线数据有限时评估不可靠。我们研究了不确定性感知选择,该选择形成模拟器集合---例如通过引导重采样获得---并选择了整体平均表现最佳的在线算法。虽然基于集合的方法已被用于减轻分布偏移并促进模拟到现实的转移,但我们正式证明,这种方法在拟合模拟器时可以减轻数据有限的影响,并且理论上相较于插件选择在多臂强盗中具有显著的后悔优势。我们还实证研究了深度强化学习实验中对机器人控制任务中选择奖励塑造超参数的不确定性感知选择方法,表明该方法能带来更可靠的选择和提升在线表现。
Noisy Test-Time Reinforcement Learning for Code LLMs
针对代码大型语言模型的噪声测试时强化学习
- Authors: Xikai Yang, Hieu Trung Nguyen, Dunyuan Xu, Yuzhi Zhao, Jinpeng Li, Wenao Ma, Pheng-Ann Heng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32172
- Pdf link: https://arxiv.org/pdf/2609.32172
- Abstract
Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage. Specifically, NTRL-Code uses conservative self-denoising to obtain a cleaner semantic anchor for target estimation, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate programs. The policy is then optimized on the original noisy prompts with a hybrid reward that combines format validity, code similarity, and anti-repetition signals. Extensive experiments on three benchmarks, each incorporating character-level, word-level, and paragraph-level perturbations, demonstrate that NTRL-Code yields robust and consistent improvements, stabilizing the predictions of various base models. Our code is available at this https URL.
- 中文摘要
大型语言模型(LLM)在各种代码相关任务中展现出了卓越的性能。然而,与通常高质量且无错误的精心策划数据集不同,现实世界的用户指令往往模糊且易出错,这对代码LLM的鲁棒性构成了重大挑战。此外,鲁棒性导向的微调依赖于成对的干净-噪声样本,这种样本的策划成本高昂,且需要复杂的噪声模拟技术。为应对这些挑战,我们提出了噪音测试时强化学习框架(NTRL-Code),该框架允许在测试阶段仅使用未标记噪声数据,实现代码LLM的稳健自我进化。具体来说,NTRL-Code采用保守的自去噪技术,获得更清晰的语义锚点以进行目标估计,并采用基于抽象语法树(AST)的结构聚合机制,从多个候选程序中估算代理目标。策略随后对原始噪声提示进行优化,采用结合格式有效性、代码相似性和抗重复信号的混合奖励。在三个基准测试上进行了大量实验,每个基准测试都包含字符级、单词级和段落级扰动,证明NTRL-Code带来了稳健且一致的改进,稳定了各种基模型的预测。我们的代码可在此 https URL 访问。
Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments
见证:互动谜题环境中的发现、解读与顿悟
- Authors: Guanghan Ning, Ping Liu, Linyi Li, Huangjie Zheng, Arjun Neervannan, Huu Nguyen, Michael Sklar, Deniz Zorlu, Nicolai Ouporov
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32208
- Pdf link: https://arxiv.org/pdf/2609.32208
- Abstract
Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rules. An agentic pipeline generates games for WitnessGym, the RL training suite, and WitnessBench, comprising public validation and private test games. The validation set separately tests new compositions of trained rule primitives and primitives absent from training. Under a shared harness, the best of 18 frontier proprietary and open-weight models solves only 24\% of private test level slots, with scores sensitive to the observation interface and agent configuration. Providing ground-truth rules raises Opus-5's validation RHAE-L5 (relative human action efficiency over the first five levels) from 59.9 to 97.8, whereas a 27B open-weight model gains only 2.1 points and remains limited even with the rules provided. RL on WitnessGym raises the 27B model's private test RHAE-L5 from 2.1 to 5.4 and yields a mean gain of 4.1 points on four external discovery benchmarks. Together, these results point to rule acquisition as a major difficulty for frontier models like Opus-5 while smaller models further struggle on rule-based execution, and indicate that RL on hidden-rule puzzles transfers to broader rules and real-world tasks beyond training. Benchmark is available at: this https URL
- 中文摘要
自动化科学需要能够通过与陌生环境互动来推断规则的智能体。交互式规则发现谜题为研究这一能力提供了受控环境:智能体通过实验推断隐藏规则,并利用其推断的内容实现既定目标。我们探究当前语言模型在这些谜题上的限制因素,以及强化学习(RL)是否能提升训练中保留规则的性能。为研究两者,我们引入了WITNESS,一个基于2D网格的谜题环境,具备真实的ASCII观测和规则的受控访问。一个代理流水线为WitnessGym、强化学习训练套件和WitnessBench生成游戏,包括公开验证和私有测试游戏。验证集分别测试训练规则原语的新组合和训练中缺少的原语。在共享框架下,18个前沿专有和开放权重模型中最好的模型仅能解决24%的私有测试层级槽位,且评分对观察界面和代理配置敏感。提供地面真实规则使Opus-5的验证RHAE-L5(前五级相对人类行动效率)从59.9提升到97.8,而27B开放权重模型仅提升2.1分,即使有规则也依然受限。WitnessGym上的强化学习将27B模型的私有测试RHAE-L5从2.1提升至5.4,并在四个外部发现基准测试中平均提升4.1分。这些结果共同表明,规则获取是像Opus-5这样的前沿模型面临的主要难题,而较小模型在基于规则的执行上也面临更大困难,并表明隐藏规则谜题上的强化学习可转移至更广泛的规则和超出训练的现实任务。基准测试可访问:此 https URL
RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process
RoboFFT:通过在线强化学习和前向过程精细化生成机器人政策
- Authors: Yu Li, Shenghe Hu, Yuhan Wang, Yaoxiang Pu, Haotong Zhang, Yuanpei Chen, Yaodong Yang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32236
- Pdf link: https://arxiv.org/pdf/2609.32236
- Abstract
Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, while further improvement typically requires additional expert data. Reinforcement learning offers a natural solution through environment interaction, but effectively finetuning generative robot policies remains challenging due to the intractability of likelihood estimation. In this work, we propose RoboFFT, a forward-process reinforcement learning framework for finetuning generative robot policies, which applies forward noising to sampled actions and uses the weighted score / flow matching loss to construct a surrogate policy ratio for PPO-style updates. We evaluate RoboFFT with popular generative robot policies on representative simulation benchmarks, including long-horizon planning and sparse reward settings. Extensive experiments and analysis demonstrate that RoboFFT consistently improves performance while achieving better stability and training efficiency. We further integrate RoboFFT into a real world RL framework and demonstrate its effectiveness in real world tasks. Project website: this https URL.
- 中文摘要
生成模型,如扩散和基于流的模型,通过从演示中捕捉复杂多模态动作分布,展现出对机器人策略学习的巨大潜力。然而,仅用模仿学习训练的策略常常存在不完美演示和分布偏移的问题,而进一步改进通常需要额外的专家数据。强化学习通过环境交互提供了自然的解决方案,但由于似然估计的难度,有效微调生成机器人策略仍具挑战性。本研究提出RoboFFT,一种用于微调生成机器人策略的前向过程强化学习框架,对采样动作应用前向噪声,并利用加权分数/流量匹配损失构建PPO风格更新的代理策略比。我们结合流行的生成机器人策略在代表性模拟基准上评估RoboFFT,包括长期视野规划和稀疏奖励设置。大量实验和分析表明,RoboFFT在实现更好稳定性和训练效率的同时,持续提升性能。我们进一步将RoboFFT集成到现实世界的强化学习框架中,并展示了其在实际任务中的有效性。项目网站:此链接 https URL。
Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
Train4Merge:一项针对强化学习与SFT教师在基于门诊教学(OPD)模式合并中的受控单一教师研究
- Authors: Jingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li, Xiaoming Zhai, Wei Chu, Ninghao Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32303
- Pdf link: https://arxiv.org/pdf/2609.32303
- Abstract
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinforcement learning (RL). Yet equally strong teachers need not be equally good teachers. We probe this choice through controlled single-teacher OPD, a building block of multi-teacher OPD: in Agentic, Reasoning, and Perception, comparably performing SFT and RL teachers are trained from Qwen3.5-9B, each guiding a student initialized from it. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points in Agentic, Reasoning, and Perception, respectively, and recover more of their teachers' performance gains over the base model. The contrast is clearest in Agentic, where the best SFT-guided student recovers only 44.44% of its teacher's gain, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our analysis points to an explanation: RL teachers stay much closer to the shared initialization in parameter space than SFT teachers and are therefore easier for their students to follow.
- 中文摘要
从共享检查点训练的领域专家可以通过策略提炼(OPD)合并为一个模型,他们作为教师监督学生的自身轨迹。一个上游选择很少被探讨:是用监督微调(SFT)还是强化学习(RL)构建每个专家。但同样强大的教师不必然同样优秀。我们通过受控单教师OPD(多教师OPD的构建模块)来探讨这一选择:在Agentic, Reasoning, and Perception中,表现相似的SFT和RL教师从Qwen3.5-9B出发,分别指导一名初始化的学生。在最佳检查点,强化学习引导的学生在能动性、推理和知觉方面分别比SFT引导学生高出4.27、1.50和0.86个百分点,且教师的表现提升比基础模型还原更多。在能动模式下,这种对比最为明显,表现最好的学生仅恢复了教师的44.44%,而强化学习引导的学生恢复率高达115.00%,超过了其教师。我们的分析指向一个解释:强化学习教师比SFT教师更接近参数空间中的共享初始化,因此学生更容易跟随。
All On-Board: Fully On-Chip Neuromorphic Q-Learning with Embedded CartPole Simulation
全程上载:全片上神经形态Q学习,内置CartPole模拟
- Authors: Steven C. Nesbit, Giovanni T. Michel, Gerd J. Kunde, Edward Kim, Andrew T. Sornborger
- Subjects: Subjects:
Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.32317
- Pdf link: https://arxiv.org/pdf/2609.32317
- Abstract
As AI models grow in size and usage, their energy demands increase dramatically, raising sustainability and economic concerns. Neuromorphic hardware, inspired by the energy efficiency of the brain, seeks to address this challenge by offering low-power, fast-processing alternatives to conventional computing. Such hardware is particularly well-suited to control systems deployed in resource-constrained environments, which are best trained via reinforcement learning (RL). This contribution presents the design and implementation of a fully on-chip, closed-loop Loihi 2 RL agent. Our neuromorphic circuit consists of a fully embedded Q-learning algorithm and an on-chip simulation of the CartPole-v0 environment on Loihi 2. Our Q-learning algorithm trained the same number of successful agents as the CPU implementation in only half the execution time and with two orders of magnitude less dynamic power. These findings demonstrate the viability of RL on neuromorphic hardware and highlight its promise for building energy-efficient, real-time, embedded AI systems.
- 中文摘要
随着AI模型规模和使用率的增长,其能源需求大幅增加,带来了可持续性和经济问题。神经形态硬件受大脑能效启发,旨在通过提供低功耗、快速处理的传统计算替代方案来应对这一挑战。此类硬件特别适合在资源受限环境中部署的控制系统,这些系统最好通过强化学习(RL)进行训练。本贡献提出了一个全片上闭环Loihi 2 RL智能体的设计与实现。我们的神经形态电路由一个完全嵌入式的Q-学习算法和Loihi 2上CartPole-v0环境的片上仿真组成。我们的Q-learning算法在仅有一半的执行时间内训练了与CPU实现相同数量的成功智能体,且动态功耗低了两个数量级。这些发现展示了强化学习在神经形态硬件上的可行性,并凸显了其构建节能实时嵌入式人工智能系统的潜力。
RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning
RLHarness:与强化学习共同演进的过程性技能,用于长视野多模态推理
- Authors: Ziqiao Shang, Zian Xu, Ji-Chen Yan, Weiming Wu, Ziyi Jia, Jie Meng, Tao Huang, Shan Huang, Lan-Zhe Guo
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32326
- Pdf link: https://arxiv.org/pdf/2609.32326
- Abstract
Multimodal reasoning requires models to preserve visual evidence through long decision chains while selecting appropriate procedures across diverse scenarios and rules. When learning is guided only by terminal verifiers, reinforcement learning (RL) reveals whether a final answer is correct but not how it should be produced. The policy must therefore discover reusable reasoning procedures while learning to execute them, creating a program cold-start problem. Skills can externalize successful procedures, reduce repeated exploration, and provide inspectable guidance. However, a fixed Skill Bank assumes that this guidance remains compatible with an evolving policy, while updating Skills alone can leave their triggers, execution protocols, and demonstrations stale or mutually inconsistent. We introduce RLHARNESS, which organizes Skills, selection and execution protocols, few-shot demonstrations, and task contracts into a unified, versioned Harness and alternates Harness evolution with policy learning. An Exploration-Distillation Harness builds the initial Harness and version-aligned verified traces for SFT and DAPO I. After the first RL block, a Post-RL Reconstruction Harness rebuilds Skills, protocols, and demonstrations from fresh success-failure rollouts, and DAPO II adapts the policy to the reconstructed program. RLHARNESS improves Accuracy from 16.25%/27.50% to 62.00%/50.00% on MetroMap/TravelMap and raises F1 score from 37.13%/45.50% to 65.81%/65.51% on Fee-VL/Cancel-VL. All four tasks achieve their best results only after reconstruction and DAPO II, showing that an evolving Harness complements RL by continually updating the external program that the policy learns to execute.
- 中文摘要
多模态推理需要模型通过长决策链保持视觉证据,同时在不同场景和规则中选择合适的程序。当学习仅由终端验证器引导时,强化学习(RL)揭示最终答案是否正确,但不确定答案应如何生成。因此,策略必须在学习执行时发现可复用的推理过程,从而形成程序冷启动问题。技能可以外部化成功过程,减少重复探索,并提供可检查的指导。然而,固定的技能库假设该指导与不断演变的策略兼容,而仅更新技能则可能使其触发条件、执行协议和演示变得陈旧或相互不一致。我们引入RLHARNESS,将技能、选择与执行协议、少数样本演示及任务合同组织成统一的版本化框架,并交替调整框架演进与策略学习。探索-蒸馏工具构建初始框架及与版本对齐的验证轨迹,用于SFT和DAPO I。在第一个强化学习模块后,后强化重建工具框架重建技能、协议和演示,基于新的成功-失败推广,DAPO II则调整策略以适应重建后的程序。RLHARNESS将MetroMap/TravelMap的准确性从16.25%/27.50%提升至62.00%/50.00%,并将F1得分从37.13%/45.50%提升至65.81%/65.51%的Fee-VL/Cancel-VL。这四项任务均在重建和DAPO II后取得最佳效果,表明不断演进的线束通过不断更新策略学习执行的外部程序来补充强化学习。
SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO
SGA-Flow-GRPO:流动GRPO的空间梯度引导学分分配
- Authors: Yunkai Yang, Yudong Zhang, Xinying Chen, Bin Luo, Jienan Lyu, Kunquan Zhang, Weitao Wan, Runmin Dong
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.32340
- Pdf link: https://arxiv.org/pdf/2609.32340
- Abstract
Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both temporal denoising steps and spatial latent dimensions, without explicitly accounting for the spatial structure of generated images, which may lead to sub-optimal policy updates. To address this, we propose a novel gradient-guided spatial credit assignment framework tailored for Diffusion Transformers (DiTs). We first reformulate the transition-level log-likelihood in Flow-GRPO into a token-wise representation natively aligned with DiT patch architectures, constructing spatially fine-grained importance sampling ratios. To allocate localized credit without rigid, boundary-sensitive segmentation heuristics, we introduce a continuous spatial credit map derived from reward gradients. Crucially, we employ an outlier-robust normalization scheme based on Median Absolute Deviation (MAD) coupled with temperature scaling, effectively eliminating gradient noise while highlighting functional prompt-aligned regions. Extensive evaluations on GenEval show that our approach delivers SOTA alignment quality, achieving a convergence rate comparable to top-tier methods like DiffusionNFT while substantially improving upon Flow-GRPO-based methods in alignment performance.
- 中文摘要
强化学习(RL)已被证明能有效将基于流的生成模型与人类偏好对齐。最近,Flow-GRPO通过计算相较于采样候选轨迹的优势,成为一种高效的无批评范式。然而,标准的Flow-GRPO在时间去噪步和空间潜在维度上均采用统一标量优势,未明确考虑生成图像的空间结构,可能导致策略更新不理想。为此,我们提出了一种针对扩散变换器(DiT)量身定制的新型梯度引导空间信用分配框架。我们首先将Flow-GRPO中的过渡级对数似然重新表述为与DiT补丁架构原生对齐的令牌级表示,构建空间细粒度的重要性采样比。为了在没有僵化、边界敏感的切分启发式的情况下分配局部信用,我们引入了基于奖励梯度的连续空间信用图。关键是,我们采用基于中位绝对偏差(MAD)的离群值稳健归一化方案,结合温度尺度,有效消除梯度噪声,同时突出功能性提示对齐区域。GenEval的广泛评估表明,我们的方法实现了SOTA对齐质量,实现了与DiffusionNFT等顶级方法相当的收敛率,同时显著提升了基于Flow-GRPO的方法的比对性能。
OpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRI
OpenMASC:一个用于加速MRI中交叉轨迹金属感知采样与校正的开源流水线
- Authors: Zhengyi Lu, Ming Lu, Chongyu Qu, Junchao Zhu, Junlin Guo, Marilyn Lionts, Yanfan Zhu, Yuechen Yang, Tianyuan Yao, Jayasai Rajagopal, Bennett Allan Landman, Xiao Wang, Xinqiang Yan, Yuankai Huo
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.32343
- Pdf link: https://arxiv.org/pdf/2609.32343
- Abstract
Metal implants corrupt MRI measurements throughout $k$-space, yet existing accelerated MRI methods assume clean data and most metal artifact reduction approaches assume fully sampled acquisitions. No public dataset provides paired $k$-space and images with and without metal for the same anatomy, and no framework jointly addresses artifact-aware acquisition and reconstruction across sampling trajectories. We present OpenMASC, an open-source pipeline covering the full workflow from data generation to deployment. A physics-based data generation module converts public CT volumes into paired clean and metal-corrupted MRI data in both Cartesian and radial formats. MA-VarNet, an unrolled reconstruction network with a per-cascade DC Rectifier, corrects artifacts that data-consistency steps reintroduce from corrupted measurements. A reinforcement learning agent actively selects $k$-space readouts and co-trains with the reconstruction network through a decoupled three-stage procedure. The framework is trajectory-agnostic except for the data-consistency operator, supporting both Cartesian and radial acquisition without architectural changes. Experiments on two datasets at $4\times$ and $8\times$ acceleration demonstrate consistent improvements over conventional and learned baselines on both trajectories.
- 中文摘要
金属植入物会破坏$k美元空间内的MRI测量数据,而现有加速MRI方法假设数据干净,大多数金属伪影减少方法假设采集完全采样。没有公开数据集提供相同解剖结构的配对$k美元空间和有金属和无金属图像,也没有框架能联合处理跨采样轨迹的伪影感知获取和重建。我们介绍OpenMASC,一个涵盖从数据生成到部署全流程的开源流水线。基于物理的数据生成模块将公共CT卷转换为成对的干净和金属损坏MRI数据,格式包括笛卡尔和径向格式。MA-VarNet是一个展开重建网络,配备每级联直流整流器,纠正数据一致性步骤因测量损坏而重新引入的伪影。强化学习代理主动选择$k$空间的读数,并通过解耦的三阶段过程与重建网络协同训练。该框架除数据一致性算符外,轨迹无关,支持笛卡尔和径向采集,无需架构更改。在两个数据集上,加速速度为$4\time$和$8\times$,显示在两种轨迹上均优于传统和学习基线的持续改进。
Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
超越脚本化搜索:通过代理黑箱优化实现样本高效奖励发现
- Authors: Minghao Li, Rui Tan, Ruihang Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32394
- Pdf link: https://arxiv.org/pdf/2609.32394
- Abstract
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.
- 中文摘要
设计低级强化学习(RL)控制的密集奖励函数仍然困难。近期工作利用大型语言模型(LLM)在脚本化搜索算法中通过策略训练反馈迭代生成和优化奖励函数。然而,评估每个候选对象需要完整的强化学习训练运行,这使得样本效率成为复杂控制任务奖励搜索的核心挑战。为解决这一限制,我们提出了一种智能奖励黑箱优化(ARBO)框架,LLM智能体在运行时从其持久工作区维护的评估历史构建搜索策略。评估历史包含两个部分:由评估预言机维护的观察数据,包括候选分数、每项训练曲线和错误追溯;以及智能体维护的信念,记录诊断和预期下一步步骤。智能体通过工具查询并生成下一批奖励候选人,而非一次性从固定提示生成。在四个对照领域中,ARBO在共享评估预算下,操作成功率提升29.9%,电网得分提升192.8%。消融分析各组件对骨干选择的贡献及敏感性。
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
重新思考大型语言模型强化学习中的训练-推理不匹配:其出现之处及纠正方法
- Authors: Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang, Porter Jenkins, Weitong Zhang, Taylor W. Killian
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32444
- Pdf link: https://arxiv.org/pdf/2609.32444
- Abstract
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrepancy in policy updates, we introduce calibrated importance sampling (CIS). CIS is motivated by an empirically supported logit-displacement characterization that expresses the mismatch as an additive displacement $\varepsilon_t$ in log-odds, determined by the per-logit perturbation before the softmax, whose distribution is approximately invariant to token confidence. This characterization motivates a confidence-aware truncation: large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases. Theoretically, we show that CIS replaces the unbounded second moment that governs the error of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess. In evaluation across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines. Diagnostic analyses show that CIS places less truncation bias on low-confidence tokens than truncated importance sampling, while upward clipping of small importance weights reduces held-out accuracy.
- 中文摘要
我们研究了大型语言模型中带可验证奖励的强化学习(RLVR)中的训练-推断不匹配,其中推理引擎采样了推广,而梯度由训练引擎计算,且两者对同一token赋予不同的概率。为解释策略更新中的这种差异,我们引入了校准重要性抽样(CIS)。CIS的动机是基于一种实证支持的logit-位移特征,该表征将错配表示为加法位移$\varepsilon_t$(对数赔率),由软极大值前的每个logit扰动决定,其分布近似不变于令牌置信度。该表征促使置信意识截断:较大的正位移在一个恒定阈值处截断,该阈值映射回重要性比上限,随着令牌置信度增加而趋紧。理论上,我们证明CIS用一个常数有界项替代了决定精确重要性抽样误差的无界第二矩,但代价是被截断剩余部分控制的偏差。在评估三种专家混合模型和五个数学推理基准测试中,CIS在评估基线中均获得了最高的五个基准平均值。诊断分析显示,CIS对低置信度代币的截断偏差比截断重要性抽样更少,而对小重要权重的向上裁剪则降低了被保留的准确性。
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
关于口头置信先验在校准大型推理模型时的陷阱
- Authors: Shuoyuan Wang, Beier Luo, Hao Zeng, Chengyao Yu, Songxin Zhang, Zejian Xie, Bingyi Jing, Hongxin Wei
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32470
- Pdf link: https://arxiv.org/pdf/2609.32470
- Abstract
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distribution, which we term confidence prior. In this work, we reveal that off-the-shelf LRMs exhibit a confidence prior heavily concentrated on a few high values, which persists throughout RL. Theoretically, we prove that this concentration suppresses policy gradient updates for rarely sampled confidence values and inflates the lower bound on expected Brier risk. To overcome this exploration bottleneck, we propose CalibSFT, a plug-and-play supervised fine-tuning stage that shapes a calibrated confidence prior with broad support before RL. For each question, CalibSFT constructs confidence targets combining its success rate with response-level correctness, which provably preserves proper-scoring optimality, and then balances training responses across the confidence spectrum to enable diverse confidence exploration during RL. To learn from incorrect responses without imitating their reasoning, CalibSFT introduces correctness-conditional supervision, guiding confidence across all responses while supervising reasoning only on correct ones. Across 16 mathematical and general reasoning benchmarks, incorporating CalibSFT reduces calibration errors and improves discrimination across five representative RL algorithms while preserving comparable accuracy. Furthermore, CalibSFT delivers practical benefits for downstream selective prediction and model routing. Our code is available at this https URL.
- 中文摘要
大型推理模型(LRM)在表达不确定性时常常存在过度自信的问题。信心感知强化学习(RL)提供了一种有前景的校准优化方法。然而,它依赖于政策内的部署,因此受限于模型的前RL置信分布,我们称之为“置信度”。本研究揭示,现成的LRMs表现出高度集中于少数高值的置信先验,这种先验在整个强化学习过程中持续存在。理论上,我们证明这种集中度抑制了很少采样置信值的政策梯度更新,并使预期Brier风险的下界膨胀。为克服这一探索瓶颈,我们提出了CalibSFT,即插即用的监督微调阶段,能够在RL前形成校准置信先验并获得广泛支持。对于每个问题,CalibSFT 构建了结合成功率与响应水平正确性的置信目标,从而可证明保持适当得分的最优性,然后在置信谱上平衡训练响应,实现在强化学习中多样化的置信探索。为了从错误回答中学习而不模仿其推理,CalibSFT 引入了正确性条件监督,引导所有回答的置信度,同时仅监督正确答案的推理。在 16 个数学和通用推理基准中,CalibSFT 减少了校准误差,提升了五种代表性强化学习算法的辨别能力,同时保持了可比的准确性。此外,CalibSFT 为下游选择性预测和模型路由带来了实际益处。我们的代码可在此 https 网址获取。
AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
AdaTutoRank:通过自适应辅导优化学习如何重新排序文档集,用于RAG和深度研究
- Authors: Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.32472
- Pdf link: https://arxiv.org/pdf/2609.32472
- Abstract
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.
- 中文摘要
文档重排序者决定哪些证据能进入RAG和深度研究的下游模型,而主流重排序者通过相关性匹配选择,且单个相关文档很少构成复杂信息需求所需的完整、互补、非冗余集合。以往的工作通过总量评分奖励集合,将目标从排序文档转向组合集。然而,该评分是集合中每个文档共享的一个标量,因此监督是稀疏的:当该文档得分高时,冗余文档获得其他评分,决定性文档则因评分不佳而被惩罚;署名分配使贡献者与搭便车者无异。政策提炼可以使这种监督变得密实,但现有方法给出的每个推广都给予相同的固定指导,对强力推广过于规范,对弱推广则过于抽象。因此,我们提出AdaTutoRank,这是一个按集合进行重新排序的软件,采用自适应辅导优化(ATO)训练,采用九个评分标准的三级层级结构,提供冷启动的银色标签、强化学习的奖励和提炼的提示。ATO从政策自身的冻结快照中提取三种日益具体的提示形式:仅评分标准、自选者根据评分标准选择的兄弟集,以及自我反思者对比该兄弟集的反映;每个推广项目都获得与其质量匹配的表单。在提示条件的冻结教师和无提示快照下重新评分该推广,将提示效果提炼为代币级优势,补充了群体相对结果优势。在涵盖RAG、深入研究和各项评估的十项基准中,AdaTutoRank在发出更少检索呼叫的同时,取得了最佳的整体表现。
From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning
从结果到策略:学习策略在数学推理中的实用性
- Authors: Ruikang Zhang, Xiao An, Xuli Shen, Jiaxing Sun, Xiaoyi Yu, Jin Zeng, Jiang Wu, Tong Lin
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32482
- Pdf link: https://arxiv.org/pdf/2609.32482
- Abstract
Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.
- 中文摘要
带有可验证奖励的强化学习显著提升了数学推理能力。然而,仅凭终端正确性,对于高层策略(如定理选择和子目标分解)的质量,若将其与后续执行分开考虑,难以提供有限的洞见。本文研究策略效用,即策略在给定执行者下支持正确下游解的可能性。我们介绍SURE,一个学习和利用相对策略效用的框架。在该框架中,高层策略与其详细推理分离。基于策略条件展开和教师生成对比构建的两两偏好,学习策略奖励模型以估计相对策略效用。在强化学习过程中,冻结奖励模型仅读取提取策略,其得分与序列级GRPO目标中的正确性和奖励格式结合。与结果与格式化的GRPO基线相比,实验显示SURE在三个政策骨干中平均pass@1提升了1.87%、2.64%和2.93%。我们的方法还在竞争性或更高的准确性上优于更强的奖励基线,同时所需的GRPO阶段计算量大幅降低。
What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor
ProcGen推广差距衡量什么?动作规则、收敛性与缺失的随机底层
- Authors: Abhisek Keshari
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32532
- Pdf link: https://arxiv.org/pdf/2609.32532
- Abstract
A generalization gap in reinforcement learning (return on training levels minus return on held-out levels) is usually reported without a reference point. We argue that the missing reference is a measured random floor: the return of a uniform-random policy on the same levels under the same harness. On ProcGen, the floor changes what several standard numbers mean. On identical checkpoints and levels across eight environments, switching between sampled and greedy (argmax) test-time actions moves held-out return in both directions, and greedy evaluation takes three environments to or below the floor: in miner, the sampled policy scores 4.9x the floor on held-out levels while its argmax scores below it. Raw policy entropy places seven of eight environments short of convergence, but 35-65% of that entropy is spread across actions with identical effects; after merging them, one to three remain short, and against the floor only heist has learned nothing that transfers. An audit of twelve prior ProcGen codebases finds that all eleven with held-out evaluation sample test-time actions for their policy-gradient agents, nine by default rather than explicit choice, and six report running in-loop averages rather than evaluating a fixed checkpoint. Applied to our own case study, the same checks grade down a statistically significant encoder effect and rule out a within-encoder train-vs-test CKA statistic. We recommend that every reported gap state its action rule, use a matched and seeded protocol, and report the random floor on both level sets.
- 中文摘要
强化学习中的泛化差距(训练水平回报减去保留水平回报)通常没有参考点地报告。我们认为缺失的参考是一个测量随机底线:在同一层级、同一框架下,均匀随机策略的返回。在ProcGen中,底线改变了几个标准数的含义。在八个环境中相同的检查点和层级上,在采样和贪婪(argmax)测试时间动作之间切换,使保留回报双向变动,贪婪评估则使三个环境达到或低于底线:在矿工中,采样策略在保留层级得分为底线的4.9倍,而其ARGMAX得分低于底线。原始策略熵使八个环境中有七个未收敛,但35%-65%的熵分布在具有相同效果的行动中;合并后,一到三个环境仍然短,且只有底线的抢劫未学到任何可转移的知识。对12个先前ProcGen代码库的审计发现,所有11个策略梯度代理的未评估测试时间测试动作,9个是默认选择而非显性选择,6个报告运行循环内平均值而非评估固定检查点。应用于我们的案例研究,同样的检查降低了统计显著的编码器效应,并排除了编码器内部的train-vs-test统计量。我们建议每个报告的差距陈述其动作规则,使用匹配和种子协议,并在两个层级集中报告随机底线。
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Code Agent RL 的分组代理分级与优势再分配
- Authors: Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32577
- Pdf link: https://arxiv.org/pdf/2609.32577
- Abstract
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
- 中文摘要
代码代理的强化学习(RL)通常使用可执行测试来提供二元奖励。基于这些奖励,组相对策略优化(GRPO)为每个推广组内测试通过轨迹赋予相同的优势,忽略实现质量和任务要求遵循度的差异。这使得该策略缺乏一个学习信号,无法优先于干净、有针对性的实现,而非包含不必要或超出范围的变更。我们引入了GAGAR,这是一个用于代码代理RL中质量感知学分再分配的框架。GAGAR基于动态抽样,保留包含通过和未通过轨迹的组,将每个组的所有轨迹置于共享工作区,由受过SFT训练的代理评分员共同检查并对通过测试候选人进行排名。基于该排名,我们对排名较低的轨迹下调权重,并按比例重新调整所有通过测试轨迹的优势,以恢复其原始总和。这种保持和的重分配保留了基于质量的降权所建立的相对权重,同时将功劳向更高质量的实现转移。我们利用MiMo-V2.6-Flash(总参数310B)和MiMo-V2.6-Pro(总参数1.02T)的前RL级SFT检查点,在工业尺度评估GAGAR。受控纯代码Flash实验显示代码代理性能提升,轨迹长度增长减少,训练更稳定。我们进一步将GAGAR应用于大规模混合任务强化学习,同时结合Flash和Pro。我们的结果支持结合基于测试的验证与分组智能体分级,以提升代码代理RL的质量和稳定性。
SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents
SWE-MILE:长视野软件工程代理的异步电位诱导里程碑学分分配
- Authors: Chaoqun Cui, Hao Zhou, Meiqi Chen, Fandong Meng, Wenji Mao
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
- Arxiv link: https://arxiv.org/abs/2609.32631
- Pdf link: https://arxiv.org/pdf/2609.32631
- Abstract
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potential-induced milestone credit assignment framework that derives fine-grained process supervision from workflow runtime, without auxiliary reward models or external evaluators. SWE-MILE quantifies task-relevant file exposure and test-state alignment as navigation and verification potentials, respectively. Differences in these potentials attribute milestone progress and regressions to individual actions, while discounted backward credit propagates supervision to preceding steps. To efficiently acquire intermediate verification states, SWE-MILE further introduces asynchronous shadow probing, which replays repository-changing actions in an isolated sandbox and runs verification in parallel with the agent's primary interaction, largely hiding verification latency. The resulting process credit augments terminal outcome advantages and provides informative learning signals. Experiments on two representative long-horizon SWE tasks demonstrate substantial improvements in agent performance, highlighting workflow runtime signals as a practical source of process supervision for long-horizon SWE agents.
- 中文摘要
采用可验证奖励强化学习(RLVR)训练的长视野软件工程(SWE)代理通常只获得终端结果监督,这使得生产性行动与冗余探索或功能回归难以区分。我们提出了SWE-MILE,一种异步潜在诱导里程碑积分分配框架,从工作流运行时推导出细粒度的过程监督,无需辅助奖励模型或外部评估器。SWE-MILE分别量化任务相关文件暴露和测试状态对齐,作为导航和验证潜能。这些潜能的差异将里程碑进展和回归归于单个操作,而折现后向认可则将监督传递到前一步骤。为了高效获取中间验证状态,SWE-MILE 进一步引入了异步影子探测,该方法在隔离的沙盒中重放存储库变更操作,并在主交互过程中并行运行验证,基本隐藏了验证延迟。由此产生的过程积分增强了终端结果的优势,并提供了有意义的学习信号。对两个具有代表性的长期视野 SWE 任务的实验显示了代理性能的显著提升,凸显了工作流运行时信号作为长期视野 SWE 代理过程监督的实用来源。
PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models
PF-RL:通过目标条件值几何学实现视觉-语言-行动模型的进步场强化学习
- Authors: Yunpeng Qing, Yilun Kong, Sixu Lin, Ming Zhou, Yiming Fei, Shuang Luo, Yixiao Chi, Haoming Gu, Jingyuan Liu, Changxu Wei, Zhi Hou, Changqing Zou
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32634
- Pdf link: https://arxiv.org/pdf/2609.32634
- Abstract
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.
- 中文摘要
强化微调~(RFT)已成为改进视觉-语言-行动~(VLA)策略的有前景范式,但任务层面的稀疏成果对中间过渡,尤其是在长期视野操作中,所提供的信用有限。一种自然的方法是对中间任务进展进行建模,并将其作为策略改进的密集反馈。尽管结构不同,现有的进度感知方法通常将任务进展表述为显式标量预测,提供了有限的结构来建模中间观察与任务目标的关系,这可能阻碍有效的过渡层级功分分配。我们引入了进展场强化学习(PF-RL),该方法通过预训练的VLA特征学习结构化的目标条件进展表示,并将其转化为策略优化的密集学分。一个轻量级共享的进度场头将当前和目标表示映射到紧凑的进度空间中,几何距离诱导目标条件值,而互补的时间和目标结构目标塑造学习后的几何。过渡层级的值变化自然带来密集的进展优势,使离线策略改进和在线强化微调都能实现细粒度的积分分配。对LIBERO、RoboTwin2.0和现实世界双手操作任务的广泛实验表明,PF-RL在强监督微调、强化微调和进度感知基线上持续提升策略性能。
Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL
时间步权权:有效基于ELBO的流量匹配强化学习的隐藏钥匙
- Authors: Qinwei Ma, Jingzhe Shi, Simin Fan, Ling Li, Mengdi Wang, Alex Lamb
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.32665
- Pdf link: https://arxiv.org/pdf/2609.32665
- Abstract
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
- 中文摘要
基于ELBO的强化学习提供了一种与采样器无关的方法,用于微调带有奖励反馈的流量匹配模型。基于ELBO的RL中时间步权权对性能影响巨大,并且它提供了统一视角(正如本研究所示),帮助理解先前工作中以启发式方式选择的预测损失,但该方法仍研究不足,且常被选用来继承预训练配置。我们研究基于ELBO的RL中时间步加权的影响和动态。我们表明有效加权取决于奖励环境和学习阶段。(1)通过对受控CIFAR图像生成的实验,结合机器人技术,我们研究加权如何影响各噪声级别的奖励驱动更新。(2)通过梯度分析,我们揭示了任务间交叉噪声协调的明显模式及其在训练过程中的演变。这些发现支持了有用权重依赖于策略当前行为与奖励偏好行为之间的差距的假设。(3)在该分析的指导下,我们研究了简单静态加权、预算配置选择和动态计划,这些方法能提升超越传统目标选择的性能。我们的结果确立了时间步权权作为流程匹配强化学习的重要设计选择,并激励进一步研究如何在学习过程中选择并调整时间步长加权。
Scaling Properties of Same-Family On-Policy Distillation
同族策略蒸馏的缩放性质
- Authors: Yuntai Bao, Qinfeng Li, Guoqing Jiang, Liwei Chen, Zhiheng Qin, Xuanping Li, Wenqi Zhang, Xuhong Zhang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32722
- Pdf link: https://arxiv.org/pdf/2609.32722
- Abstract
Reinforcement learning (RL) can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular useful-transfer regime, in which held-out accuracy (the gold score, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit power laws for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
- 中文摘要
强化学习(RL)可以在大型语言模型(LLMs)中诱导显著的推理能力,但这些能力在不同模型尺度间转移的程度以及转移的速度尚不明确。我们研究了策略上提纯(OPD)在弱到强、同基和强到弱师生设置下的尺度性质。我们发现早期OPD训练动态一致地呈现出规律的有用转移状态,其中保留的准确率(金分数,$G$)在$d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(即token级反向KL发散与学生初始化的平方根)中大致线性上升。在观察到的每对弱到强的组合中,学生的峰值金分数都高于其教师,因此紧凑型强化学习专家可以通过门诊能力转移给更强学生。为了估计门诊结果,我们对$G_{\mathrm{peak}}$和有用转移机制尺度的斜率拟合幂律,以学生和教师参数计数及教师金评分为基础。这些定律表明,峰值金分数在教师量表下仅大致提升到学生的量表,且在匹配金分数下,较小的教师转移效果更好,因此教师评分本身并不决定其监督价值。我们还研究了两种OPD变体的尺度效应:弱到强OPD的自助法,以及政策指导的程度。
Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
分布稳健平均奖励强化学习:弱交流下的有限样本保证
- Authors: Chenyu Lu, Zijun Chen, Nian Si
- Subjects: Subjects:
Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.32727
- Pdf link: https://arxiv.org/pdf/2609.32727
- Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for learning an $\epsilon$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_{\delta}^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_{\delta}^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.
- 中文摘要
我们研究在弱通信下的平均奖励环境下的分布鲁棒强化学习(DR-RL)。我们的主要结果为估计稳健最优平均奖励和学习近似最优策略提供了有限样本保证,涵盖SA矩形和S矩形结构,并具有基于散度和距离的不确定性集。具体来说,对于Kullback-Leibler和$f_k$-散度球,我们建立了显式半径条件,使得稳健平均奖励Bellman方程具有常数增益解,而对于全变差和Wasserstein球,任意正半径即可,而无需名义MDP弱传递。我们的算法是无先验知识的,样本复杂度为$\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$用于估算稳健的最佳平均奖励和$\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta} }^{\ast})\epsilon^{-2})$ 用于学习 $\epsilon$-最优策略。这里,$p_{\wedge}$ 是最小的正名义转移概率,$u_{\delta}^{\ast}$ 是一个稳健的最优偏置函数。我们还给出了 $\operatorname{Span}(u_{\delta}^{\ast})$ 的几乎紧密显式上界。最后,我们通过数值实验验证预测的$n^{-1/2}$收敛率。
Retrospective Distillation Attribution via Normalized Response Similarity
通过归一化反应相似度进行回溯性蒸馏归因
- Authors: Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang, Jungseul Ok
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.32749
- Pdf link: https://arxiv.org/pdf/2609.32749
- Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring syntactic patterns into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated syntactic signatures along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
- 中文摘要
模型蒸馏通过对教师回答进行监督微调(SFT)转移能力,通常从商业API收集,这引发了模型来源性的问题。现有的蒸馏归因方法大多在SFT步骤后立即对学生进行评估。然而,提炼模型在发布前可能还需进一步SFT、偏好优化或强化学习,而审计员可能无法访问基于引用的归因所需的预蒸馏检查点。为弥合这一差距,我们提出了SCOUT方法,这是一种仅输出的方法,将重复出现的句法模式聚合为候选人轮廓,过滤低对比度模式,并校准学生-候选者距离与候选人间距离。SCOUT支持仅使用当前文本的归因和省略,不使用模型权重、符号似然或历史检查点。SCOUT审计了涵盖多种培训后目标的公开发布的蒸馏模型后代,持续识别蒸馏源。此外,沿培训轨迹追踪教师相关的句法特征,发现它们在蒸馏过程中出现,并在后续偏好优化和强化学习中持续存在。
CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
CUA沙盒:计算机使用代理强化学习的高效环境
- Authors: Xin Yan, Zhengbo Jiao, Jiaqi Liu, Zhenglin Wan, SiYuan Ma, Xuliang Yu, Tianyi Jiang, Chubin Zhang, Pengfei Zhou, Wangbo Zhao, Xingrui Yu, Bo An, Yang You, Ivor Tsang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32750
- Pdf link: https://arxiv.org/pdf/2609.32750
- Abstract
Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications. However, conventional deployments replicate an initialized runtime for each independent rollout, even when trajectories use the same software, incurring repeated memory and initialization costs as the number of parallel environments grows. Does an independent computer-use environment require an independent execution runtime? Our key observation is that trajectories require independent mutable state, while initialized application runtimes can be reused across concurrently evolving environments, making state the natural unit of environment independence. Guided by this observation, we introduce CUA-Sandbox, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators. Experiments show comparable or improved task success relative to Docker, while substantially reducing rollout and resource costs. CUA-Sandbox achieves up to a 6.20x increase in rollout throughput, a 9.2x reduction in per-environment memory, and a 504x reduction in incremental storage.
- 中文摘要
强化学习使计算机使用代理通过与真实软件环境(包括网站和桌面应用)的交互来改进。然而,传统部署每次独立部署都会复制初始化运行时,即使轨迹使用相同软件,随着并行环境数量增加,仍会产生重复内存和初始化成本。独立的计算机使用环境是否需要独立的执行时?我们的关键观察是,轨迹需要独立且可变的状态,而初始化后的应用运行时可以在同时演进的环境中重用,使状态成为环境独立的自然单位。基于这一观察,我们引入了CUA-Sandbox,通过状态范围执行和事务生命周期操作(包括重置和分支)将私有状态胶囊与共享运行时分离,同时保留原始软件接口和任务评估器。实验显示任务成功率与Docker相当甚至更好,同时显著降低了部署和资源成本。CUA-Sandbox实现了最多6.20倍的扩展吞吐量提升,每环境内存减少9.2倍,增量存储减少504倍。
CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models
CT-OPD:扩散视觉语言模型中的反事实追踪策略提炼
- Authors: Long Qian, Bingke Zhu, Jiaqi Wei, Yu Li, Yingying Chen, Jinqiao Wang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32781
- Pdf link: https://arxiv.org/pdf/2609.32781
- Abstract
Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal decisions. Its trajectories capture these decisions, yet their provisional visible tokens can conflict with the target response. Outcome-based reinforcement learning follows these trajectories but provides only response-level feedback, which loses contrast when sampled rewards tie. To align coherent token-level supervision with the model's reveal decisions, we introduce Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student. CT-OPD retokenizes each teacher response in the student's vocabulary and extracts unresolved-position masks at successive stages of the student's reverse process. For each mask, it discards provisional rollout values and reconstructs the partial state from the teacher endpoint, so the supervised positions follow the current trajectory while the visible context and targets remain consistent with the same response. The student is trained on these reconstructed states with its native categorical loss, and trajectories are refreshed as the model evolves. Across dense and sparse diffusion architectures, CT-OPD consistently enhances multimodal understanding and reasoning capabilities, with gains of up to 9.80 points on the nine-benchmark average. On the unified understanding-and-generation architecture, it also improves both visual understanding and image generation, showing that the same principle transfers across architectures and modalities. Ablations further attribute these gains to coherent reconstruction and current-model trajectory masks.
- 中文摘要
扩散视觉语言模型通过逐步解析掩蔽的标记生成答案,使部分解决状态下的准确条件预测成为训练后的核心。掩蔽已完成的答案可产生连贯的上下文和目标,但规定的掩码不反映模型的揭示决策。其轨迹捕捉了这些决策,但其暂时可见标记可能与目标反应冲突。基于结果的强化学习遵循这些轨迹,但仅提供响应级反馈,当抽样奖励相匹配时,反馈会失去对比。为了使连贯的标记级监督与模型揭示决策保持一致,我们引入了反事实追踪策略蒸馏(CT-OPD),将完成的教师回答与当前学生的轨迹掩码结合起来。CT-OPD在学生词汇中重新标记每个教师的反应,并在学生逆向过程的后续阶段提取未解决的位置掩体。对于每个掩码,它丢弃了临时的展开值,重建教师端点的部分状态,使监督位置遵循当前轨迹,而可见上下文和目标保持一致。学生在这些重构状态上接受训练,并伴随着其原生的类别损失,轨迹随着模型演进而更新。在密集和稀疏扩散架构中,CT-OPD持续提升多模态理解和推理能力,较九个基准点平均提升最高9.80分。在统一理解与生成架构上,它还提升了视觉理解和图像生成,表明同一原理在不同架构和模态间可迁移。消融进一步将这些收益归因于相干重建和当前模型轨迹掩膜。
Understanding and Exploiting Anisotropy in Post-Training
理解和利用后期培训中的各向异性
- Authors: Samyak Jha, Harshvardhan Saini, Yizhen Liao, Yiming Tang, Dianbo Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.32792
- Pdf link: https://arxiv.org/pdf/2609.32792
- Abstract
LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large activations. Anisotropy is widely documented and usually treated as a defect, yet its function and its interaction with post-training remain unclear. We first analyze it. A label-free outlier rule isolates about 5\% of residual channels that are essential for language modeling: removing them raises perplexity from 10 to over $10^6$, versus 35 for count-matched random channels. Yet they barely distinguish correct from incorrect reasoning. SFT reshapes them, whereas RL leaves them largely intact and adapts the complementary channels. These channels therefore form the model's \emph{coherence substrate}, and reasoning adaptation happens elsewhere. We then exploit this. \textsc{SphereGate} learns one bounded gain per residual channel on a frozen backbone. Its activation-weighted gradients provably limit movement of high-energy coherence channels and leave the remaining channels free. With 0.1M trainable parameters, \textsc{SphereGate} outperforms parameter-efficient baselines by 2.0--7.3 points on MATH-500 across Qwen2.5 (0.5B--7B) and Llama-3-8B, is comparable or exceeds full-model GRPO. Anisotropy is not a defect but a division of labor that post-training can exploit.
- 中文摘要
LLM后训练结合了监督微调(SFT),一种覆盖模式的前向基准目标,和强化学习(RL),一种寻求模式的反向KL目标。频率加权似然训练留下一个众所周知的特征:\emph{各向异性},其中少数残余通道携带不成比例的激活。各向异性被广泛记录,通常被视为缺陷,但其功能及其与后训练的相互作用仍不明确。我们首先分析它。一个无标签离群值规则分离出约5/%的残余通道,这些通道对语言建模至关重要:移除它们使困惑度从10%提升到超过10^6美元,而计数匹配的随机通道则为35。然而它们几乎无法区分正确与错误推理。SFT重塑它们,而强化学习则基本保持不变,并适应互补通道。因此,这些通道构成模型的\emph{相干基底},推理适应则在其他地方进行。然后我们利用这一点。\textsc{SphereGate}在冻结主链上每个残差通道学习一个有界增益。其激活加权梯度可证明限制高能相干通道的移动,并保持剩余通道自由。在0.1M可训练参数下,\textsc{SphereGate}在QWEN2.5(0.5B--7B)和Llama-3-8B上比参数效率基线高出2.0-7.3分,与全模型GRPO相当甚至超过。各向异性不是缺陷,而是后训练可以利用的分工。
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的现实基准
- Authors: Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32810
- Pdf link: https://arxiv.org/pdf/2609.32810
- Abstract
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
- 中文摘要
多学科肿瘤委员会通过专家讨论整合多模态临床观察和纵向患者病史,但基准很少能捕捉这些真实世界的轨迹。我们引入了OpenTumorBoard,这是一个包含611个患者病例和19,157次讨论轮次的基准,涵盖十个专家角色,内容来自YouTube上12,534分钟的公开肿瘤委员会录像。基准评估两种环境:专家转向,即大型语言模型在真实讨论中回答临床重要问题;以及董事会模拟,生成完整的来回讨论,达成治疗建议、手术计划、下一步行动和临床试验匹配的共识。对14个通用前沿和医学LLM的评估揭示了显著局限性:最佳模型在临床等效性上得分为3.43分(满分5分),与记录的委员会结论一致为2.78分(满分5分)。监督微调和强化学习提升了在测试集中的表现,表明现实讨论轨迹可以支持模型适应。三位医学博士专家回顾了基准的一个子集,发现患者病例信息覆盖率高且事实性强,且提取的共识结论高度准确。我们将发布OpenTumorBoard及其自动化策展流程,以支持LLMs的开发和评估,以支持多学科、个性化癌症决策。
Improving LLM Collaboration via Multi-Agent Preference Learning
通过多智能体偏好学习提升LLM协作
- Authors: Shuo Liu, Xinzichen Li, Tianle Chen, Christopher Amato
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.32827
- Pdf link: https://arxiv.org/pdf/2609.32827
- Abstract
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI feedback. Yet, its extension to multi-agent systems remains underexplored. To address this gap, we formulate preference-based multi-agent systems (MAS) from decentralized and centralized collaboration perspectives. We also introduce a general multi-agent preference learning framework (MAPL) to solve these problems. MAPL allows iterative updates by comparing the current solution with decentralized or centralized solutions generated by various agents. We instantiate MAPL using MARL from human feedback (MARLHF) with a learned reward model and multi-agent direct preference optimization (MADPO). Experiments on collaborative writing, coding, tool use, and travel planning show that MAPL can improve collaboration quality and efficiency while approaching the performance of MARL with fixed, well-defined rewards. Within MAPL, MARLHF generally outperforms MADPO on most tasks but remains sensitive to data coverage, agent and comparator models, and the underlying MARL algorithms.
- 中文摘要
已有多项研究探讨了大型语言模型协作中的多智能体强化学习(MARL)。然而,构建可靠奖励在实际中较为困难,因为完整且准确的指标往往不可得且难以汇总。偏好学习通过比较人类或人工智能反馈提供了一种替代方案。然而,其向多智能体系统的扩展仍未被充分探索。为弥补这一空白,我们从去中心化和集中式协作视角提出了基于偏好的多智能体系统(MAS)。我们还引入了通用的多智能体偏好学习框架(MAPL)以解决这些问题。MAPL通过将当前解决方案与不同智能体生成的去中心化或集中式解决方案进行比较,实现迭代更新。我们通过学习奖励模型和多智能体直接偏好优化(MADPO)使用人类反馈(MARLHF)实现MARL实例化MAPL。关于协作写作、编码、工具使用和出差规划的实验表明,MAPL能够提升协作质量和效率,同时以固定且明确的奖励接近MARL的性能。在MAPL中,MARLHF在大多数任务上通常优于MADPO,但仍对数据覆盖率、代理和比较器模型以及底层MARL算法保持敏感。
Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping
重新定时的钟声流:逃离速度自助的不可能三角
- Authors: Boyang Xu, Shengzhe Chen, Hao Yan
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32828
- Pdf link: https://arxiv.org/pdf/2609.32828
- Abstract
Flow critics learn return distributions by transporting Gaussian noise to Bellman endpoints via continuous velocity fields. While velocity bootstrapping stabilizes training by querying a successor teacher, existing methods face a structural dilemma: on straight paths, no residual-free same-time affine mapping can preserve Gaussian initial noise while maintaining an unbiased target. To overcome this limitation, we introduce Retimed Bellman Flows (ReBF). ReBF queries the teacher critic at a dynamically shifted earlier flow time, aligning intermediate student and teacher trajectories. By combining this retimed clock with fresh, decoupled noise generation, ReBF constructs a provably conditionally unbiased velocity target that preserves the Bellman fixed point and contracts under Wasserstein distances. Empirically, ReBF reduces $W_1$ distance to ground-truth return distributions by up to $7.7\times$ on synthetic MRPs and outperforms existing flow critics across 38 challenging OGBench and D4RL offline reinforcement learning tasks.
- 中文摘要
流批评者通过连续速度场将高斯噪声传输到贝尔曼端点来学习返回分布。虽然速度自举通过查询继任教师稳定训练,但现有方法面临结构性难题:在直线路径上,没有残差自由同时间仿射映射能在保持无偏目标的同时保持高斯初始噪声。为克服这一限制,我们引入了重新时序贝尔曼流(ReBF)。ReBF在动态偏移的早期流时间查询教师批评者,对齐中间学生和教师轨迹。通过将重新计时的时钟与全新解耦噪声生成结合,ReBF构建了一个可证明的条件无偏速度目标,保持贝尔曼不动点,并在瓦瑟斯坦距离下收缩。实证上,ReBF在合成MRP$W上将地面真实返回分布的距离缩短了最多7.7美元,并且在38个具有挑战性的OGBench和D4RL离线强化学习任务中,表现优于现有流程批评者。
Allspark: Weak to Strong Transfer via Alternating Chain of Thought
全火种:通过交替思维链由弱到强转移
- Authors: Kaizhao Liang, Junxiong Wang, Chen Liang, Zhendong Wang, Qiang Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.32913
- Pdf link: https://arxiv.org/pdf/2609.32913
- Abstract
Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. We introduce Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought. A weak teacher is trained alongside a frozen copy of the same model; the two alternate reasoning segments, and the frozen model produces the final answer. At inference time, a stronger student replaces the frozen training partner, while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. We study Allspark at two scales: controlled Qwen experiments across math and reasoning, and larger-scale Inkling experiments on ARC-AGI-2. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi and Nemotron, with benefits that vary across inference settings. These findings motivate reusing a trained weak teacher across strong students and examining the resulting accuracy--token tradeoff.
- 中文摘要
前沿模型的最新进展重新激发了对大规模强化学习(RL)的兴趣,但生成大型模型展开的成本使得即使是测试强化学习配方的成本也高昂。我们探讨,一个小型、弱模型所学到的推理改进是否能在不使用强模型的展开的情况下,惠及更大、更强的模型。我们介绍了Allspark,一种通过交替思维链实现弱到强转移的训练与推理框架。弱教师与同一模型的冻结副本一起训练;两个交替推理片段,冻结模型产生最终答案。在推理时,一个更强的学生替代被冻结的训练伙伴,而两个模型保持固定。由于它们通过文本交流,教师可以引导来自不同模型家族和不同标记器的学生。我们在两个尺度上研究Allspark:一是数学和推理领域的对照Qwen实验,另一是ARC-AGI-2上的更大规模Inkling实验。Inkling实验显示,在家庭内和跨家庭环境中,包括转移到Kimi和Nemotron,且各推理场景带来的益处不同。这些发现促使人们在强学生中重复使用受过训练的弱教师,并检验所得的准确性——象征性权衡。
Counting on Thinking: Tracing Evidence Integration in Language Models
依靠思维:在语言模型中追踪证据整合
- Authors: Jingming Xue, Robert C. Wilson, Huadong Xiong
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32932
- Pdf link: https://arxiv.org/pdf/2609.32932
- Abstract
Finite computational resources force a tradeoff between automatic System 1 processes and costly System 2 thinking. Large language models (LLMs) can spend extra computation on hard problems, yet direct answers struggle even with counting, an elementary operation humans and animals perform automatically. We ask why this requires thinking in LLMs. Evidence integration has long been used in psychology and neuroscience to probe decision-making. Our evidence-integration task presents one letter per conversational turn and asks which of two target letters appeared more often. A running count difference solves the task optimally by weighting every letter equally; tokens at each turn could represent and update this difference. Direct responses instead weighted evidence unevenly, with strong recency effects, and assigned less probability to the correct answer as difficulty increased. Thinking improved performance and made integration weights nearly uniform, yet final-query attention remained concentrated on the sequence ends in both modes. Reasoning trajectories showed models revisiting input, recounting letters, and checking intermediate counts that informed the answer, suggesting that thinking constructs the accumulated count that direct responses lack rather than reading out one already formed. Reasoning-token costs grew with the number of letters far more than with coherence. Outcome feedback did not bring this computation into direct responses: under in-context reinforcement learning (ICRL), performance deteriorated over repeated games and recency effects strengthened, yet models grew more confident. Humans and animals amortize such computations into automatic processes, whereas current LLMs still pay for them with thinking on every trial. Which operations learning can make directly available remains central to how future models allocate computation.
- 中文摘要
有限的计算资源迫使系统1自动处理与昂贵的系统2思维之间做出权衡。大型语言模型(LLMs)可以在难题上投入额外计算,但直接回答即使在计数(人类和动物自动执行的基本操作)时也难以回答。我们提出问题:为什么这需要LLMs的思考。证据整合长期以来一直被用于心理学和神经科学中探究决策。我们的证据整合任务每回合对话中出现一个字母,询问两个目标字母中哪个出现频率更高。通过均衡权重每个字母,进行持续计数差异以最优方式解决任务;每个回合的标记可以表示并更新这一差异。直接回答则不均匀地加权证据,产生强烈的近代效应,且随着难度增加,正确答案的概率降低。思维提升了性能,使积分权重几乎均匀,但最终查询的注意力仍集中在两种模式的序列端端。推理轨迹显示模型会重新审视输入、重新计数字母,并检查中间计数以形成答案,表明思维构建的是直接回答所缺乏的累积计数,而非读出已形成的计数。推理代币成本随着字母数量的增加而增加,远大于连贯性。结果反馈未能将计算直接应用于回应:在上下文强化学习(ICRL)下,性能随着重复游戏下降,近期效应增强,但模型却增强了信心。人类和动物将此类计算分摊为自动过程,而当前的大型语言模型仍通过每次试验的思考来支付这些计算。运算学习能直接提供哪些计算,仍是未来模型计算分配的核心。
Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs
结构化受限MDP中在线强化学习的最后迭代保证
- Authors: Nam Phuong Tran, Trinh Ha Mai Huynh, Tuyen Pham Le, Van-Truong Nguyen, Quan Nguyen, Long Tran-Thanh
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32933
- Pdf link: https://arxiv.org/pdf/2609.32933
- Abstract
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.
- 中文摘要
在安全关键应用中,部署使用单一策略,其性能和约束满足应直接成立,而非仅对平均或混合训练策略保持。这激励了约束强化学习中的最后迭代保证。近期进展已在精确梯度或表格在线环境中建立了此类保证,但结构化大状态问题的可扩展结果仍未确定。我们开发了一个通用且统计高效的结构约束MDP(CMDP)末次迭代收敛框架。我们的分析将正则化原始对偶动力学的收缩与参与者近似及在线探索策略评估中的统计误差区分开来。这使得无模型的on-off策略学习和结构化函数近似成为可能:乐观策略评估避免显式转换模型构建,而紧凑参数演员避免维持过去策略的混合或历史。我们实现了线性CMDP和一般函数近似的框架,获得了与表示相关的复杂度和比以往积极正则化原始对偶分析更优的目标准确性依赖性。我们进一步验证了我们理论预测的合成线性CMDP的稳定效应:正则化方法表现出稳定的最后迭代行为,而其非正则化方法则表现出更大的振荡。
Communication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task Allocation
用于去中心化多人多机器人任务分配的通信感知异构图学习
- Authors: Ziqin Yuan, Ruiqi Wang, Baijian Yang, Byung-Cheol Min
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32935
- Pdf link: https://arxiv.org/pdf/2609.32935
- Abstract
Multi-human multi-robot (MH-MR) teams combine robotic autonomy with human expertise, but effective task allocation requires coordinating scarce, dynamically available human support with distributed robot execution. Limited robot-robot and human-robot communication further complicates this coupling by delaying information exchange and supervisory intervention. We introduce CommHG, a communication-aware heterogeneous graph learning framework for decentralized MH-MR task allocation. CommHG represents coupled human-robot-task interactions through local graphs conditioned on information availability and age. Learned communication actions enable robots to decide when and which operator to query, and when to share information with peers. Allocation and communication are jointly optimized through cooperative multi-agent reinforcement learning, allowing the team to acquire useful information while managing limited communication and supervisory resources. We also introduce a benchmark integrating heterogeneous humans and robots, dynamic tasks and operational states, and constrained robot-robot and bidirectional human-robot links. Experiments across heterogeneous teams with up to 16 robots, 6 humans, and 112 tasks show that CommHG improves timely weighted mission completion as coordination scale increases, outperforming the strongest baseline in the Large scenario.
- 中文摘要
多人多机器人(MH-MR)团队结合了机器人自主性和人类专业知识,但有效的任务分配需要协调稀缺且动态可用的人类支持与分布式机器人执行。有限的机器人-机器人和人-机器人通信进一步加剧了这种耦合,延迟了信息交换和监督干预。我们介绍了CommHG,一种用于去中心化MH-MR任务分配的通信感知异构图学习框架。CommHG通过基于信息可用性和年龄的局部图表示人机任务的耦合交互。学习到的通信动作使机器人能够决定何时、何时向哪个操作员查询,以及何时与同行共享信息。分配和通信通过协作多智能体强化学习共同优化,使团队在管理有限的通信和监督资源的同时获取有用信息。我们还引入了一个整合异构人机、动态任务与操作状态,以及受限机器人与双向人机连接的基准测试。跨异构团队(最多16台机器人、6名人类和112项任务)的实验显示,随着协调规模扩大,CommHG提升了加权任务的及时完成率,超过了大型场景中最强基线。
Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View
受限流策略更新:广义薛定谔桥视图
- Authors: Boyang Li, Matthew Kim, Sylvia Herbert
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.32952
- Pdf link: https://arxiv.org/pdf/2609.32952
- Abstract
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schrödinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.
- 中文摘要
在线安全强化学习(RL)寻求在满足安全约束的同时最大化奖励的策略。奖励和安全性可以诱导多模态动作分布,挑战了主流的原始对偶方法:高斯演员可能崩溃到单一次最优模式,且对非凸拉格朗日景观的优化可能不稳定。扩散策略和流策略可以表示此类分布,但近期扩散演员的工作依赖于估计并匹配增强拉格朗日目标策略的得分。相反,我们通过流量策略的生成路径直接区分增强目标,因此无需估计分数。由于流策略缺乏用于熵正则化的现成动作对数密度,我们基于FLAC的无密度动能正则化法(FLAC)这一近期的奖励方法,提出了重参数化增强拉格朗日流演员(RAFALE),这是一种非策略的actor-critic方法,用于安全强化学习。我们将其更新表述为一个受约束单端广义薛定谔桥,并证明对于每个源绘制,该路径空间问题正好是作用空间中的熵正则化问题。在正噪声下,其解仅在估计成本超过拉格朗日乘数阈值时重新加权仅奖励动作分布。随着噪声消失,最优值收敛到流策略直接优化的最低能量映射目标值。在七个安全体育馆任务中,RAFALE在每个任务中均以平均最终成本在预算内实现了竞争性奖励,而强基线则是相互交换;消融支持其增强目标和流动行为者的必要性。
Self-Confirming Superposition Traps in Reinforcement Learning
强化学习中的自我确认叠加陷阱
- Authors: Dai Shi, Andi Han, Feng Chen, Yiqun Duan, Junbin Gao, José Miguel Hernández-Lobato
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.32966
- Pdf link: https://arxiv.org/pdf/2609.32966
- Abstract
Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap, every optimal code assigns overlapping directions to features that rarely occur together under the current policy. An alternative action brings them together, causing interference that lowers its return and reinforces avoidance, although refitting to that action would yield more return at the same capacity. We characterize the dimensions admitting a trap in a tied two-step model and show separately that equal feature frequencies, continued visitation, and independent controller learning need not prevent it. Because fitting weights errors by visitation, an avoided action can lose its return advantage at little cost to the objective. In a finite-action model, we bound this distortion and derive a replay condition: sufficient training weight on the best separately adapted action preserves its ranking despite residual error. Neural PPO experiments show how the feedback develops during learning: agents initialized toward different actions develop different interference patterns, opposite mean return rankings, and different final policies at the same capacity. We therefore test whether retaining access to neglected states can improve control. Keeping these states in training reduces measured interference and improves sequential return, with gains even when the encoder is frozen. Related interventions on state access, replay weights, and feature overlap improve control on MiniGrid and DMControl. For agents that learn through a world model, protected fitting improves DreamerV3--Crafter's cumulative training scores at unchanged capacity.
- 中文摘要
强化学习(RL)训练由智能体策略选择的数据表示,智能体随后利用所得回报指导下一步选择。我们证明即使表征拟合在这些数据上是全局最优的,该循环仍能维持较低的回报策略。在自我确认叠加陷阱中,每个最优代码都为在当前策略下很少同时出现的特征分配重叠方向。另一种动作将它们聚集在一起,造成干扰降低返回并强化避免,尽管重新调整到该动作会在相同容量下获得更多回报。我们刻画了平局两步模型中允许陷阱的维度,并分别证明相等特征频率、持续访问和独立控制者学习不必阻止陷阱。由于通过访问来拟合错误,避免的动作可能会以极小的成本失去回报优势。在有限动作模型中,我们对这种扭曲进行了界限,并推导出一个重放条件:对最佳独立调整动作的足够训练权重,即使存在残余误差,其排名依然保持。神经PPO实验展示了反馈在学习过程中的发展:初始化到不同动作的代理在相同容量下发展出不同的干扰模式、相反的平均回报排名和不同的最终策略。因此,我们测试保留对被忽视状态的访问是否能改善控制。保持这些状态在训练中减少测量干扰,改善顺序返回,即使编码器冻结时也有提升。相关干预措施包括状态访问、重放权重和特征重叠,有助于提升MiniGrid和DMControl的控制。对于通过世界模型学习的代理,保护拟合提升DreamerV3——Crafter在容量不变时的累计训练分数。
The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration
模型还懂另一种方式:策略切换以实现有效RLVR探索
- Authors: Jin Cui, Xinyue Long, Boran Zhao, Pengju Ren, Hao Dong
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33085
- Pdf link: https://arxiv.org/pdf/2609.33085
- Abstract
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
- 中文摘要
带有可验证奖励的强化学习(RLVR)通常受限于探索不足:困难问题可能导致均匀错误的展开组,从而减少学习信号。我们证明此类失败不必反映能力缺失。相反,有限抽样通常专注于特定问题的主导推理策略,而保留模型已支持的替代策略未被探索。此外,这些策略的可及性在强化学习过程中不断变化:有些策略内化为自主行为,另一些则难以引发,难以被吸收。基于这些观察,我们引入了问题-策略展开分配(PSRA),将无引导和策略条件提示视为竞争的探索臂,利用贝叶斯顺序分配将固定的推广预算引导至最有可能产生信息丰富、非饱和群体的臂。保留目标保持有用的策略条件路径可访问,同时成功引导行为转移至无引导策略。在Qwen2.5模型中,从1.5B到7B,加上两个强化学习训练语料库,PSRA持续提升推理性能,减少死饱和,强化分布外转移,并在推理预算增加下保持更大收益。
Evolving Dexterous Robots from Scratch
从零开始进化灵巧机器人
- Authors: Zihan Guo, Shuzhe Zhang, Muhan Li, Peiyang Li, Sam Kriegman
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33101
- Pdf link: https://arxiv.org/pdf/2609.33101
- Abstract
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
- 中文摘要
关于如何手动设计具备灵巧操作能力的智能体,知之甚少。通过仔细观察动物如何操控物体,推断出了一些设计原则,但这些结构和行为迄今为止对仿生学具有抵抗力,可能不适合人工机器。在这里,我们进化出自由形态机器人,能够拾取、握持、旋转和使用各种物体。与其他优化机器人手的方法不同,我们不假设身体任何部位的存在、关节或几何形状。虽然熟悉的抓握形态如尾巴、喙、爪子和爪子在某些条件下可能自发出现——虽然这些条件可能引起进化生物学家的兴趣——但全新操作手设计也能揭示全新的解决方案,这些结构可能更适合当前任务。我们利用对比学习创建高度可搜索的设计空间遗传嵌入,利用自回归发展模型解码设计,进化策略寻找良好设计,以及强化学习训练每个进化设计。获胜设计自动转换为可制造蓝图,在现实世界中以零样本方式打印、组装和测试。这些结果代表了进化机器人学在性能、多样性和复杂性方面的顶尖水平。
MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
MedRouter:解析基于路由推理的医学大型语言模型知识差异
- Authors: Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33119
- Pdf link: https://arxiv.org/pdf/2609.33119
- Abstract
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any single model. However, existing LLM routing methods primarily seek to balance answer quality and inference cost, leaving open how to exploit differences in specialist competence to improve medical reasoning. In this paper, we introduce MedRouter, an agentic system that uses an embedding-based multi-label router to select and query specialist LLMs, then passes their responses to a generator to produce the final answer. We further propose SCALE (Specialist Competence-Aware Learning), a two-stage training framework that first trains the Router with specialist correctness supervision and then optimizes its selections through reinforcement learning. The second stage uses a Performance Gain Reward (PGR) that measures how specialist information affects the generator's answer correctness relative to answering without that information. Experiments on eight text-based and multimodal medical QA benchmarks show that MedRouter outperforms the strongest routing baseline by 8% in average accuracy. Our analysis of specialist outputs further reveals distinct strengths and complementary question-level coverage, motivating learned routing to combine these capabilities for more comprehensive medical reasoning.
- 中文摘要
医学问答涵盖了多种专业和模式,单个医学大型语言模型(LLM)在任务和领域中展现出明显优势。这种异质性表明,结合专家或许能比依赖单一模型更广泛地覆盖医学问题。然而,现有的LLM路由方法主要致力于平衡答案质量和推理成本,留下了利用专家能力差异提升医学推理能力的开放空间。本文介绍了MedRouter,一种基于嵌入的多标签路由器的代理系统,利用其选择和查询专业LLM,然后将其回答传递给生成器生成最终答案。我们还进一步提出了SCALE(专家能力感知学习),这是一个两阶段训练框架,首先在专业正确性监督下训练路由器,然后通过强化学习优化其选择。第二阶段使用绩效提升奖励(PGR),衡量专家信息如何影响生成器的答案正确性,相对于无该信息的回答。对八个基于文本和多模态的医疗质量保证基准测试的实验显示,MedRouter在平均准确率上比最强的路由基线高出8%。我们对专家输出的分析进一步揭示了其独特优势和互补的问题层级覆盖,激励学习到的路由结合这些能力,实现更全面的医学推理。
Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface
一起训练还是以后合并?通过共享动作界面统一VLA专家
- Authors: Zhizhen Zhang, Yuxia Fu, Zijian Wang, Helen Huang, Yadan Luo
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.33125
- Pdf link: https://arxiv.org/pdf/2609.33125
- Abstract
Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining independently trained experts through model merging is a natural approach, yet strong individual experts do not necessarily yield a strong merged policy. We identify one source of this incompatibility: task-specific changes to the action interface, comprising action normalization and the action encoder and decoder. We propose PolicyWeave, combining merge-compatible post-training with context-guided sparse merging. During post-training, all experts retain the common base policy's action interface, while task adaptation is restricted to LoRA updates in the hidden layers of the action model. This makes the experts more compatible with existing model merging methods. However, merging all experts can still introduce interference from unrelated tasks at deployment. PolicyWeave scores each expert's LoRA updates using the initial visual-language context, determines the expert set through leave-one-layer-out ranking stability, and forms a sparse weighted merge of the selected updates that remains fixed for current task. We evaluate PolicyWeave with GR00T N1.5 on 18 RoboCasa365 tasks, using only 10% of the target-task demonstrations for supervised fine-tuning (SFT). Preserving the shared action interface raises the average success rate across four static merging methods from 17.0% to 52.8%. PolicyWeave achieves 64.7% success with these SFT experts and 74.1% after task-specific reinforcement learning (RL), compared with 60.7% for joint RL. Further evaluations on LIBERO-10 and an AgileX Piper arm support the deployment of independently learned skills in long-horizon and real-world manipulation.
- 中文摘要
协同训练提供了一种直接构建多任务视觉-语言-动作(VLA)策略的方法,但可能无法达到独立训练每个任务所获得的性能。挑战在于如何在多任务策略中保留这些任务专属的提升,而无需联合后期训练。通过模型合并组合独立训练的专家是自然的做法,但强大的个别专家并不一定能带来强有力的合并策略。我们指出了这种不兼容的一个原因:任务特定的动作界面变更,包括动作规范化以及动作编码器和解码器。我们提出PolicyWeave,将合并兼容的后训练与上下文引导的稀疏合并结合结合。在后训练阶段,所有专家保留共同基础策略的动作界面,而任务适配仅限于动作模型隐藏层中的LoRA更新。这使得专家与现有模型合并方法更兼容。然而,合并所有专家仍可能在部署时引入无关任务的干扰。PolicyWeave 利用初始可视化语言上下文对每位专家的 LoRA 更新进行评分,通过保留一层的排名稳定性确定专家集,并形成一个稀疏加权的合并,保持当前任务的固定状态。我们用 GR00T N1.5 在 18 个 RoboCasa365 任务中评估了 PolicyWeave,仅使用了目标任务演示的 10% 进行监督微调(SFT)。保留共享动作接口使四种静态合并方法的平均成功率从 17.0% 提高到 52.8%。PolicyWeave 在这些 SFT 专家中成功率为 64.7%,任务特定强化学习(RL)后成功率为 74.1%,而联合强化学习为 60.7%。对 LIBERO-10 和 AgileX Piper 臂的进一步评估支持了独立学习技能在长期和现实操作中的应用。
Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
保存你饱和的数据:在基于群体的强化学习中超越奖励饱和
- Authors: Ziyuan Yang, Yike Wang, Shangbin Feng, Yulia Tsvetkov
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33126
- Pdf link: https://arxiv.org/pdf/2609.33126
- Abstract
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, where the group-relative learning signals vanish and leave previously useful data obsolete. In this work, we investigate whether useful learning signals can be recovered from such saturated data. We study interventions at four levels of group-policy RL pipelines---data, rollout, reward, and advantage---and conduct extensive RL training on saturated reasoning data only. While standard GRPO on saturated data would almost always yield near-0 advantages and near-noise signals, diverse interventions successfully recycle and repurpose such data: among the proposed strategies, interventions at rollout generation are consistently most effective: nudging the policy to generate ``high-quality'', incorrect solutions introduces rollouts with poor rewards into saturated groups as negative samples, which turns out to improve GRPO by 6.4% to 9.0% across Qwen3-1.7B and 4B. Other interventions such as increasing rollout temperature or adding auxiliary rewards can also restore non-zero advantages, but yield less consistent gains. Further analyses show that effective negative rollouts require informative negative trajectories, that the method remains effective alongside unsaturated data, and that it supports iterative recycling of newly saturated examples. While increasingly stronger LLMs would render more data as saturated, our results demonstrate that don't waste your saturated data: with the right strategies they can be recycled into useful RL training signals in an increasingly data-scarce world.
- 中文摘要
群体相对强化学习(RL)依赖抽样反应之间的奖励变异来估算信息性相对优势。随着语言模型能力的提升,现有训练数据可能变得奖励饱和:同一问题的所有抽样反应可能获得同等高的奖励,而群体相对学习信号消失,使得先前有用的数据变得过时。本研究中,我们探讨是否可以从此类饱和数据中恢复有用的学习信号。我们研究了四个层级的群体政策强化学习流程干预---数据、推广、奖励和优势---并仅对饱和推理数据进行大量强化学习训练。虽然标准GRPO在饱和数据上几乎总是带来接近0的优势和近乎噪声信号,但多样化的干预措施能够成功回收和再利用此类数据:在提出的策略中,推广生成时的干预始终最为有效:推动政策生成“高质量”但错误的解决方案,将奖励较差的推广引入饱和群体为负样本,这实际上在Qwen3-1.7B和4B中使GRPO提升了6.4%至9.0%。其他干预措施如提高推广温度或增加辅助奖励,也能恢复非零优势,但收益不那么一致。进一步分析显示,有效的负向推广需要信息性的负轨迹,该方法在未饱和数据下依然有效,并支持对新饱和样本的迭代循环利用。虽然越来越强大的大型语言模型会让更多数据变得过饱和,但我们的结果表明,不要浪费饱和数据:通过正确的策略,它们可以被回收利用,成为在日益稀缺数据的世界中有用的强化学习训练信号。
Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation
政策可塑性在离线到线上强化学习中至关重要:为在线适应调整线下政策
- Authors: Yuheng Huang, Yunpeng Qing, Yixiao Chi, Yilun Kong, Changqing Zou
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33127
- Pdf link: https://arxiv.org/pdf/2609.33127
- Abstract
Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while generally treating the offline-trained policy as a given initialization. We instead study O2O adaptation from the perspective of network plasticity, asking whether the offline-trained policy remains sufficiently adaptable for online learning. Controlled experiments show that prolonged optimization on static offline data progressively reduces network plasticity even after offline performance has largely saturated, and that lower plasticity is associated with weaker subsequent online improvement. Motivated by these observations, we propose REstoring plasticity via Fresh Initialization and policy Transfer (REFIT), a lightweight model-level method for the O2O transition. Before online fine-tuning, REFIT distills the offline policy into a freshly initialized student while temporarily freezing a random subset of student units, transferring the learned offline behavior to a more plastic policy initialization. Extensive experiments on D4RL and OGBench demonstrate that REFIT consistently achieves higher aggregate performance than existing O2O plug-in methods across both Cal-QL and IQL backbones, while plasticity diagnostics and ablations provide further evidence of restored network plasticity.
- 中文摘要
离线到在线强化学习(O2O RL)已成为一种实用范式,利用静态离线数据集预训练策略,随后通过在线交互调整策略。现有的O2O方法主要通过价值校准来处理过渡,通常将离线训练策略视为给定初始化。我们转而从网络可塑性的角度研究O2O适应,探讨离线训练策略是否足够适应在线学习。受控实验表明,即使离线性能基本饱和,长期对静态离线数据的优化会逐步降低网络可塑性,且较低可塑性与后续在线改进较弱相关。基于这些观察,我们提出通过Fresh initialization和策略转移(REFIT)恢复可塑性,这是一种轻量级模型级O2O过渡方法。在在线微调之前,REFIT将离线策略提炼到新初始化的学生中,同时暂时冻结随机学生单元子集,将学习的离线行为转移到更具可塑性的策略初始化。在D4RL和OGBench上的大量实验表明,REFIT在Cal-QL和IQL骨干链上始终比现有O2O插件方法实现更高的总体性能,而可塑性诊断和消融则进一步证明了网络可塑性恢复。
Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?
任务的天花板:变压器什么时候能在没有思维链的情况下取得成功?
- Authors: Jiashu He, Jinxuan Fan, Xiao Xiao, Radu Marculescu, Alejandro Ribeiro
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33134
- Pdf link: https://arxiv.org/pdf/2609.33134
- Abstract
Reasoning models generate long chains of thought before they answer, yet it is debated whether the content of these chains does real computational work or is largely decorative. We study this question by viewing a transformer as a shallow circuit. One forward pass through a fixed number of layers has constant depth, so any procedure that runs the model a constant number of times is a shallow circuit. We call the best accuracy that a shallow circuit can reach on a task the ceiling of the task, and a task is serial if its ceiling lies below one. We prove three results on serial tasks that hold for every transformer, no matter how it was trained. Necessity: replacing the chain by anything that does not depend on its content, such as filler tokens or a restatement of the question, drives the accuracy down to the ceiling, and on a maximally serial task down to chance. Depth: no shallow computation can write the chain of a model whose accuracy exceeds the ceiling, not even approximately. Locality: the answer is one shallow pass away from the finished chain, so all of the serial reasoning happens in the chain. On word problems of finite groups, whose ceilings are known, small transformers trained from scratch, with or without reinforcement learning, attain the predicted numbers: chain-trained models solve every input length and fall to chance when the chain is erased, chainless models collapse to the ceiling as the input length grows, and open-weight reasoning models given the same problem in words return to the baseline without their chain. On MATH-500 and AIME, erasing the chain costs open reasoning models 0.52 to 0.82 accuracy, a sentence shuffle is harmless, and a token shuffle is as harmful as erasing; the same holds for checkpoints trained by GRPO with a correct or a random reward. The ceiling of a task therefore answers when a transformer can succeed without its chain of thought.
- 中文摘要
推理模型在回答前会产生长串思考链,但关于这些链的内容是否真正进行计算工作还是主要装饰性存在争议。我们通过将变压器视为浅层电路来研究这个问题。一次前向通过固定层数的路径深度恒定,因此任何运行模型恒定次数的过程都是浅电路。我们称浅电路在任务中能达到的最佳精度为任务上限,任务的上限低于1则称为串行。我们证明了三个串行任务的结果,无论其如何训练,都适用于每个变换器。必要性:用任何不依赖其内容的东西替代链,比如填充标记或问题的重述,会将准确率降低到上限,而最大串行任务则降低到偶然。深度:没有任何浅层计算能写出准确度超过上限的模型链,甚至近似也不行。局部性:答案是距离完成链仅一个浅层路径,因此所有串行推理都发生在链中。在有限群的文字题中,天花板已知,从零训练的小变换器(无论是否进行强化学习)都能获得预测数值:链训练模型解决所有输入长度,链被抹除时随机,无链模型随着输入长度增加崩溃至天花板,开权推理模型在词中遇到相同问题时会返回基线,链条不带链。在MATH-500和AIME中,擦除链条的开逻辑模型精度降低0.52到0.82,句子洗牌无害,令牌洗牌与擦除同样有害;同样适用于由GRPO训练的检查点,且有正确或随机奖励。因此,任务的上限决定了变压器何时能在没有思维链的情况下取得成功。
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
测试时间的缩放和训练在个人站姿预测中有哪些不足?
- Authors: Yuyang Zhao, Xuan Liu, HaoYang Shangm Haojian Jin
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33155
- Pdf link: https://arxiv.org/pdf/2609.33155
- Abstract
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately.
- 中文摘要
测试时间尺度和后训练提升了LLM在编码和数学推理方面的表现,但它们对个体姿态预测的有效性仍不明确。我们通过预测一个人在新讨论中的立场,从其历史中研究该问题。我们评估了广泛使用的测试时尺度策略和训练后方法,如监督微调和强化学习,识别出生成、选择和学习中的四种失败模式:(1)错误共识,重复样本一致认同错误姿态;(2)选择失败,生成覆盖观察姿态但选择未及;(3)反应过拟合,监督微调提升模仿但损害预测;(4)早期平台期,强化学习显示初步获益适度,随后改进有限。我们使用STANCE-BENCH暴露这些失败,该系统包含来自500名Hacker News用户的2499个预测任务。在本次分析的指导下,我们探索了一种简单方法,将所有候选人立场的直接评分与个人历史支持的明确评估相结合。在781任务测试集中,该方法在Qwen3-8B中获得了21.83的讨论专项宏观F1,而直接评分为19.27。我们的结果激励我们分别评估候选人生成、最终选择和个体特定证据的使用。
Downstream-Aware Context Selection for Online In-Context Reinforcement Learning
在线上下文强化学习的下游感知上下文选择
- Authors: Ruihan A. Li, Shangtong Zhang, Rohan Chandra
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33166
- Pdf link: https://arxiv.org/pdf/2609.33166
- Abstract
In-context reinforcement learning (ICRL) enables large language model agents to adapt to new environments using their interaction history without updating model parameters. However, repeatedly conditioning on growing histories can lead to substantial token cost. We propose a bounded-history context-management framework that predicts the task-dependent downstream effect of removing historical interactions to guide history selection and determine a decision-dependent context budget. Formally, our framework uses the full rolling history as a reference. The predictor evaluates removal effects, defines a deletion ordering, and applies a shared selection criterion to determine how much history to retain at each decision. We evaluate the method in closed-loop SUMO driving under held-out in-distribution, unseen-domain, and unseen-route settings, and in ScienceWorld under a continual ICRL protocol. Relative to a baseline using the full context, our method reduces total token usage by 25.7%, 25.8%, and 23.2% across the three driving settings while maintaining comparable closed-loop driving performance. In ScienceWorld, it reduces total token usage by 52.1% compared to full context and uses 30.2% and 37.8% fewer tokens than the Recent and Similarity baselines, respectively, while maintaining performance.
- 中文摘要
上下文强化学习(ICRL)使大型语言模型代理能够利用交互历史适应新环境,而无需更新模型参数。然而,反复条件于增长的历史可能导致显著的令牌成本。我们提出了一个有界历史上下文管理框架,预测移除历史交互以指导历史选择并确定决策依赖上下文预算的任务依赖性下游效应。形式上,我们的框架以完整的滚动历史作为参考。预测器评估去除效应,定义删除顺序,并应用共享选择准则来决定每次决策应保留多少历史。我们在封闭循环SUMO驱动下,在未显示的分布、未见域和未见路径环境中评估该方法,并在ScienceWorld中采用持续ICRL协议。相较于全上下文基线,我们的方法在三种驾驶设置中分别减少了25.7%、25.8%和23.2%的总代币使用,同时保持了相当的闭环驾驶性能。在ScienceWorld中,它将总代币使用量减少了52.1%,且分别比近期和相似度基线少用30.2%和37.8%代币,同时保持性能。
When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
模型何时承认自己错误?在强化学习下,失败披露是不稳定的
- Authors: Steven Y. Feng, Noah D. Goodman, Michael C. Frank, Evan Hubinger, Paul C. Bogdan, Andrew Lampinen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33220
- Pdf link: https://arxiv.org/pdf/2609.33220
- Abstract
Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy. The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in an instruction-conditioned 32B setting. We also find that small floating-point and sampling differences during training can redirect reporting behavior even when the task objective and earlier training history are held fixed. Additional tests show that failure disclosure is not a single decision: Checking the answer, entering a report, and completing the admission can separate, and the weak point depends on the task and response format. Further, experiments with neutral controls show more broadly that behaviors left weakly constrained by training are especially likely to vary across runs, of which failure disclosure is an example. We can reduce variability in failure disclosure by discouraging the model from drifting from its starting policy on failed, well-formed responses. This makes reporting substantially more consistent, though its effect on task performance depends on the setting. Stable task accuracy therefore does not guarantee stable safety-relevant behavior: Researchers should measure these behaviors directly across runs and design training methods that keep them reliable.
- 中文摘要
基于结果的强化学习可以生成任务表现相似但沟通错误方式截然不同的模型。我们研究失败披露:模型是否承认尝试的解决方案失败,而不是保持沉默或呈现为成功。在重复的仅结果GRPO训练中,失败披露的差异远大于任务准确率。该模式扩展到第二推理任务和稳定的PPO,持续在7B,并在指令条件的32B环境中出现。我们还发现,即使在任务目标和早期训练历史固定的情况下,训练过程中微小的浮点和抽样差异也能重新引导报告行为。其他测试显示,失败披露并非单一决策:检查答案、录入报告和完成承认可以分开,弱点取决于任务和响应格式。此外,中性对照实验更广泛地表明,训练限制较弱的行为在不同运行间尤为可能变化,失败披露即是一个例子。我们可以通过阻止模型偏离其对失败且良好响应的起始策略来减少失败披露的变异性。这使报告更加一致,尽管其对任务表现的影响取决于具体环境。因此,稳定的任务准确性并不保证安全相关行为的稳定:研究人员应直接在各运行间测量这些行为,并设计保持其可靠性的训练方法。
RMB: Reward Model Boosting Mitigates Reward Hacking
RMB:奖励模式提升缓解奖励黑客行为
- Authors: Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33221
- Pdf link: https://arxiv.org/pdf/2609.33221
- Abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.
- 中文摘要
人类反馈强化学习(RLHF)是一种强大的技术,用于将大型语言模型(LLM)与人类偏好对齐。然而,它常常存在奖励黑客问题,即策略优化会改善代理奖励模型,但由于代理的不完美,实际上会降低真实人类偏好的性能。为此,我们提出了奖励模型提升(RMB)这一新颖方法,旨在增强RLHF奖励信号的鲁棒性和可靠性。RMB首先用促进多样性的正则化器训练一组奖励模型。这鼓励每个模型学习奖励景观中的互补方面。然后,RMB学习一个基于提升原理的轻量级聚合器,将多样化奖励模型的输出聚合成更准确、更稳健的奖励信号。我们广泛的实验表明,RMB显著提升了分布内和非分布数据集的奖励准确性,显著缓解了奖励黑客问题,最终提升了RLHF性能。
CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents
CodeSkill:长期代码代理的潜在技能抽象
- Authors: Song-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan, Weiwen Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33243
- Pdf link: https://arxiv.org/pdf/2609.33243
- Abstract
Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads to inefficient exploration and weak credit assignment under sparse rewards. Moreover, while large-scale agent trajectories often contain recurring multi-step behavioral patterns, their noisy token-level representations hinder effective experience reuse. To address these challenges, we propose CodeSkill, a framework that adapts hierarchical latent skill modeling to the code agent domain. CodeSkill first leverages a teacher model to distill both successful and failed trajectories into multi-level textual abstractions. It then integrates temporal variational inference with reinforcement learning to map these discrete semantics into continuous latent variables, while an adaptive boundary mechanism dynamically gates skill transitions based on execution feedback. The learned skills are injected into a frozen LLM policy as latent semantic prefixes, enabling optimization in a compact semantic space rather than over raw token sequences. By shifting RL from token-level exploration to experience-level reasoning, CodeSkill improves optimization efficiency and long-horizon behavioral coherence. Extensive experiments demonstrate that CodeSkill achieves highly competitive performance against strong open-weight baselines across diverse general and industrial coding benchmarks. Furthermore, the learned skills exhibit strong transferability and robust cross-domain generalization, highlighting the effectiveness of explicit behavioral abstraction for scalable agentic code generation.
- 中文摘要
代码代理需要在复杂交互轨迹上进行长期决策。然而,现有的强化学习(RL)方法通常在令牌层面优化行为,导致低层生成与高层行为推理之间存在不匹配。这一局限导致探索效率低下,在奖励稀疏下信用分配较弱。此外,尽管大规模代理轨迹常包含重复的多步行为模式,但其嘈杂的令牌级表示阻碍了有效的经验再利用。为应对这些挑战,我们提出了CodeSkill框架,该框架将层级潜在技能建模适配到代码代理领域。CodeSkill首先利用教师模型,将成功与失败的轨迹提炼为多层文本抽象。随后,它将时间变分推断与强化学习相结合,将这些离散语义映射为连续的潜在变量,同时通过自适应边界机制动态地基于执行反馈对技能转换进行门禁。所学技能作为潜在语义前缀注入冻结的LLM策略中,使优化能够在紧凑的语义空间中实现,而非对原始令牌序列进行优化。通过将强化学习从令牌级探索转向经验级推理,CodeSkill提升了优化效率和长期行为一致性。大量实验表明,CodeSkill在多种通用和工业编码基准中,面对强开权基准时表现出极高竞争力的性能。此外,所学技能表现出强的可迁移性和稳健的跨域泛化能力,凸显了显式行为抽象在可扩展代理代码生成中的有效性。
ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents
ActiveMem:面向长视界代理的动态潜在记忆树
- Authors: Song-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33244
- Pdf link: https://arxiv.org/pdf/2609.33244
- Abstract
Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-step execution. As memory scales, such flat retrieval introduces context fragmentation and cross-task interference, leading to structurally inconsistent reasoning trajectories. We propose ActiveMem, a hierarchical memory framework that recursively organizes agent experiences into dependency-aware latent execution trees. ActiveMem abstracts trajectories into reusable subtask nodes while explicitly preserving execution transitions, enabling coherent reasoning-path retrieval conditioned on the current execution state. To support continual adaptation, ActiveMem further learns dynamic memory expansion, retrieval, and pruning policies through reinforcement learning. Experiments across various agent benchmarks demonstrate that ActiveMem consistently improves task completion, reasoning stability, and memory efficiency over existing memory-based agents. Moreover, ActiveMem enables compact open-weight models to achieve competitive performance with substantially larger proprietary systems.
- 中文摘要
大型语言模型(LLM)代理越来越依赖外部内存来支持长视野推理和决策。现有内存系统通常以独立的上下文片段形式检索历史轨迹或摘要,忽略多步执行背后的过程依赖。随着内存扩展,这种扁平检索引入上下文碎片和跨任务干扰,导致结构不一致的推理轨迹。我们提出了ActiveMem,一种层级记忆框架,递归地将代理体验组织为依赖感知的潜在执行树。ActiveMem 将轨迹抽象为可复用的子任务节点,同时显式保留执行转移,支持基于当前执行状态的连贯推理路径检索。为支持持续适应,ActiveMem 通过强化学习进一步学习动态内存扩展、检索和剪枝策略。跨多个智能体基准测试的实验表明,ActiveMem 在任务完成率、推理稳定性和内存效率方面持续提升,优于现有基于内存的智能体。此外,ActiveMem 使紧凑的开放权重模型能够在更大规模的专有系统中实现竞争性能。
Minimax-Optimality of Posterior Sampling for Reinforcement Learning
后置采样的极小极大最优性用于强化学习
- Authors: Taewon Goo, Kihyuk Hong
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.33246
- Pdf link: https://arxiv.org/pdf/2609.33246
- Abstract
Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.
- 中文摘要
强化学习的后验抽样(PSRL)是最简单且最有效的探索方法之一,但一个基本问题仍未解决:未经修改的PSRL是否能在不依赖先验结构假设的情况下实现极小最大遗憾?我们的答案是肯定的。在任意相关先验下,精确的原版PSRL在前导贝叶斯遗憾中是极小极大最优。难点在于后验采样的转移模型与其自身的延续值耦合。我们通过一个共同的经验转移参考(隔离结果值不匹配)和基于Bellman的方差论证克服,后者控制该模型且不依赖额外的前导阶状态空间因子。对于有限视界、时间不齐次且随机奖励未知的表格MDP,这给出了在任意联合先验下奖励和转移下的极小最大$\widetilde{O}(\sqrt{SAH^3K})$。同样的证明原理也给出了在任意联合参数先验下线性混合MDP的极小极大$\widetilde{O}(d\sqrt{H^3K})$%率。
GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences
GTRL:扎根、分化并征服时间差异的价值学习
- Authors: Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33259
- Pdf link: https://arxiv.org/pdf/2609.33259
- Abstract
In offline goal-conditioned reinforcement learning (GCRL), divide-and-conquer scales to long horizons by joining two shorter segments at a subgoal. However, under stochastic dynamics, the base case of this rule values the luckiest trajectories through the data. The subgoal must also lie on a shared trajectory, so a state-goal pair that no trajectory connects gets no value update at all. To address both, we present Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target. Over a single step, TD is correct, as its target averages over the successors and needs no subgoal. GTRL adds this target to the composition rather than replacing it, so every pair receives an update, and the composition still carries the long horizon. GTRL also corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors. We evaluate our algorithm on nineteen OGBench tasks spanning stochastic, deterministic, and stitching environments, where it achieves the highest average success rate. Code will be released soon.
- 中文摘要
在离线目标条件强化学习(GCRL)中,分而治之通过在子目标处连接两个较短的段子,扩展到较长的视野。然而,在随机动力学下,该规则的基础情形会以数据中最幸运的轨迹为值。子目标还必须位于共享轨迹上,因此没有轨迹连接的状态-目标对则不会获得任何值更新。为解决这两点,我们提出了基于基层传递RL(GTRL)的离线GCRL值学习算法,它通过一个一步TD目标来对分化与征服的更新进行基础化。在单一步内,TD是正确的,因为其目标对后续目标的平均值,无需子目标。GTRL将该目标添加到组合中而非替换,因此每对目标都会获得更新,组合仍保持长视野。GTRL还通过根据其他继任目标的可达性进行权重,纠正了事后诸葛亮重新标记带来的偏见。我们在19个OGBench任务中评估了我们的算法,这些任务涵盖随机、确定性和拼接环境,其中平均成功率最高。代码将很快发布。
Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space
接下来的思考是分布:潜在空间中的生成自回归推理
- Authors: Yang Li, Yi Wang, Shiyuan Huang, Yang Liu, Hao Wang, Chengzhi Mao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33271
- Pdf link: https://arxiv.org/pdf/2609.33271
- Abstract
Reasoning problems often admit multiple valid ways to proceed. Continuous reasoning promises to move computation beyond language tokens into a more compact latent space, but representing several plausible ways to think next remains difficult. We introduce Autoregressive Thought Flow (ATF), which models the next continuous thought as a multimodal distribution. A causal autoregressive model performs the reasoning computation, while a lightweight diffusion head generates a plausible next thought from the resulting condition. The sampled thought is fed back into the model, allowing continuous reasoning to unfold for a variable number of steps while preserving the pretrained backbone. Across mathematical reasoning tasks, ATF improves accuracy with compact latent traces and benefits from reinforcement learning and additional test-time thinking. Multi-sample evaluation shows broader solution coverage, indicating that its multimodal predictions capture useful diversity among reasoning paths. Our results suggest that continuous reasoning is more effective when multiple possible next thoughts remain available rather than being collapsed into a single prediction.
- 中文摘要
推理问题通常允许多种有效推进方式。连续推理有望将计算从语言符号扩展到更紧凑的潜在空间,但要表达多种合理的“下一步思考方式”仍然困难。我们介绍了自回归思维流(ATF),将下一个连续思考建模为多模态分布。因果自回归模型执行推理计算,而轻量级扩散头则从结果条件中生成合理的下一个思考。采样的思维被反馈回模型,使连续推理在可变步数内展开,同时保持预训练骨干。在数学推理任务中,ATF通过紧凑的潜在迹提升准确性,并受益于强化学习和额外的测试思维。多样本评估显示出更广泛的解覆盖率,表明其多模态预测捕捉了推理路径间有用的多样性。我们的结果表明,当多个可能的下一个想法仍然存在时,连续推理比被压缩成单一预测更有效。
Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
学习销售:多产品市场中战略性大型语言模型代理的强化学习
- Authors: Shuze Daniel Liu, Claire Chen, Jiuqi Wang, Thorsten Joachims
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33289
- Pdf link: https://arxiv.org/pdf/2609.33289
- Abstract
Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.
- 中文摘要
在多产品市场中运行的自主大型语言模型(LLM)代理必须在信息不对称和资源限制下做出顺序决策。我们开发了一种机器学习方法,用于训练此类代理在多项目议价环境中有效地作为卖方行动,即卖方同时在独立买家池中协商可替代资产目录。买家持有跨产品私有、异构的估值,每个买家最多只能购买一件商品。面对总沟通回合的限制,卖方必须动态地将买家与考虑私人估值后最有利可图的产品匹配,同时战略性地将有限的互动预算分配给潜在价值更高的组合。我们将该问题形式化为部分可观察的马尔可夫决策过程,使用结构化的四部分消息协议,将自然语言映射到可解析且可调控的决策空间。基于这种形式化,我们设计了一种基于可验证奖励强化学习(RLVR)的训练后方法。为评估该框架,我们构建了一个多维度量套件,量化约束遵循度、卖方剩余提取和配置质量。我们受过训练的卖方代理学习更有效地将有限库存与买方匹配,在卖方剩余提取和买方-产品配置质量上均匹配甚至超越万亿参数前沿模型。最后,这些学习到的策略能够稳健地推广到未见的市场结构、相关估值分布和培训中未遇到的价格区间。
Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
校准,而非答案选择:推理模型中内部信心的提炼
- Authors: Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33290
- Pdf link: https://arxiv.org/pdf/2609.33290
- Abstract
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.
- 中文摘要
带有二元正确性的强化学习奖励的是正确性,而非校准置信度。推理模型口头表达的信心系统性地过于自信,问题不仅仅是尺度问题:口头化信心跟踪模型承诺答案的意愿,而非答案正确的概率。因此,事后重标度只拟合一个分布,但很少转移。我们转而观察模型内部。在事实性问题回答中,对思考链与答案之间隐藏状态的线性探针校准得更好:其预期校准误差比四个基准测试和两个模型族的口头评分低5到38倍。然而,当用它从N个抽样答案中选择时,该探针几乎能平分,但远低于预言机。内部状态在“我有多确定”和“哪个答案正确”方面回答得不好,因此信号应作为置信度报告,而非用来选择答案。因此,我们引入了探针引导自蒸馏(Probe-SD):用探针对模型自身采样的痕迹进行评分,覆盖每个痕迹所表达的置信度,并微调同一族的基础检查点,使测试时除了模型本身外,其他都不存在。在Qwen3-14B中,Probe-SD将ECE从域内0.178降至0.024,域外将0.542降至0.113,同时优于事后再校准和自一致性蒸馏。所得置信度校准良好,适用于加权投票,这些行为此前归因于在线强化学习,而此时仅通过监督微调获得。
Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning
安全评分匹配:传播政策与汉密尔顿-雅各比可达性,用于在线安全强化学习
- Authors: Boyang Li, Matthew Kim, Sylvia Lee Herbert
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.33337
- Pdf link: https://arxiv.org/pdf/2609.33337
- Abstract
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.
- 中文摘要
在线安全强化学习(RL)寻求在满足安全约束的同时最大化奖励的策略。安全强化学习中一种流行的研究方向将安全性放宽为软期望成本约束,并通过仅在平均情况下强制执行安全的原始对偶拉格朗日更新解决受限马尔可夫决策过程。为解决这一限制,引入了硬的状态级约束,通常通过哈密顿-雅可比(HJ)可达性强加。然而,这些约束要求在可行和不可行区域解决不同的目标:前者是奖励最大化,后者是向可行区域的恢复。由此产生的目标动作分布本质上是多模态的,这种结构对现有基于HJ的安全强化学习中所用的高斯或确定性行为者构成根本挑战,因为它们常常崩溃到次优模式。扩散策略提供了表示此类分布所需的表达性,近期关于Q分数匹配的研究为在线强化学习(按分数回归)提供了训练策略的途径——但仅应用于奖励最大化。我们提出了安全分数匹配(SSM),这是一种非策略的行为者-批判者方法,通过对具有HJ可达性的双分支分数目标进行门控,将Q分数匹配适配到硬约束安全强化学习:在可行集合内,去噪过程进化为对HJ批评者归类为可行行为的Q分数匹配;外部,恢复分支偏向最坏情况违规较低的区域去噪。在四旋翼和固定翼轨迹跟踪及稳定避让基准测试中,SSM以低假安全率实现最佳或接近最佳任务表现,而原始-双基线允许更多不安全行为,基于可达性的基线则更保守;在安全-体育馆速度任务中,SSM以竞争性回报实现最低成本。
Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
超越时间戳:长期战略代理的决策对齐政策提炼
- Authors: Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang, Shiji Zhou
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33391
- Pdf link: https://arxiv.org/pdf/2609.33391
- Abstract
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at this https URL
- 中文摘要
带有可验证奖励的强化学习(RLVR)通常依赖稀疏的结果奖励,为长视野代理提供粗略监督。策略上自我蒸馏(OPSD)则通过密集的特权反馈补充这一信号。然而,我们识别出\emph{Decision--Timestamp Mismatch}:特权指导可能与学生的功能决策不匹配,因为相应决策可能发生在不同时间步,而学生的决策可能跨越多个时间步,而非绑定于单一时间戳。因此,时间戳-局部监督可能同时导致学分的上下文和时间范围不匹配。为解决这种不匹配,我们引入了\textsc{AlignOPSD},遵循在分配学分前对齐监督的原则。决策对齐督导纠正在功能匹配的情境下重新评分同一学生抽样反应,以校准本地教师证据。半马尔可夫层级学分分配随后从函件变化中推导变量时长决策跨,并利用纠正证据在跨跨及其组成回合分配基于结果的学分。我们在ALFWorld、WebShop和Search-QA上对QWEN2.5-3B和Qwen2.5-7B评估\textsc{AlignOPSD},并以代表性基线进行评估。\textsc{AlignOPSD}在所有八个骨干-聚合-指标比较中均优于GRPO和StepOPSD,GRPO提升5.5-8.7%,排名六个中第一。额外分析还考察了任务间的两个对齐阶段及超参数敏感性。我们的代码可在此 https URL 获取
COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
COEVO:递归自我提升的共演化背景与参数
- Authors: Siwei Chen, Xinping Bao, Xinyu Cai, Yuan Cao, Wan Jiang, Shaohong Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33398
- Pdf link: https://arxiv.org/pdf/2609.33398
- Abstract
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
- 中文摘要
递归自我改进(RSI)旨在将大型语言模型从静态训练管道转向能够参与改善自身未来行为的系统。现有方法主要遵循两个方向:通过在线学习更新模型参数,或通过搜索、反思和提示优化改善外部语境。虽然这两种机制都能支持持续改进,但通常独立研究。这种分离忽略了一个重要交互作用:情境塑造模型学习的体验,而演变中的模型可能随着时间推移以不同的方式解释和利用同一语境。因此,我们将RSI表述为参数-语境共进化问题,即模型参数与学习语境在共享反馈循环中适应。我们介绍COEVO,这一框架根据政策经验更新模型参数,同时根据策略演变状态调整上下文指导。策略熵和提示条件关注作为辅助信号指导这种适应。实验表明,COEVO在任务表现上优于固定上下文强化学习,并产生对系统提示变化更具鲁棒性的策略。更广泛地,我们的结果表明,外部上下文应不仅仅视为大型语言模型的固定接口,更应视为递归自我提升的自适应组成部分。
SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
SciGen-Verifier:科学图像生成中可解释验证的多模态推理器
- Authors: Jiali Chen (South China University of Technology, The Hong Kong Polytechnic University), Zhengteng Lin (South China University of Technology), Zuqi Wang (South China University of Technology), Shirong Lin (South China University of Technology), Xi Yu (South China University of Technology), Xusen Hei (South China University of Technology), DingBa Fu (South China University of Technology), Jiayuan Xie (The Hong Kong Polytechnic University), Yi Cai (South China University of Technology)
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.33399
- Pdf link: https://arxiv.org/pdf/2609.33399
- Abstract
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.
- 中文摘要
在现实教育中,解决方案不仅通过文字表达,还通过图画——电路、几何构图、函数图——教师必须像对文本一样仔细评分绘图。统一多模态模型的最新进展使得科学图像生成成为可能,但验证这些专业视觉输出的正确性仍是一个关键瓶颈:错误常源于复杂的领域知识、结构推理和多步骤指令,而非表面伪影。现有验证器主要针对自然图像,将判断压缩为标量分数,导致科学覆盖和可解释的错误纠正反馈未被充分探索。为弥合这一差距,我们提出了三项主要贡献。(1)我们构建了SciGen-Verify,这是一个专注于科学图像生成可解释验证的基准工具,涵盖指令跟踪、多学科推理和世界知识领域。它包含三层分层协议,涵盖二元判断、支持解释和纠正编辑指令。(2)我们开发了SciGen-Verifier,这是一种基于推理的多模态验证器,通过冷启动监督微调训练,随后是基于课程的两阶段强化学习流程。评分标准引导的过程首先强化科学推理探索,结果奖励随后将输出与真实注释对齐。(3)在SciGen-Verify上,SciGen-Verifier在与更大专有模型竞争中表现出色。它还作为迭代图像校正的实用在线批评者。
VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings
VaME:探索多模嵌入的变分潜在推理
- Authors: Peixi Wu, Mingzhou Jiang, Feipeng Ma, Biao Yang, Yunhao Zhou, Wei Yuan, Bosong Chai, Huizu Lin, Jie Chen, Zhangchi Hu, Fan Yang, Wenwu Ou, Hebei Li, Xiaoyan Sun
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.33402
- Pdf link: https://arxiv.org/pdf/2609.33402
- Abstract
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
- 中文摘要
通用多模态检索需要紧凑的嵌入,以保留跨多种模态的任务相关语义信息。此前的研究已将潜性推理纳入多模嵌入学习,以在嵌入提取前对信息进行细化。然而,大多数现有方法仍局限于确定性潜在路径,未探索替代路径以发现更好的嵌入。因此,我们提出了VaME(变分多模嵌入)框架,将潜在推理建模为轨迹上的可学习分布。具体来说,我们首先引入变分潜在推理(VLR),以实现在潜空间中的自回归探索,通过轻量级解码器进行答案重建。同时,我们用潜在融合嵌入增强原始嵌入符号读出,以便后续强化学习中的探索。最后,我们通过强化学习优化潜推理,利用语义解码奖励(SDR)优化具有语义意义且解码结果可解释的轨迹。在涵盖图像、视频和视觉文档检索的78任务MMEB-V2基准测试中,VaME表现优于大多数显式基于CoT的模型及所有潜在推理基线。VaME在推理密集型基准测试如MRMR上表现出稳健性能,强化学习后显著提升。值得注意的是,VaME的推理速度至少比确定性潜在自回归基线快4.25倍。代码将公开发布。
TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
TTRSD:视觉语言模型的测试时间强化学习与自我蒸馏
- Authors: Shuning Wang, Zhiheng Wu, Xun Zhou, Chongyang Cui, Chen Jia, Bowen Liu, Chuanjie Li, Xiang Chen, Yi Yang, Yumeng Zhang, Wenjie Huang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33414
- Pdf link: https://arxiv.org/pdf/2609.33414
- Abstract
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
- 中文摘要
测试时强化学习使视觉语言模型(VLM)能够利用无标签输入进行适应。然而,在固定视觉条件下的重复抽样可能强化共享的感知错误,而序列级奖励未能隔离视觉感知——这是多模态推理的核心瓶颈,存在预训练推理能力退化的风险。我们提出了TTRSD,一种测试时强化学习框架,结合了多视角答案层级的自我蒸馏与视觉对比标记选择。共享策略将教师预测在原始、裁剪和降采样视图中聚合为答案分布。学生从原始图像生成的轨迹根据其最终答案的支持度获得奖励。为了精确地将反馈分配给感知瓶颈,我们比较同一抽样标记在原始和视觉吸收输入下的对数概率,同时保持文本前缀固定,选择对视觉敏感的位置进行策略梯度更新。TTRSD将由群体相对优势决定的更新方向与由视觉敏感度决定的更新位置分离,无需真实标签、外部验证器或单独教师。仅用20个未标记的适应样本,TTRSD在七个基准测试和三个VLM上提升了性能,将InternVL3-2B的MMMU准确率从35.79%提升至49.32%(+13.53%),展示了跨数据集泛化性,同时保持了内在推理完整性。
TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
TeacherGRPO:通过教师对齐缩小推理提炼能力差距
- Authors: Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, Jundong Li
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33426
- Pdf link: https://arxiv.org/pdf/2609.33426
- Abstract
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at this https URL.
- 中文摘要
从强大教师模型中推理提炼到较小学生面临“差距诅咒”:随着教师日益精细,其复杂分布越来越偏离学生的近似水平,导致表现下降。现有缓解策略要么通过数据选择过滤有挑战性的例子,要么引入较弱的中级助理模型,本质上损害监督覆盖率或质量。我们提出教师对齐机制,直接调整教师以适应学生的分布,同时不丢弃数据或降低推理质量。然而,通过标准知识蒸馏进行朴素对齐会引发教师推理能力的灾难性崩溃。为此,我们将教师对齐重新表述为强化学习,并引入基于群体相对策略优化、并有两项关键创新的TeacherGRPO:(i)课程选择性对齐采用双重代币级和分布级课程,将奖励聚焦于高信号推理差距,同时过滤琐碎代币和不确定尾部分布的噪声;(ii)重要性-自适应长度正则化选择性惩罚冗长冗余,同时保留教学上关键的推理步骤。对齐后的教师随后通过标准流程向学生提炼知识。大量实验显示,TeacherGRPO在多种推理基准和提炼方法中显著优于基线。我们的代码可在此 https URL 获取。
Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning
阐明基于回归的扩散强化学习设计空间
- Authors: Toyota Li, David Zhao, Alan Zhao
- Subjects: Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.33444
- Pdf link: https://arxiv.org/pdf/2609.33444
- Abstract
A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.
- 中文摘要
一种新兴的方法家族放弃策略梯度,转而重权监督回归,在扩散和流模型的强化学习中获得了动力。DiffusionNFT、FlowAWR和RAM是具有不同动机的代表性模式。它们之间是否共享什么尚不明确。我们证明,每一种都是对一个发散约束奖励最大化问题的解,且仅通过定义约束的凸生成器进行区分。在统一建模框架下,我们揭示了先前技术在构建优势嵌入回归目标时所做的松弛:分别近似DiffusionNFT和FlowAWR线性和指数倾斜形状的KKT条件和后验归一化,同时保持线性倾斜概率单纯形的精确稀疏极大投影,从而在本工作中引入另一种更优模型类型。除了理论基础,我们还通过实证方式研究了设计空间,并揭示了回归式扩散强化学习的训练配方。保留探索中发现的优点催生了扩散RFT,这是我们收敛更快、训练更稳定、性能更优的范式。
DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
DISCO:带有基础推理分解的分布式长上下文尺度
- Authors: Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav Kveton, Seunghyun Yoon, Franck Dernoncourt, Qizhe Xie, Trung Bui
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33485
- Pdf link: https://arxiv.org/pdf/2609.33485
- Abstract
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
- 中文摘要
虽然大型语言模型(LLM)宣传百万令牌上下文窗口,但随着输入增加,推理质量常常崩溃——这种现象称为上下文腐烂。这种失败源于单体架构中的结构纠缠,上下文基础的巨大搜索负担耗尽了复杂推理所需的表征能力。为解决这个问题,我们提出了通过DIStributed长文本缩放(DISCO)进行基础-推理拆分。DISCO受Apache Spark等分布式计算框架启发,将长上下文划分为一组专门用于并行、局部基础的Worker LLMs。一个通过强化学习(GRPO)训练以优化规划的中央驱动LLM,通过动态映射查询到原子提取任务并减少收集的证据来协调执行,最终得出答案。通过将推理与原始上下文噪声隔离开来,DISCO有效消除了上下文腐烂。在RULER-QA(100万个代币)上,它在标准基线崩溃时保持78.4%的准确率。此外,它在LongBench v2上比全上下文模型高出高达9.8个百分点,并与Gemini-3-Pro-Preview等前沿模型匹敌,同时将推理成本降低80%以上,建立了高效的长上下文推理范式。
Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning
利用多智能体强化学习实现的联合多模态人类活动识别
- Authors: Debasmita Dey, Tanmay Sen, Himel Mallick
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33492
- Pdf link: https://arxiv.org/pdf/2609.33492
- Abstract
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
- 中文摘要
来自异构可穿戴传感器的人体活动识别(HAR)是物联网(IoHT)的基础,支持康复、老年护理和智能医疗。现有多模态融合方法通常对传感器流赋予固定的相等权重,忽略模态重要性、获取成本和传感器质量的差异,这些差异可能因移动、错误位置或暂时阻塞而变化。我们提出了基于多智能体强化学习的自适应且成本意识型多模态HAR框架,用于集中化HAR,并将其扩展到联邦学习,称为FedMHAR。在集中化环境中,多模态融合被表述为合作式多智能体强化学习(MARL)问题,每个感测模态被分配一个基于PPO的智能体,学习每样本融合权重,使模型能够强调信息型态,同时在廉价替代方案提供足够信息时降低昂贵传感器的权重。在联邦环境中,我们引入了BiFL-PPO,这是一种双向联邦优化策略,服务器端PPO策略学习客户端特定信任权重并反馈以适应局部学习率和近端正则化。与轮级优化不同,BiFL-PPO采用密集的批次级奖励,在异构客户端数据下提供更频繁的反馈和稳定训练。对MEx修复和UTD多模态人类行动数据集的评估显示,该集中式框架分别实现了87.30%和94.98%的准确率,优于传统融合方法和最先进的HAR模型。FedMHAR在联邦环境中分别达到79.74%和77.49%,持续超越FedAvg、FedProx、FedBN、FedNova和AdaFedProx,同时提供更稳定的性能并降低传感器获取成本。
MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
MA-JEPA:多智能体强化学习的联合嵌入世界模型
- Authors: Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2609.33563
- Pdf link: https://arxiv.org/pdf/2609.33563
- Abstract
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
- 中文摘要
世界模型通过对想象轨迹进行训练策略提升样本效率,但其实用性依赖于能够捕捉未来控制所需信息的学习表征。我们研究自监督联合嵌入预测(JEPA)是否能为多智能体强化学习提供这种学习信号。我们引入了MA-JEPA,一种随机世界模型,用目标表示的预测取代了观察重建,实现基于模型的多智能体强化学习,实现集中训练和去中心化执行。一个类别潜态和一个因果变换器通过后验和动作条件动力学预测目标训练,然后用于潜在想象中的行为者-批判者学习。仅训练的联合预测器对所有智能体的局部状态和动作进行条件,以预测每个智能体下一次本地观察嵌入。这些预测会通过与中心化批评者实际交互时使用的同一局部后验传递,该后验仅用于价值学习,执行则保持去中心化。我们的实验表明,该架构在SMAC上表现优异,在八张评估地图中有四张中,匹配或超过报告的最强比较器平均胜率。
OpenFC: Learning Verification Policies towards Open-Search Fact Checking
OpenFC:学习验证政策以实现开放搜索事实核查
- Authors: Xinming Wang, Kaixiang Qiu, Yansong Lin, Chunji Lv, Yi Chen, Boran Wang, Hong-Ming Yang, Xu-Yao Zhang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33579
- Pdf link: https://arxiv.org/pdf/2609.33579
- Abstract
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39\% average accuracy and 63.30\% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
- 中文摘要
开放搜索事实核查不仅仅是检索后分类,而是一个顺序决策问题,每一次查询、来源访问和停止决策都会重塑可供验证的证据。然而,现有系统往往将这些决策分散在预设的流程或单独提示的模块中,而非作为统一的任务特定策略学习。我们引入了 \textbf{OpenFC},一个统一的验证-策略培训框架,后期训练 Qwen3-8B 作为紧凑的下一步行动控制器,负责推理、证据获取和停止。OpenFC 分两个阶段学习该策略。\textbf{Stepwise-Calibrated 冷启动(SCCS)} 使用强有力的培训时间监督者,在执行前审查后期推理、工具使用和停止提案,生成可靠的监督微调轨迹,无需获得黄金判决。\textbf{验证感知强化学习(VA-RL)} 通过预算感知工具奖励、标签感知优势重权和局部响应掩蔽,改进了未解决索赔的冷启动政策。在六个事实核查基准中,OpenFC 实现了70.39%的平均准确率和63.30%的宏观F1准确率,是所有评估方法中最高的整体平均值。分阶段消融进一步显示,SCCS和VA-RL提供了互补的收益,支持了两阶段训练框架的设计。这些结果使OpenFC成为一个强大且有效的开放搜索事实核查框架。我们将开源代码并发布模型检查点以支持可重复性。
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
TGRL:温度分组强化学习,用于高效探索大型语言模型
- Authors: Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Wei Lin, Guojun Yin, Ran He
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33589
- Pdf link: https://arxiv.org/pdf/2609.33589
- Abstract
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at this https URL.
- 中文摘要
高效探索往往仍是可验证奖励强化学习(RLVR)的核心瓶颈。尽管温度控制和测试时间缩放策略可以增加大型语言模型(LLM)的展开多样性,但它们要么在推广时扩大样本预算,要么使探索的收益未被量化。为此,我们提出了温度分组强化学习(TGRL),将温度诱导的多样性转化为显式训练信号。对于每个提示,TGRL将其推广组划分为低温子集和高温子集,通过奖励对比估计探索收益,并利用同一logit诱导的相应温度尺度次代币分布之间的Jensen-Shannon(JS)发散,将该组级信号分配为代币级信用。值得注意的是,TGRL在不扩展推广预算的情况下,达到比强RLVR基线快达36%的等效精度。在来自不同领域的11个基准测试中,TGRL总体上优于强劲的RLVR基线:它在32B的六个基准数学平均值中提升了1.6%,使CodeForces评分提升了196.7分,LiveCodeBench Pass@16提升了4.4%,同时ALFWorld/WebShop成功率提升了6.3%/4.9%。全面的烧蚀和墙钟分析确认了所有提议组件的有效性。代码可在此 https 网址获取。
You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
你只需编辑一次:通过本地演示细化激励LLM的上下文内能力
- Authors: Jiarong Wen, Qi Wang, Yun Qu, Yixiu Mao, Heming Zou, Haoang Chi, Lizhou Cai, Yiqin Lv, Kaiyu Zhang, Yuhang Jiang, Xiangyang Ji
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33609
- Pdf link: https://arxiv.org/pdf/2609.33609
- Abstract
In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likelihood proxies to implicitly assess ICL quality. Making repeated queries to the target LLM with these strategies can incur substantial costs. This work simplifies selection by framing it as a constrained local search problem and presents local demonstration editing (LDE). Starting with an initially retrieved set of demonstrations, LDE employs a single structured edit to explore its surrounding neighborhood while balancing performance gains with search costs. Technically, LDE is reduced to a policy search problem, for which we train a small LLM, referred to as Jev-LDE. This model as the System-1 modifies the retrieved demonstration set by performing actions such as \texttt{Keep}, \texttt{Delete}, or \texttt{Replace} elements, all within a framework of reinforcement learning with verifiable rewards. At test time, Jev-LDE executes a single edit of the retrieved demonstration set, followed by one inference from the target LLM, avoiding the need for iterative context scoring or subset searches. Across standard classification benchmarks, various target LLMs with Jev-LDE as the plug-and-play module consistently improve ICL performance, and Jev-LDE shows transferability to held-out benchmarks and models without retraining. These findings indicate that the LDE approach offers an efficient and adaptable method for harnessing the ICL capabilities of target LLMs.
- 中文摘要
上下文学习(ICL)对于提升大型语言模型(LLM)的推理性能至关重要。然而,ICL在LLM中的有效性很大程度上取决于演示集的选择。对这些集合的穷尽搜索是组合性的,现有选择器通常依赖相关性或似然代理来隐式评估ICL质量。用这些策略对目标LLM进行反复查询可能会带来巨大成本。这项工作通过将选择框架为受限局部搜索问题,简化了选择过程,并提出了局部演示编辑(LDE)。从最初检索的演示集出发,LDE采用单一结构化编辑来探索其周围邻域,同时平衡性能提升与搜索成本。从技术上讲,LDE被简化为一个策略搜索问题,我们为此训练一个小型LLM,称为Jev-LDE。该模型作为System-1通过执行诸如 \texttt{Keep}、\texttt{Delete} 或 \texttt{Replace} 元素等动作,修改检索到的演示集,所有这些操作均在强化学习框架内,并可验证奖励。测试时,Jev-LDE 执行检索演示集的一次编辑,随后从目标 LLM 进行一次推断,避免了迭代上下文评分或子集搜索的需求。在标准分类基准测试中,采用 Jev-LDE 即插即用模块的各种目标大型语言模型持续提升了 ICL 性能,Jev-LDE 显示出无需重新训练即可迁移到已保留基准测试和模型的可迁移性。这些发现表明,LDE 方法为利用目标 LLM 的 ICL 能力提供了一种高效且可适应的方法。
ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
ParaAgent:在开放世界工具环境中强化并行行动
- Authors: Shengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen, Wei Chen, Zhongyu Wei
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33618
- Pdf link: https://arxiv.org/pdf/2609.33618
- Abstract
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration $\rightleftharpoons$ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
- 中文摘要
语言模型代理越来越多地部署在开放世界工具环境中,这需要在探索未知能力和利用已知能力之间取得平衡。现有方法面临性能效率的权衡:它们要么严格解耦探索与执行,要么在不协调的情况下交错。我们认为关键不在于是否解耦或交错,而在于如何跨粒度协调它们。我们介绍ParaAct,一个结构化的并行动作循环,结合了阶段级探索(Exploration $\rightleftharpoons$)执行与动作级并行性。为学习该循环,ParaAgent结合了多智能体冷启动演示与多层优势解耦下的强化学习,使规划结构显现,并以步级、阶段级和轨迹级奖励监督。学习由我们的ToolEnv支持,这是一个基于50,011个真实工具接口的可扩展模拟器。在两个开放世界工具基准测试中,ParaAgent-4B在所有基线中平均成功率最高,包括GPT-4.1系统,在多工具任务中取得最大优势。行为分析显示,这些提升源自这种行动组织,凸显了其对高效且高效开放世界代理的重要性。
Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
攀登高峰:提示注入红队对抗前沿模式与课程强化学习
- Authors: Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia
- Subjects: Subjects:
Machine Learning (cs.LG); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2609.33628
- Pdf link: https://arxiv.org/pdf/2609.33628
- Abstract
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8\% and 45.0\% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0\% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at \href{this https URL}{here}.
- 中文摘要
提示注入是大型语言模型及其基于大型语言模型的应用(如代理)的主要安全风险。最先进的提示注入红团队方法利用强化学习(RL)训练攻击方大型语言模型生成有效的注入提示。然而,针对前沿大型语言模型如GPT-6-Luna时,冷启动问题是主要挑战:攻击者LLM的每次攻击尝试均失败,因此没有任何奖励,无法提供学习信号。本研究提出一种基于课程学习的方法来解决冷启动问题。特别是,我们提出针对一系列越来越稳健的目标LLM进行训练攻击者LLM,每个阶段均从前一个LLM获得的攻击者LLM热启动。然而,仅仅针对弱目标(如GPT-4o-mini)进行训练,可能无法充分准备攻击方LLM,从而获得对前沿LLM(如GPT-5.6-Terra)有用的学习信号。相反,我们发现课程设计至关重要:每完成一个阶段后,攻击方LLM需要部分成功对下一个目标LLM进行攻击,以便从成功攻击新目标的尝试中学习。我们的广泛评估表明,我们的方法能够有效地对前沿LLM进行红队化,在AgentDyn上对GPT-5.6-Luna和GPT-5.6-Terra的攻击成功率(ASR@10%)分别为93.8%和45.0%,而最先进的RL方法如RL-Hammer和PISmith在相同条件下的ASR为0/%。此外,我们发现攻击方LLM会跨目标转移,例如,一个训练用来击败一个强LLM(GPT-5.6-Terra)的攻击者LLM,也能成功击败其他六个未曾训练过的前沿LLM(如GPT-6-Luna)。我们的代码可在\href{this https URL}{here}获取。
Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss
基于通信丢失下仓库机器人协调的层级多智能体强化学习
- Authors: Weihao Sun, Gehui Xu, Andreas A. Malikopoulos
- Subjects: Subjects:
Systems and Control (eess.SY); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.33637
- Pdf link: https://arxiv.org/pdf/2609.33637
- Abstract
In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups. Each group uses a recurrent predictor to estimate unavailable interaction information due to communication loss. A higher-level policy then generates a compact coordination reference that conditions the local control policy within each group. A predictive safety filter evaluates and modifies the proposed controls when they violate safety constraints. Simulation results show improved task completion under communication loss, reduced communication growth as the team size increases, and safe operation in the tested scenarios.
- 中文摘要
本文提出了一种分层多智能体强化学习框架,用于在仓库环境中协调机器人团队,在通信丢失时进行协调。我们将机器人团队划分为若干组,每个组内集中协调,组间分布式协调。每个组使用循环预测变量估算因通信中断导致的交互信息不可用。更高层策略生成紧凑的协调参考,为每个组内的本地控制策略提供条件。预测安全过滤器在拟议控制违反安全约束时评估并修改。模拟结果显示,在通信中断下任务完成度提升,团队规模增加时通信增长减缓,在测试场景中运行安全。
Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
无演示成功概率奖励学习适用于通用机器人政策
- Authors: Duo Wu, Haifeng Wang, Rongwei Lu, Jinghe Wang, Tianyi Xiong, Zhimin Wang, Chao Yu, Shuai Ma, Zhi Wang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33653
- Pdf link: https://arxiv.org/pdf/2609.33653
- Abstract
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: this https URL.
- 中文摘要
强化学习(RL)使通用机器人策略通过试错互动改进,但其有效性从根本上受限于任务奖励稀疏。现有通用奖励模型通常通过专家演示学习任务进展来缓解这一问题,但与策略优化过程中遇到的混合质量推广存在分布不匹配,使得策略必须从中学习的次优和失败行为的估计不可靠。本研究引入了一种无演示的奖励学习范式,其中可直接从稀疏任务结果和政策经验中学习密集奖励反馈。理论上,我们表明终端任务结果隐含定义了中间时间步的密集成功概率反馈,这些反馈可以通过自助递归学习。基于这一见解,我们引入了eVTA$_0$,它通过时间差分式自助法从混合质量政策推广中学习成功概率,无需专家演示或中间注释。我们进一步引入了带有演化奖励(RLER)的强化学习,这是一个闭环框架,利用新收集的政策推测数据调整eVTA$_0$。实验显示,eVTA$_0$在相同RL训练预算下,提供更多信息的奖励,并在所有LIBERO任务套件中实现最佳平均政策性能,成功率比初始策略提升5.4%-13.8%。在现实操作中,RLER进一步提升整体成功率20%-26%,在非分配条件下提升35%-36%。这些结果展示了无示范奖励学习和随着政策演变调整奖励的有效性。项目网页:此 https URL。
LieDiscover: Adaptive Symbolic Library Construction for Explicit Open-form Symmetry Discovery
LieDiscover:显式开放形式对称发现的自适应符号库构造
- Authors: Xinxin Li, Jianming Ma, Xingyu Cui, Da Li, Juan Zhang, Junping Yin
- Subjects: Subjects:
Symbolic Computation (cs.SC); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33663
- Pdf link: https://arxiv.org/pdf/2609.33663
- Abstract
Discovering underlying symmetries from data has emerged as a crucial challenge in scientific discovery. Existing data-driven methods for symmetry discovery fail to determine the exact number and mathematical form of unknown infinitesimal generators. Recent explicit methods represent generators using a predefined function library and identify them through algebraic optimization, but they often struggle to capture complex symmetries involving high-order polynomials or transcendental functions. To address this limitation, we formulate symmetry discovery as a joint optimization problem over the function library and coefficients. We propose a novel framework that leverages an encoder-decoder architecture to dynamically generate symbolic expressions and expand the library. This generation process is optimized via reinforcement learning, which accelerates the exploration of the symbolic search space through step-wise rewards. Experiments demonstrate that LieDiscover can successfully uncover open-form infinitesimal generators involving high-order polynomials or transcendental functions, which remain intractable for existing methods. The discovered symmetries also improve performance in downstream PDE solving and discovery tasks.
- 中文摘要
从数据中发现潜在对称性已成为科学发现中的关键挑战。现有的数据驱动对称发现方法未能确定未知无穷小生成元的确切数量和数学形式。近期显式方法表示使用预定义函数库的生成元,并通过代数优化识别,但它们常常难以捕捉涉及高阶多项式或超越函数的复杂对称性。为解决这一限制,我们将对称性发现提出为函数库和系数的联合优化问题。我们提出了一种新框架,利用编码器-解码器架构动态生成符号表达式并扩展库。该生成过程通过强化学习得到优化,通过逐步奖励加速符号搜索空间的探索。实验表明,LieDiscover能够成功发现涉及高阶多项式或超越函数的开式无穷小生成元,而这些生成元在现有方法中仍难以处理。发现的对称性还提升了下游偏微分方程求解和发现任务的性能。
CompoWorld: Compositional Environment Scaling for General Agents
CompoWorld:通用代理的合成环境尺度
- Authors: Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33665
- Pdf link: https://arxiv.org/pdf/2609.33665
- Abstract
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
- 中文摘要
自动生成的环境为训练通用代理提供了可扩展的交互数据来源。然而,现有方法主要在单一环境中生成任务,而现实世界的工作流程则要求代理跨多个服务连接信息和操作。我们引入了组合环境缩放(\textbf{CompoWorld}),通过组合有限的可重用服务库来扩展任务空间。编码代理将工具规格转化为带有类型状态和共享接口的验证服务,而世界模型则处理无法可靠实现的工具。随机游走过程通过依赖图连接服务,使需要信息在服务间流动的任务能够生成和验证。经过验证的轨迹支持监督微调(SFT),而我们的完成导向评分标准通过强调每个推广组内较低通过率的标准,引导强化学习(RL)实现任务的完整完成。我们构建了448个服务,暴露了10,130个工具,并使用3K的SFT轨迹和1K的强化学习任务来训练Qwen3.6-35B-A3B。实验结果显示,CompoWorld在八个基准测试中平均提升了9.17个基准。在AutomationBench上,它超越了Claude Opus 4.6等前沿模型,并在所有针对智能体的35B-A3B模型中领先。
Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs
共享经验,独立学习:伴随置信度校准用于大型语言模型
- Authors: Shiyu Ni, Keping Bi, Jiafeng Guo, Yilong Xu, Jingtong Wu, Zengxin Han, Xueqi Cheng
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33721
- Pdf link: https://arxiv.org/pdf/2609.33721
- Abstract
Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.
- 中文摘要
可靠的自我评估对于大型语言模型(LLM)至关重要,但它们在答案错误时往往保持高度自信。我们研究\emph{并发置信度校准},其中信心与能力提升同时学习,而非仅在训练后校准。可验证奖励强化学习(RLVR)为该范式提供了自然环境,因为它持续生成与可验证正确性反馈相结合的响应。然而,现有并发方法通过在共享策略参数内的强化学习同时学习能力和信心,可能将两种根本不同的学习问题耦合。我们提出 \emph{共享经验但独立学习}:能力和信心从相同轨迹学习,但通过不同的优化机制和参数。基于这一原则,我们引入了 \textbf{CoCal(伴随信心校准)},它通过扩展隐藏状态和验证者衍生的正确性监督训练轻量级伴随者,同时保持任务优化不变。Qwen3-8B和Qwen3-14B的实验显示,CoCal在不牺牲任务表现的前提下提升了置信估计,优于基于强化学习的并发方法和匹配的事后校准。该学习伴随者进一步在不同领域和政策转变中进行了推广,而CoCal的益处在两个尺度上均存在。
DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
DuraS2ST:针对时长对齐的语音转语音翻译的思维链与强化学习
- Authors: Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu, Yuanyuan Wang, Weidong Chen, Helen M. Meng, Shujie Liu, Xixin Wu
- Subjects: Subjects:
Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- Arxiv link: https://arxiv.org/abs/2609.33742
- Pdf link: https://arxiv.org/pdf/2609.33742
- Abstract
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: this https URL.
- 中文摘要
在视频配音等时间敏感应用中的语音转语音翻译(S2ST)不仅需要语义准确性和说话者保持,还需要严格的时长一致性以避免视听错位。然而,现有的S2ST系统大多生成目标语音时没有明确的时间规划,这使得时长控制成为一个尚未解决的挑战。我们介绍了DuraS2ST,一种时长对齐的推理框架,使单个语音语言模型能够先生成明确的思维链(CoT),用于规划目标措辞和语音长度,然后合成相应的语音符号。为支持这一范式,我们构建了DuraSet-440K,一个高质量、时长对齐的CoT语料库,用于监督初始化。我们进一步优化模型,采用多模态多维强化学习,利用持续时间保证金(Duration Margin Reward)来平衡翻译质量和持续时间一致性,并利用模态感知奖励归因(Modality-Aware Reward Attribution)将奖励分配到合适的代币区间。在CVSS-T上的实验显示,DuraS2ST在翻译质量与持续时间一致性之间实现了强平衡,优于竞争激烈的开源和商业基线。项目页面:此链接链接。
Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies
冻结生成机器人策略在线适配的主要引导子空间
- Authors: Jialeng Ni, Nathan Zhao, Kunpeng Song
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.33765
- Pdf link: https://arxiv.org/pdf/2609.33765
- Abstract
Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.
- 中文摘要
生成机器人策略提供表达行为先验,但通过在线交互更新大型扩散或流匹配模型成本较高。潜在空间强化学习通过控制初始采样噪声避免更新预训练生成器,但高维噪声对解码动作会产生强烈各向异性影响。我们引入主引导子空间(PSS),一种前向查询接口,从有限差分解码器响应构建固定的低维控制基。软演员-批判控制前导响应方向,正交补集则在每次查询中独立从高斯先验中重新采样。在三个具有扩散和流匹配策略的机器人模拟任务中,响应谱显示出显著的集中度。在五对匹配的任务生成器中,训练曲线表明PSS通常收敛更快,晚期训练行为比全潜控制更稳定,同时整体表现更强。受控扩散-平方消融进一步表明,领先响应方向优于同维随机和响应最低的子空间。我们进一步将PSS与冻结的闭源3B参数视觉-语言-动作(VLA)策略集成到具备同步过渡收集、重置时间优化和延迟感知异步部署的人形学习系统中。在探索性螺丝刀放置评估中,冻结VLA策略成功率为2/10,SAC+PSS适配后成功率为6/10。这些结果支持解码器响应几何作为冻结生成机器人策略在线适配的实用基础。
Selecting Diverse SFT Traces Improves Post-RL Generalization
选择多样化的SFT迹可以提升强化后推广能力
- Authors: Dylan Zhang, Mingyuan Wu, Jinning Li
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33780
- Pdf link: https://arxiv.org/pdf/2609.33780
- Abstract
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.
- 中文摘要
经过验证的解在强化学习(RL)推理模型准备方面并不同等有用。我们提出了对路径多样性的综合研究,即监督式微调(SFT)数据中推理步骤序列的变异,并提出了一种轻量级、基于规则的指纹选择方案。从同一预算的池中,结合匹配的训练配方和检查点,选择多样化而非相似路径,提升了在谜题和数学上的后续问题覆盖率,包括比任一训练阶段更难的问题。在综合实验中,路径多样化SFT在SFT保留的环境中提高了OLMo3-7B的pass@8 16.9个点。在单模型条件下,即一个模型编写每个候选人,多样性选择在10个数学基准测试中可提升最多6.2个平均pass@8。前强化学习的诊断表明原因:多样化SFT即使平均准确率略低,仍能对更多提示产生成功和失败的尝试,从而使组相对强化学习获得更多带有学习信号的提示。在3个开源语料库中,我们的仅CPU选择器在不使用模型调用的情况下,在所有平均后强化表现比较中都优于成本更高的替代方案。这些结果表明推理路径多样性是选择更适合强化学习模型的SFT数据的实用标准。
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
意外成功,反复失败:熵引导学分作业用于探索大型语言模型推理
- Authors: Woongyeong Yeo, Minki Kang, Chanuk Lee, Sangwoo Park, Jinheon Baek, Sung Ju Hwang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.33781
- Pdf link: https://arxiv.org/pdf/2609.33781
- Abstract
Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
- 中文摘要
带可验证奖励的强化学习(RLVR)通过结果级反馈增强大型语言模型(LLMs)的推理能力,但近年来更细粒度的信用分配方法通常需要辅助模型、额外抽样或特权信息。虽然策略熵提供了易于获得的信号,但在强化和惩罚下优先考虑不确定位置会集中惩罚,即失败反应仍保留恢复的选项,从而抑制探索的机会。为此,我们引入熵优势策略优化(EAPO),这是一种以熵为导向的信用分配方法,对成功和失败进行非对称处理。具体来说,基于观察到不确定性下的成功较难重复,而信心失败往往反复出现,EAPO将归一化的策略熵与响应优势符号结合,以强化意外成功并纠正反复失败。它在成功响应中对高熵决策施加更强的强化,在失败反应中对低熵决策施加更强的惩罚,同时在不确定位置减弱惩罚以保留恢复机会。通过将响应优势重新分配到不同代币,EAPO直接从现有的推广信号中获得代币级的信用,无需额外监督。我们在基础和推理骨干上的多种推理任务中验证了EAPO,证明其实现了最佳的整体性能。我们还进一步证明,EAPO促进了更有效的探索,拓宽了问题覆盖范围,并生成了更多样化的候选答案。
RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation
RAISE:通过符号评估强化LLM中的访问控制策略综合
- Authors: Yingming Zhou, Adarsh Vatsa, William Eiers
- Subjects: Subjects:
Software Engineering (cs.SE); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33796
- Pdf link: https://arxiv.org/pdf/2609.33796
- Abstract
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.
- 中文摘要
将自然语言访问控制需求转化为策略需要对权限、约束和异常进行仔细推理,甚至前沿大型语言模型也常常生成违反预期授权语义的策略。据我们所知,我们构建了CedarInstruct,这是首个支持训练和语义评估的数据集,用于正式验证的Cedar策略综合。它包含了跨44个域的5800个场景和1408个代表单一合成组织的场景,每个场景都有已验证的目标策略和可执行的验证计划。基于这些数据,我们介绍了RAISE,它通过两阶段从形式验证训练策略合成器,先是验证监督微调(SFT),随后是从验证器信号学习的强化学习(RL)阶段。我们发现SFT之所以成功,是因为让模型表达已有的授权逻辑,因为未训练的模型很少写出有效的Cedar,但通常能正确推理。SFT之后,验证者信息的使用方式比使用多少更为重要。在六个消耗越来越丰富的验证者信号的强化学习实例中,只有RAISE-OC在SFT上有显著提升;它将失败的检查和符号反例转化为引导探索,并通过非上下文的GRPO从结果中学习。通过约5.4K个已验证的场景和LoRA微调,RAISE-OC训练Qwen3.5-9B在保留场景中语义成功率提升13.33个百分点,超越Claude Opus 5,语义成功率提升13.33个百分点,提升16.26个百分点,并将训练转移到独立构建的CedarBench。
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
QwenGyre:用于训练xLong-Horizon代理的弹性强化学习框架
- Authors: Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2609.33848
- Pdf link: https://arxiv.org/pdf/2609.33848
- Abstract
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
- 中文摘要
大型语言模型(LLM)代理越来越多地承担极长(xlong)视野任务,单次执行可能持续数小时,涉及数百次模型-环境交互,每次部署需近100万个令牌。将在线强化学习(RL)应用于此类执行面临两个根本挑战:(1)严重的执行方差和延长的推展延迟导致GPU大规模空闲;(2)复杂的非线性分支产生巨大的轨迹冗余,严重削弱训练效率。为此,我们介绍了QwenGyre,一个面向xlong-horizon在线强化学习的端到端框架。QwenGyre在不中断实时执行的情况下,弹性地重新分配GPU于推展和训练之间,其轨迹处理器则重建分支历史,对部分进度进行评分,并将冗余路径去重,以限制训练成本。以我们的旗舰模型Qwen~3.8 2.4T为基础,每次推出70万令牌,QwenGyre在NL2RepoBench上实现了6.0%的绝对提升(52.5%的$兑$ 58.5%),共48步。在我们对多个训练数据集领域的评估中,QwenGyre相比Coloc和Async分别提升了1.85美元和1.78美元的时间美元。
R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution
R$^2$ 流:通过递归技能进化实现递归自我提升
- Authors: Mingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang, Qika Lin, Xiaoying Tang, Tiesunlong Shen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.33867
- Pdf link: https://arxiv.org/pdf/2609.33867
- Abstract
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at this https URL.
- 中文摘要
基于LLM的代理可以通过重用和修订他们编排的技能到可执行过程来提升自身。基于流程的训练符合这一循环:它按奖励比例采样过程,每个技能的流程会将其归功于下一个库的修订。让这种自我提升的可靠性面临三个障碍:流程训练在树状结构历史上容易出现策略崩溃;非负的基于流程的积分奖励频繁使用,仿佛是受益;而库的编辑则依赖于策略优化的任务奖励。我们引入了R$^2$ Flow,一种递归自我提升框架,交替在共享状态编排图上进行策略学习、独立验证和版本化技能库更新。该图合并了仅在独立步骤顺序上不同的历史,使流训练能够在等效执行间汇集证据。一个对训练流程的流分读出,对后退策略保持不变,以及一个独立的带名效用秩,说明哪些技能需要更改,验证者证据决定是否需要编辑,剩余方差平台则决定何时更新。承诺的编辑会重塑下一策略学习的图,实现递归技能演进。在问答、数学推理、交互式决策和代码生成等方面,R$^2$ Flow提升了任务的准确性和库编辑的精准度,优于启发式编排、强化学习和技能演化基线,并实现执行者间的转移。代码可在此 https URL 访问。
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
反事实推广重放:可分叉环境作为软件工程代理的免费流程奖励
- Authors: Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang
- Subjects: Subjects:
Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33875
- Pdf link: https://arxiv.org/pdf/2609.33875
- Abstract
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
- 中文摘要
仅结果强化学习为软件工程(SWE)代理提供终端成功信号,但对中间决策的直接指导有限。我们引入了反事实展开重放(CRR),这是一种训练时间过程,利用可分叉的可执行环境获得步级返回对比。CRR选择一小组决策点,恢复每个状态,采样一个备选动作,并在策略下将分支向前滚动。它保留实现的训练轨迹,并用终端回报与采样反事实回报之间的差额替代选定步骤的优势。该方法无需人工过程标签或学习过程奖励模型;免费指的是监督成本,而非重放计算。采用14B策略,CRR改进了SWE-bench Verified、SWE-bench Live和SWE-rebench的 pass@1,并结合过程奖励和轨迹搜索方法。在SWE工作台Verified上,同一硬件上的等壁钟比较结果为41.7%,而仅限扩展结果的GRPO为36.7%,含分叉开销后提升5.0点。这些结果适用于具有经济且可靠状态恢复的环境;随机延续和昂贵或不完美的重放仍是局限。
DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation
DexTaG:灵巧操作强化学习中的触觉指导
- Authors: Han Yang, Yian Wang, Yunlong Song, Zhenjia Xu, Chuang Gan
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.33882
- Pdf link: https://arxiv.org/pdf/2609.33882
- Abstract
Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving in-hand reorientation. Prior work bridges this gap in simulation through reinforcement learning (RL) or trajectory optimization, but the human contact pattern is hard to preserve under such formulations, often producing unnatural manipulation and unstable functional grasps. These methods also train a separate policy or solve a separate optimization for each reference trajectory, which is inefficient. To solve these problems, we propose DexTaG, a tactile-guided RL framework for dexterous manipulation. During training, tactile signals captured by the glove guide policy search toward the measured human contact pattern, reducing reliance on precise reference geometry for contact supervision. To improve efficiency, we train a single generalizable retargeter jointly on all training trajectories of the same object. The retargeter is further distilled into a tactile-free student controller conditioned on the target object trajectory for real-world deployment. On marker-pen and hammer manipulation tasks, DexTaG learns natural, contact-rich behaviors that baselines with distance-based contact heuristics fail to learn, generalizes to held-out trajectories of the same object and task, and outperforms single-trajectory baselines on OakInk2.
- 中文摘要
基于手套的动作捕捉正逐渐成为一种可扩展的灵巧手演示数据收集方法。然而,由于人手与机器人手之间的运动学差距,记录的人类动作无法直接在机器人上执行,尤其是在涉及手部重新定向的接触丰富工具使用任务中。以往的研究通过强化学习(RL)或轨迹优化弥合了这一模拟空白,但在此类表述下,人类接触模式难以保持,常导致不自然的操作和功能抓握不稳定。这些方法还会为每个参考轨迹训练独立策略或解决独立优化,效率较低。为解决这些问题,我们提出了DexTaG,一种触觉引导的灵巧操作强化学习框架。在训练过程中,手套引导策略捕获的触觉信号会搜索到测量到的人际接触模式,减少对精确参考几何的依赖。为了提高效率,我们联合训练一个可推广重定向器,针对同一物体的所有训练轨迹。该重定向器进一步简化为无触觉的学生控制器,基于目标物体轨迹,用于实际部署。在记号笔和锤子操作任务中,DexTaG学习了基于距离的基线接触启发式无法学会的自然、接触丰富的行为,推广到同一物体和任务的持续轨迹,并且在OakInk2上优于单一轨迹基线。
No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
无免费效率:重新审视训练效率与模型脆弱性之间的权衡
- Authors: Yiyong Liu, Jun Sakuma, Michael Backes, Rui Wen
- Subjects: Subjects:
Machine Learning (cs.LG); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2609.33898
- Pdf link: https://arxiv.org/pdf/2609.33898
- Abstract
Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified reinforcement learning pipelines. While these strategies drastically reduce overhead, they prompt a critical, yet neglected question: Is efficiency achieved at the expense of model robustness and security? To our knowledge, we present the first systematic cross-domain investigation of the efficiency-vulnerability trade-off. Across vision and language models, we show that efficiency-oriented training increases susceptibility to adversarial and privacy attacks. We characterize this vulnerability by analyzing the models' internal geometry and functional representations, demonstrating that the evaluated efficient variants consistently exhibit sharper loss geometry together with systematic changes in representational structure. We further extend our analysis to "zero RL training", finding that models trained using simplified RL recipes exhibit substantially greater susceptibility to catastrophic forgetting and more pronounced overconfidence than those trained through conventional alignment pipelines. Our findings suggest that training efficiency is rarely a "free lunch"; rather, the mechanisms that minimize computation can inadvertently compromise safety. We conclude by calling for a paradigm shift toward multi-objective training that jointly optimizes for performance, cost, and security.
- 中文摘要
训练效率已成为基础模型近期进展的核心驱动力。为克服大规模训练庞大的计算和数据需求,研究人员越来越多采用选择性数据采样、高效预训练和简化强化学习流程等策略。虽然这些策略大幅降低开销,但也引发了一个关键但被忽视的问题:效率是否是以牺牲模型的稳健性和安全性为代价实现的?据我们所知,我们首次系统地跨域研究了效率与脆弱性权衡。在视觉和语言模型中,我们表明效率导向训练增加了遭受对抗性和隐私攻击的脆弱性。我们通过分析模型的内部几何和函数表征,来描述这一脆弱性,证明被评估的高效变体持续表现出更锐利的损失几何,同时表示结构也发生系统性变化。我们进一步将分析扩展到“零强化学习”,发现使用简化强化学习配方训练的模型比传统对齐流程训练的模型更易发生灾难性遗忘和更明显的过度自信。我们的发现表明,训练效率很少是“免费午餐”;相反,最小化计算的机制可能无意中损害安全性。我们最后呼吁实现范式转变,转向多目标训练,共同优化性能、成本和安全。
ICMAPE: In-Context Multiagent Pure Exploration
ICMAPE:上下文多智能体纯探索
- Authors: Xinyi Hu, Alessio Russo, Aldo Pacchiano
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33986
- Pdf link: https://arxiv.org/pdf/2609.33986
- Abstract
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.
- 中文摘要
在某些多智能体系统中,需要优化的数量不是外部指定的奖励,而是通过主动顺序假设检验(ASHT)问题获得的关于环境未知属性的信息。然而,ASHT文献往往侧重于具有明确定义模型的有限单智能体问题,而目前对于能够执行主动顺序测试的实用多智能体方法存在空白。我们用ICMAPE填补了这一空白,ICMAPE是一个基于贝叶斯学习的框架,用于由推理目标驱动的去中心化多智能体纯探索。ICMAPE将固定置信度识别目标转换为由推理置信度导出的奖励,使标准强化学习机制能够应用于去中心化纯探索。它共同学习一个集中式神经推断网络,从全球轨迹数据中估算假设的后验分布,以及去中心化策略,从局部观测历史中选择行动,并学习在达到目标置信度后何时停止收集数据。在两个合成基准测试和基于真实世界数据的马里兰硝酸盐浓度监测任务中,ICMAPE-TD3以更少的探索步骤实现目标精度。
RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
奖励解释器:从反事实偏好反馈中学习奖励模型解释
- Authors: Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du, Xia Hu
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33989
- Pdf link: https://arxiv.org/pdf/2609.33989
- Abstract
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.
- 中文摘要
奖励模型(RM)是大型语言模型后训练中的关键组成部分,为后续强化学习提供奖励信号。然而,传统的判别性RM通常只输出标量分数,这使得识别与评分决策相关的反应行为变得困难。现有的解释方法通常依赖预定义的高层属性,并且需要对每个反应对反事实干预来验证候选解释,缺乏利用RM反馈来训练可重复使用的解释器的闭环机制。为此,我们提出了RewardExplainer框架,该框架通过反事实重写从目标奖励模型中获取反馈,并利用这些反馈进一步优化解释器。RewardExplainer生成开放式、原子化且可交互的自然语言评分机制,使解释更具体、易读且可操作。它进一步将反事实反馈转化为偏好监督,使解释器比单次生成更忠实地捕捉目标 RM 的评分偏好和敏感行为。跨多个目标 RM 和解释者骨干的广泛实验显示出持续的改进。除了解释外,我们还利用生成的机制识别潜在偏差模式,构建针对性的去偏数据以微调奖励模型,提升奖励黑客基准的稳健性。
ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim
ZeroBot:通过生成式Real2Sim几分钟内从零开始学习
- Authors: Ivan Kapelyukh, Xiaohan Zhang, Stephen James, Laura Herlant, Edward Johns
- Subjects: Subjects:
Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.34010
- Pdf link: https://arxiv.org/pdf/2609.34010
- Abstract
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
- 中文摘要
我们介绍了ZeroBot,一个real2sim框架,可在零人工演示、零策略预训练和零已知物体模型的条件下,几分钟内从零开始学习机器人操作任务。在仅给定一个对象的单一视图和目标姿态的情况下,ZeroBot利用图像到三维生成模型获得完整的物体网格,用于大规模并行强化学习的仿真。为加快训练,我们引入了一个动作空间,利用生成的几何和学习到的值函数采样涉及机器人与物体接触的状态。在实际任务中评估,包括抓取、推动、关节物体交互和多阶段操作,ZeroBot实现了87%的成功率,平均训练时间为119秒。这些结果展示了在real2sim框架中使用图像到三维模型在快速自主机器人学习中的价值。
Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
视觉——受限强化学习中的语言信号:安全优势而无预期
- Authors: Samuel Tetteh, Cody Fleming
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34041
- Pdf link: https://arxiv.org/pdf/2609.34041
- Abstract
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6\% to 19.4\%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.
- 中文摘要
安全强化学习寻求在满足安全约束的同时最大化任务表现的策略。然而,在驾驶基准测试中,碰撞成本通常仅在碰撞时出现,无法提前预警接近的危险。冻结视觉语言模型可以提供密集的语义反馈,但其评分是否预见碰撞以及哪个组件推动安全改进仍不明确。情节成本也可能有利于任务进展缓慢的策略。为弥补这些空白,我们提出了VLM-Safe-RL,该框架通过奖励塑形和增强乘数更新将冻结的CLIP信号整合为PPO-拉格朗日量。在结合最密集流量和最大地图的MetaDrive Hard中,灾难率从31.6%降至19.4%。FormulaOne-L2分析未发现CLIP信号预见碰撞的证据,且VLM项对拉格朗日乘数的影响微乎其微。这些发现显示了灾难率的有条件下降,但无碰撞预期的证据。
Learning Perturbation Robust Policies for LLM Agents with Stable Optimization
学习具有稳定优化的LLM代理的扰动稳健策略
- Authors: Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang, Jingdi Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34064
- Pdf link: https://arxiv.org/pdf/2609.34064
- Abstract
Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve perturbation robustness during policy optimization. We first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates preserve stable monotonic improvement. Based on this analysis, we introduce Stable Perturbation-Robust Policy Optimization (SPrPO), which applies adaptive and sensitivity-aware perturbations during RL training. We evaluate SPrPO on ALFWorld and WebShop and conduct systematic experiments across multiple perturbation types and scales, showing improved perturbation robustness while maintaining stable policy optimization.
- 中文摘要
强化学习(RL)已成为长视野大型语言模型(LLM)代理的有效训练后范式。然而,我们发现所得策略对隐藏状态噪声、剪枝和量化等多种策略扰动具有敏感性。本研究中,我们研究如何在策略优化过程中提升扰动鲁棒性。我们首先引入了扰动鲁棒策略的概念,并分析扰动策略更新在何种条件下保持稳定的单调改进。基于该分析,我们引入了稳定扰动-鲁棒策略优化(SPrPO),在强化学习训练中应用自适应且敏感度感知的扰动。我们在ALFWorld和WebShop上评估SPrPO,并系统地在多种扰动类型和尺度上进行实验,显示扰动鲁棒性提升,同时保持策略优化的稳定。
PhysFieldBench: Can Multimodal Models Understand Physical Fields?
PhysFieldBench:多模态模型能理解物理场吗?
- Authors: Yuezhou Ma, Huikun Weng, Jialong Wu, Chenyi Zhao, Hang Zhou, Haonan Shangguan, Jianmin Wang, Mingsheng Long
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34072
- Pdf link: https://arxiv.org/pdf/2609.34072
- Abstract
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.
- 中文摘要
多模态大型语言模型(MLLM)越来越被视为科学和工程代理的核心组成部分,但它们解释物理场的能力仍然了解不足。现有物理基准大多强调教科书式的问题解决或直观的物理推理,导致MLLM是否能从连续的现场观察中推断出具有物理意义的信息。我们介绍PhysFieldBench,这是一个包含24个任务和1160个评估示例的基准测试,涵盖受控方程场、模拟物理场和观测物理场。这些任务评估三种推断形式:识别物理机制、比较潜在控制变量和预测结果属性。在代表性的开源和专有MLLM中,零样本性能较低:最佳模型的概率归一化得分为29.3,而部分开源模型仍保持接近偶然性。相比之下,任务特定监督视觉转换器表现显著更好,证明输入包含可学习的物理信息。为诊断这些失败,结构化自我解释分析将大多数错误归因于视觉模式遗漏和视觉到物理映射错误。进一步,为探讨后训练是否能提升物理推断并推广到看不见任务,我们比较监督微调与最终答案或思维链监督与强化学习。最终答案监督整体表现最佳,但转移效果较差,而思维链监督后的强化学习则实现最佳泛化。综合而言,这些发现强调了MLLM在科学和工程工作流中可靠解释物理场时,需要改进视觉到物理的基础和跨任务推广。
Evolution of fairness in multi-objective reinforcement learning framework
多目标强化学习框架中公平性的发展
- Authors: Jingyi Zhang, Xin Ou, Guozhong Zheng, Shengfeng Deng, Jiqiang Zhang, Li Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn)
- Arxiv link: https://arxiv.org/abs/2609.34114
- Pdf link: https://arxiv.org/pdf/2609.34114
- Abstract
Fairness, as a fundamental social norm, continues to pose a longstanding puzzle regarding its emergence. Traditional game-theoretic models largely rely on the assumption of \emph{Homo economicus}, wherein individuals are purely rational and self-interested, acting solely to maximize material payoffs. Such accounts, however, overlook the multidimensional nature of human decision-making, which is often shaped also by other considerations beyond economic incentives. To address this gap, we propose a multi-objective reinforcement learning framework that models the evolution of fairness as a dynamic trade-off between material payoff maximization and fairness-driven moral behavior, regulated by a fairness pressure coefficient. Using simulations of a two-objective Q-learning ultimatum game, we find that increased fairness pressure promotes fair outcomes, as expected. Strikingly, however, under moderate pressure, responder behavior reverses: responders become ``forgiving" by accepting low offers -- a pattern in line with our daily experience. Microscopic analyses reveal that this strategy reversal stems from competition between payoff-maximizing and fairness-oriented preferences. We further extend our framework to an asymmetric setting, where proposers and responders assign different weights to the two objectives. Overall, our work expands the reinforcement learning paradigm from a single-objective to a multi-objective formulation, offering a versatile tool for elucidating a broader range of human social behaviors.
- 中文摘要
公平作为一种基本的社会规范,至今仍是一个长期以来关于其出现的谜团。传统的博弈论模型主要依赖于\emph{经济人}的假设,即个体纯粹理性和自利,仅为最大化物质回报而行动。然而,这些说法忽视了人类决策的多维本质,而这种决策往往还受经济激励之外的其他因素影响。为解决这一差距,我们提出了一个多目标强化学习框架,将公平的演化建模为物质收益最大化与公平驱动的道德行为之间的动态权衡,并由公平压力系数调节。通过模拟一个双目标Q学习最后通牒游戏,我们发现增加的公平压力促进了公平结果,正如预期的那样。然而,令人惊讶的是,在适度压力下,回应者行为会逆转:回应者通过接受低报价变得“宽容”——这与我们的日常经验相符。显微分析显示,这种策略逆转源于收益最大化偏好与公平取向偏好之间的竞争。我们进一步将框架扩展到非对称环境,提出者和回应者对这两个目标赋予不同权重。总体而言,我们的工作将强化学习范式从单一目标扩展为多目标,提供了阐明更广泛人类社会行为的多功能工具。
Learning to Optimize through Solver-Grounded Self-Play
通过解算器基础的自我游戏学习优化
- Authors: Xia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu, Wim P.M. Nuijten, Yingqian Zhang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34205
- Pdf link: https://arxiv.org/pdf/2609.34205
- Abstract
Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This dependence introduces a Generalization Ceiling, where models overfit to narrow data distributions, and Capability Anchoring, where models' reasoning is bounded by annotator proficiency and teacher model capability. In response, we propose OPT-Zero, the first fully self-play training framework for optimization modeling that requires zero external training data. OPT-Zero employs a single LLM in a dual-role closed loop: a Proposer that synthesizes increasingly challenging optimization problems alongside their mathematical formulations and solving code, and a Solver that attempts to resolve the problems given only natural-language problem descriptions. Grounded in execution feedback from external optimization solvers, we alternately train both roles using reinforcement learning. This process fosters an auto-curriculum in which the Proposer and Solver co-evolve: generating harder valid problems by the Proposer seamlessly enhances the structural reasoning ability of the Solver. Extensive results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
- 中文摘要
优化建模在许多决策场景中至关重要,但传统上需要丰富的领域专业知识。虽然大型语言模型(LLM)在自动化这一过程方面展现出潜力,但当前的训练范式主要依赖人工注释或教师生成的数据集。这种依赖引入了泛化上限,即模型对狭窄数据分布的过度拟合,以及能力锚定,即模型的推理受注解员熟练度和教师模型能力的限制。为此,我们提出了OPT-Zero,这是首个完全自玩的优化建模训练框架,无需任何外部训练数据。OPT-Zero采用单一LLM进行双重闭环:一个提案者,综合越来越难的优化问题及其数学表述和求解代码,另一个求解器,试图解决仅给出自然语言问题描述的问题。基于外部优化求解器的执行反馈,我们交替使用强化学习训练这两个角色。这一过程促进了自律课程,提议者与求解者协同进化:提案者无缝生成更难的有效问题,增强求解者的结构推理能力。大量结果表明,在零策划数据的情况下,OPT-Zero能够匹配最先进的数据依赖方法,同时展现出显著更强的泛化性,确立了自玩训练作为推动LLM推理建模和解决优化问题的高度可扩展范式。
GlyphBench: A Playground for Language-Model Reinforcement Learning
GlyphBench:语言模型强化学习的游乐场
- Authors: Roger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker, Glen Berseth, Pablo Samuel Castro
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34214
- Pdf link: https://arxiv.org/pdf/2609.34214
- Abstract
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.
- 中文摘要
我们介绍了 GlyphBench,一套用于语言模型代理后训练强化学习(RL)的环境套件,包含涵盖多样游戏的 360 多个任务。GlyphBench 将空间观测渲染为二维 Unicode 网格,并通过统一界面连接训练、评估和轨迹重放,以支持高效且可重复的研究。我们使用 GlyphBench 研究观察接口、推理努力和代理工具如何影响性能,以及强化学习配置如何塑造学习动态。我们的结果显示,字形观测在 Craftax 实验中优于原生文本和像素,并在多个 BALROG 环境中取得进一步提升。在 100 个 GlyphBench 任务中,强化学习提升了 Qwen3.5-4B 在未完成的 Reasoning Gym 问题上,准确率达到 63.48%,并优于基础模型、数学训练的基线和代码训练的基线。这些实验提供了实证证据,表明游戏性推理带来的提升比数学或代码更强的迁移。这些结果共同凸显了 GlyphBench 作为系统研究语言模型代理学习、交互和泛化的试验平台的价值。
MAS-OPD: On-Policy Distillation for Multi-agent Systems
MAS-OPD:多智能体系统的策略提炼
- Authors: Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su, Huwei Ji, Li Zhang, Junfeng Fang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.34234
- Pdf link: https://arxiv.org/pdf/2609.34234
- Abstract
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.
- 中文摘要
多智能体系统(MAS)将任务分配到不同专业角色,在复杂任务上前景看好,但主流方法仅依赖推理时间编排。通用API成本高且难以定制,而带有角色提示的小型模型很少发展稳定的角色能力或可靠协作,因此联合训练MAS是核心。大多数尝试使用强化学习,其团队级奖励不确定哪个代理的哪个步骤产生了结果,而局部奖励则需为每个任务重新设计。策略提炼(OPD)为学生样本的轨迹提供代币级教师监督,信号更密集且无需局部奖励,但对MAS中相互依赖的代理却未被充分探索。出现两个难题:从判断行为归属角色构建互补专业化,同时保留所有角色所需的知识;以及将跨代理协作信息转化为OPD可利用的监督。我们介绍MAS-OPD,其中角色优势专精定义角色优势为教师在目标与非目标角色条件下信号的差异,协调特权归因则将互动冲突归因于其源头,并仅作为特权信息提供给教师。对代码和数学基准的广泛实验表明,MAS-OPD在学生两个尺度上均得分最高,并促使代理发展出更清晰的角色专精和更有效的协作行为。
AdaGuard: An Adaptive Guard Model with User-defined Policies
AdaGuard:一个带有用户自定义策略的自适应保护模型
- Authors: Yunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li, Mingrui Lao, Zeyuan Wang, Yanming Guo
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34241
- Pdf link: https://arxiv.org/pdf/2609.34241
- Abstract
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at this https URL
- 中文摘要
守护模型支持语言模型代理的安全部署,但固定风险分类法限制了其适应不同应用和任务需求的能力。在用户定义的策略下,检测违规需要同时解读适用规则和代理行为,因为相同行为在不同策略下可能获得不同判断。为支持学习这一能力,我们引入了AdaptiveSafety数据集,包含10,939个训练示例和1,000个测试示例,涵盖1-100条规则的策略。该数据集结合了来自多个来源的轨迹与策略和行为反事实,每个示例配有解释和完整的违规规则集合。这些反事实揭示了改变合规性的变更,而结构增强则为规则重排序和标识符重映射的一致性提供监督。基于该监督,我们提出了SafePO,一种强化学习算法,用于优化违规识别,同时平衡解释性推理与最终判决。SafePO使用结构化奖励评估预测正确性,保留响应层级的群体相对优势,并采用独立训练的价值模型调制解释和裁决区域内的标记权重。分离归一化控制它们在训练中的相对贡献,尽管长度存在差异。通过监督初始化和SafePO,我们开发了AdaGuard,这是一类0.6B、4B和8B防护模型,用于在推断时评估策略下的代理轨迹。我们的4B模型在AdaptiveSafety上实现了89.30%的二进制准确率,在DynaBench上为71.82%。项目仓库可在此 https URL 访问
SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
SALMONN-duo:全双工语音代理的自适应双系统协调
- Authors: Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Chao Zhang
- Subjects: Subjects:
Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- Arxiv link: https://arxiv.org/abs/2609.34247
- Pdf link: https://arxiv.org/pdf/2609.34247
- Abstract
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $\tau$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $\tau$-Voice with an acceptable increase in the delegation rate.
- 中文摘要
全双工语音大语言模型(LLMs)实现低延迟、自然的语音交互。然而,现实世界的代理还必须使用工具并执行深思熟虑的推理操作,其可变的延迟和计算成本与实时对话严格的时序要求相冲突。为调和这些需求,我们提出了SALMONN-duo,一种受认知双进程理论启发的自适应双系统语音代理。SALMONN-duo通过将始终在线、思维快速的全双工语音LLM(系统1)与强大的异步慢思考LLM代理(系统2)结合,将实时交互与深思计算区分开来。除了处理实时交互外,系统1还能学习何时直接回答、何时委派,在后端执行中保持响应,并将返回的信息无缝整合进持续对话,同时不暴露工具痕迹或失去对话上下文。单回合口头问答(QA)和多回合对话的评估表明,自适应委派显著提升了知识密集型和多跳推理问题的准确性,而知识边界感知训练则避免了不必要的系统2调用。在定制版的$\tau$-Voice上,SALMONN-duo进一步展示了其通过多回合交互完成环境基、策略约束任务的能力,适用于真实的业务场景。最后,成本意识强化学习进一步提升了任务性能与后端使用率在QA和对话任务中的权衡,同时通过可接受的委派率提升,提升了$\tau$-Voice的任务成功率和响应安全性。
Evolving Support Priorities in Empathetic Reinforcement Learning
同理强化学习中支持优先事项的演变
- Authors: Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34249
- Pdf link: https://arxiv.org/pdf/2609.34249
- Abstract
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
- 中文摘要
我们发现了同理强化学习中的根本不匹配:支持优先级随着对话状态演变,而现有方法通常优化预设奖励规格,且这些奖励在回合中保持固定。为建模这些演变的支持优先级,我们将同理心支持组织为认知、情感和主动同理心,并提出了情境自适应评分标准演化(CARE)。在每个环节,CARE通过调整这三个同理心维度的权重及其细粒度评估标准,生成情境适应性评分标准。评分标准生成器通过回合级评分标准监督和人类偏好数据进行监督微调训练,随后进行基于偏好的强化学习,随后作为在线同理心强化学习的自适应奖励接口。CARE与RLVER和MICA集成,在三位独立LLM评审下,在SentientBench、EQBench3和EMPA中实现了最先进的性能。值得注意的是,在EMPA测试中,CARE在三位评审下均提升了EPM-Idx至少13分,包括Gemini-2.5-Pro的28.11提升至83.54分。进一步分析显示,学习后的评分标准优先级在对话阶段和用户情绪中系统性变化,表明CARE会随着支持需求的变化调整奖励内容。
Rotated Manifold Optimization for Low-Rank Adaptation
低秩适应的旋转流形优化
- Authors: Yuhui Ding, Javier Zazo, James Hensman
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34264
- Pdf link: https://arxiv.org/pdf/2609.34264
- Abstract
We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpreting them as normalization under a rotated basis. We show how rotation and normalization can be integrated with the fixed-rank manifold efficiently. Our optimizer converges faster to lower held-out loss and achieves better or comparable downstream performance on both supervised finetuning and reinforcement learning tasks.
- 中文摘要
我们提出了一种新颖的低秩适应优化器(LoRA),明确包含低秩分解的规范对称性。我们的优化器将最新的全参数训练矩阵优化器扩展到固定秩矩阵流形,将其解释为旋转基下的归一化。我们展示了如何高效地将旋转和归一化与固定秩流形整合。我们的优化器收敛速度更快,降低保留损耗,并在监督微调和强化学习任务中实现更好或相当的下游性能。
ZeroCode: On-demand Error-Correcting Code Construction from the Zero Matrix via Reinforcement Learning
ZeroCode:通过强化学习从零矩阵构建按需纠错代码
- Authors: Ju-Hyeong Lee, Yongjune Kim, Sang-Hyo Kim, Dae-Young Yun, Hee-Youl Kwak
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34265
- Pdf link: https://arxiv.org/pdf/2609.34265
- Abstract
Error-correcting codes (ECCs) are essential across diverse applications, from wireless communications and storage to quantum computing, yet each application imposes distinct design requirements on the parity-check matrix (PCM). To address these on-demand requirements in a unified framework, we propose ZeroCode, a reinforcement learning (RL)-based approach that constructs PCMs sequentially from the all-zero matrix. ZeroCode formulates construction as a discrete sequential decision-making problem and uses proximal policy optimization with action masking to select valid edges. ZeroCode achieves a gain of approximately 1 dB over the prior RL-based construction method at a bit error rate (BER) of $10^{-4}$ for the (32,16) code and outperforms existing genetic, differentiable, and classical code-design methods in our experiments. Beyond optimizing decoding performance, the masking mechanism allows on-demand structural constraints, such as a maximum degree, 4-cycle-free structure, and quasi-cyclic structure, to be flexibly incorporated. Moreover, a single policy rollout yields a library of PCMs with varying edge counts, offering trade-offs between decoding performance and complexity without retraining. Overall, ZeroCode addresses diverse code-design requirements within a unified framework, providing solutions with optimized decoding performance under given constraints.
- 中文摘要
纠错码(ECC)在无线通信、存储到量子计算等多种应用中至关重要,但每个应用对奇偶校验矩阵(PCM)都施加了不同的设计要求。为了在统一框架中满足这些按需要求,我们提出了ZeroCode,这是一种基于强化学习(RL)的方法,从全零矩阵顺序构建PCM。ZeroCode将构造表述为离散顺序决策问题,并利用带有动作掩蔽的近端策略优化来选择有效边。ZeroCode在(32,16)代码的比特错误率(BER)为$10^{-4}$时,比之前基于RL的构建方法获得约1 dB的增益,并且在实验中优于现有的遗传、可微和经典代码设计方法。除了优化解码性能外,掩码机制还允许灵活地整合按需结构约束,如最大度数、4周期无结构和准循环结构。此外,单一策略推出即可生成一个边缘计数不同的PCM库,在解码性能与复杂度之间提供权衡,无需重新训练。总体而言,ZeroCode在统一框架内满足多样化的代码设计需求,在特定约束条件下提供优化解码性能的解决方案。
See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
参见、测量与推理:学习病理学中的视觉基础推理
- Authors: Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Xinyu Liu, Jiaming Yang, Jie Chen, Zhang Zhang, Yuhao Yi, Hong Bu, Jiancheng Lv
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34277
- Pdf link: https://arxiv.org/pdf/2609.34277
- Abstract
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
- 中文摘要
病理评估依赖于识别组织图像中的细粒度视觉细节。视觉语言模型(VLMs)越来越支持病理解读,但其感知这些细节的能力仍然不足。这一弱点导致细胞观察不准确,即使最终答案正确也可能持续存在。本文提出通过明确监督细胞外观和丰度来提升视觉基础推理。ASPECT 通过病理特征重建、细胞特征比对和计数监督训练中间视觉符号。三阶段监督式微调教授模型感知、生成视觉符号和推理,随后进行强化学习,奖励报告的测量结果的正确性和一致性。我们还引入了 PathoVernier,这是一个由五个病理数据集中 759 个专家评审问题组成的基准测试,涵盖四个细胞组成任务。它评估最终答案和中间测量,以揭示答案准确性隐藏的误差。在PathoVernier上,ASHD相较最强基线Gemini-3.1-Pro实现了约19.2%的相对准确率提升,Qwen3-VL-8B骨干链提升了99.3%,同时将测量正确反应中错误计数的RAWR降低了28.1%和42.7%。ASPECT 在涵盖分类和问答的三个外部病理基准上也相较于其骨干有所提升,涵盖了细胞组成任务之外的分类和问答。
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
博士学分:深度研究代理的评分标准基础过程学分作业
- Authors: Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang, Shen Huang, Xunjie Zhu, Pengjun Xie, Shiming Xiang
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34296
- Pdf link: https://arxiv.org/pdf/2609.34296
- Abstract
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. this http URL uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that this http URL outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
- 中文摘要
基于评分标准的任务越来越多地通过强化学习(RL)来解决,评分标准分数作为训练奖励。然而,这些奖励通常监督最终答案,而不区分中间决策的贡献。许多现有的学分分配方法依赖基于实地的答案来定义过程奖励,限制其适用范围为无规范解法的开放式任务。为解决这一限制,拟议的基于评分标准的学分以任务需求作为最终答案评估和过程监督的共享参考。工具返回的信息会根据该评分标准的接受支持历史,评估其满足每个评分标准的额外支持程度。通过引用这些历史,信用区分了新的支持与轨迹中已有的证据,同时认可每个评分标准的部分支持。该HTTP URL利用基于评分标准的信用,在强化学习框架中监督深度研究代理的中间工具轮换。由此产生的流程优势与GRPO结果优势相结合,指导研究决策,同时保持对最终报告质量的监督。对四个领域内外基准的评估显示,该http URL在所有主要指标和子指标上均优于已评估的开放深度研究基线。同时,采用8B参数骨干,训练有素的代理在与评估前沿专有模型中均有竞争力。进一步分析表明,在有限的研究预算下,证据获取效率更高,报告质量更高,促使基于评分标准的过程监督扩展到更广泛的基于评分标准的任务。
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
知道何时思考还不够:教导小推理模型超越其参数化知识
- Authors: Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.34327
- Pdf link: https://arxiv.org/pdf/2609.34327
- Abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.
- 中文摘要
测试时计算的扩展是提升语言-模型推理的强大方法,尤其适合服务成本低廉的小型推理模型(sRM)。然而,额外的思考是否总是正确的操作?通过在两个模型家族和多尺度的中间推理状态进行干预,我们发现自我精炼在很大程度上将概率质量集中到当前状态已可达的解上,而非让新的解决方案变得可达。这些干预揭示了两种失败模式:执行瓶颈,即正确路径可达且通过反思能恢复;以及知识瓶颈,即相关外部信息使其可达。基于这一区别,我们引入了FlyBy,一种选择性查询框架,训练4B和8B变体,先推理,诊断未解决问题,并在知识瓶颈时查询参数知识超越自身的强大模型。监督微调引导执行多深度查询动作,成本感知强化学习校准是否查询、问什么以及花费多少。在六个基准测试的1158个难题中,FlyBy-4B实现了45.96%的难pass@8,超过Qwen3-14B(41.64%),服务成本低2.7倍,同时在pass@1上也超过Qwen3-8B(16.85%对15.31%)。扩展到FlyBy-8B进一步提升pass@8达到51.81%。
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
学习引导,引导观察:通过可训练向量揭示大型语言模型中RLVR的几何结构
- Authors: Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34344
- Pdf link: https://arxiv.org/pdf/2609.34344
- Abstract
Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: this https URL
- 中文摘要
强化学习(RL)已成为提升大型语言模型推理的关键范式,但参数更新的高维度使其训练动态难以分析。我们研究可验证奖励强化学习(RLVR),并利用向量引导识别与激活空间中与强化学习诱导增益相关的低维有效流形。我们揭示了两个几何性质。(1)有效流形容量:重现强化学习增益所需的容量可以非常小,但不可无限压缩;在极低容量下,干预维度和依赖输入的表达性成为关键约束,且该要求随注入深度变化。(2)控制流形分离:有效控制方向主要存在激活主子空间的低方差补空间。在任务和基础模型中,学习的几何在训练配置间基本保持一致,且任务间几何对齐与能力转移相关。对5个大型语言模型和6个可验证奖励任务的实验支持了这些发现。随后,我们提出了Alpha-Stabler,一个即插即用框架,带有预测器,用于监测主子空间入侵以识别早期崩溃警告,以及一个控制器,在反向传播过程中去除激活梯度的主子空间分量,同时保持正交补集。Alpha-Stabler稳定训练可达2000步,持续提升强化学习的收益,为稳健的后训练提供实用见解。代码:此 https URL
Improving Large Language Models for Code through Runtime Program-State Reasoning
通过运行时程序状态推理改进代码大型语言模型
- Authors: Hongwei Li, Spandan Garg, Yufan Huang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34359
- Pdf link: https://arxiv.org/pdf/2609.34359
- Abstract
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
- 中文摘要
大型语言模型在对运行时程序状态的推理中获得有限的显式训练。我们研究训练模型推理运行时程序状态是否能提升下游软件工程能力。我们引入了两个互补的程序-状态推理任务。有缺陷的输入输出推理需要模型生成一个具体输入,揭示有缺陷程序与隐藏正确实现之间的行为差异,并预测最终的执行行为。前置条件-后置条件推理要求代理符号化触发错误的前置条件,预测预期的后置条件,解释其因果关系,并将推理实例化为可执行的回归测试。通过将这两种任务整合进分阶段的训练后流水线,我们开发了Comet-9B,这是一个基于Qwen3.5-9B Base的9B语言模型。我们评估了存储库级补丁生成、回归测试生成和安全PoC生成的检查点。将两个程序-状态推理任务加入问题解决的监督微调(SFT),在SWE-bench Pro上提升了7.25个百分点,在SWT-Bench Verified上提高了9.70个百分点。对这两个任务进行顺序强化学习,在SWE-bench Pro、SWT-Bench Verified和CyberGym上分别提升了7.25、26.79和4.67个百分点。尽管只有9B参数,Comet-9B的得分与SWE-bench Pro报告的GPT-5.2结果相当,并且与基于GPT-4o的代理在SWT-Bench Verified上报告的成功率相当。
Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
涌现,而非带宽:物理耦合与多智能体通信学习的局限性
- Authors: Mihir Chauhan, Aniket Bera
- Subjects: Subjects:
Machine Learning (cs.LG); Multiagent Systems (cs.MA); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.34373
- Pdf link: https://arxiv.org/pdf/2609.34373
- Abstract
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
- 中文摘要
速率限制多代理团队提出了三个问题,涌现通信文献仅通过实证方法回答了:最优消息应编码什么、压缩成本、以及何时学习的协议足够独特,队友能读取。我们回答速率限制的Dec-POMDPs,然后测量强化学习落后最优的程度。我们的定理固定了独立于任何学习者可实现的目标,因此在相同比特预算下,工程发送者与学习发送者之间的差距是优化事实,而非信息理论。我们将此实例化在三个MuJoCo竞技场上,涵盖零、部分和刚性物理耦合,每个条件每个决策都恰好收费2位,并通过在一个竞技场内关闭物理侧信道,固定身体、任务和奖励来创建辨别机制。通信值由耦合决定:在通过共享对象的刚性耦合下,没有信道能胜过静默(+0.001 +/- 0.001,p = 0.982,n = 25),因为本体感觉已经承载了这些信息;没有耦合时,所有条件都能解决任务;在部分耦合下,工程化的2位发送端达到四分位均值1.000,但学习的发送端达到0.482,与静默无异(p = 0.400,n = 25)。在共享字母表中,带宽无法解释这一间隙。工程接收机的热启动能够定位故障:同一信道达到0.857,而冷启动时达到0.562(p < 0.001),因此既非表征性,也非维护性;强化学习未能发现该协议。跨平台节目学习的协议在各个场馆是有意义但彼此无法理解的:自玩0.980会在种子间压缩到0.144,而我们最优的对齐方案至少保留了77%的差距。所有头条结果每个场馆使用25个种子和7条已发布的基线,且按匹配率计算。
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
ABC-Align:预测驱动的对齐与自适应偏倚控制
- Authors: Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.34374
- Pdf link: https://arxiv.org/pdf/2609.34374
- Abstract
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at this https URL .
- 中文摘要
语言模型训练后常常因需要人工收集的偏好数据而成为瓶颈,而这种数据既昂贵又难以扩展。利用伪标签的AI反馈强化学习(RLAIF)方式提供了丰富的替代方案,但会引入系统性偏差,从而削弱下游对齐。近期通用的半监督方法通过少量人工标注示例校正教师偏差,但当人工注释稀少时,方差较大。为此,我们提出ABC-Align,利用丰富的伪标签信号以最小化方差,并基于人工标记子集应用轻量级、自适应的修正。校正强度在训练过程中通过插件估计的相关偏差——方差量——自动调整。在人类反馈稀少的情况下,我们通过实证证明,ABC-Align在一系列规模递增的实验中,与RLHF、DPO和GRPO的LLM对齐表现优于以往半监督基线。我们的代码可在此 https URL 获取。
Marathoner: Ultra-Long-Horizon Autonomous Intelligence
马拉松者:超长视野自主智能
- Authors: Zhang Ruiyang, Ou Jinpeng, Xie Yifan, Zhou Jingang, Pan Lirui, Guo Qingpei, Zheng Zhedong
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.34378
- Pdf link: https://arxiv.org/pdf/2609.34378
- Abstract
Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.
- 中文摘要
人类天生具备持续朝着长期目标努力的能力。面对具有挑战性的任务,人类可以连续工作数月甚至数年以完成特定目标。本文提出Marathoner,一种具备超长视野执行能力的自主代理模型。具体来说,我们提出了一个全面的训练后流程,将这一关键能力注入基础模型。在超长视野任务综合中,我们利用来自多个GitHub仓库的1000+行新代码的主要发布PR作为综合挑战性任务级数据的主要来源。此外,我们引入了多任务链,将多个生成任务串联成一个更具挑战性的任务,使得具有前沿难度的任务能够综合。在拒绝抽样微调中,我们将强教师模型与多种工具结合,生成综合任务轨迹,并在基于拒绝采样轨迹的基础模型上进行监督式微调。在强化学习方面,冷启动模型在独立沙盒中通过线束执行真实世界执行,有效促进了真正的超长视野执行能力的获得。我们还提出了一种新颖的奖励策略——后期阶段奖励奖励,明确鼓励模型在执行后期阶段执行有意义的操作。通过对包含超长视野任务的5个基准测试进行广泛评估,Marathoner在基础模型上实现了持续且显著的性能提升,甚至超过了强专有模型的性能。进一步分析显示,Marathoner能够持续工作10+小时,并执行1000+次工具调用。
CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving
CAR-VLA:自动化的复杂性感知和风险适应性推理
- Authors: Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Shi Fan, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.34387
- Pdf link: https://arxiv.org/pdf/2609.34387
- Abstract
Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: this https URL
- 中文摘要
现有用于驱动视觉-语言-行动(VLA)模型的自适应推理方法主要关注是否推理,忽视了推理应如何不同驾驶情境下的不同。我们的关键见解是,场景复杂度决定推理深度,动态风险同样关键,对决定在时间关键情境下如何推理至关重要。因此,我们提出了CAR-VLA,一种统一的驾驶VLA模型,结合场景复杂性和动态风险,指导推理的深度、紧迫性和聚焦度。CAR-VLA将四个复杂性-风险类别映射为三种推理模式:\textit{快速直觉}用于简单低风险场景中的直接轨迹生成,\textit{慢思考}用于复杂低风险场景中的刻意推理,以及\textit{反射反应},用于高风险场景中紧凑且聚焦危险的推理,无论复杂度如何。反射反应不仅仅缩短思考时间,还将推理聚焦于最关键的危险和即时的安全响应。我们通过渐进式监督学习训练CAR-VLA,将场景评估、推理模式选择和轨迹生成相结合,随后通过推理增强强化学习提升驾驶质量和推理行为。在NAVSIM v1(91.1 PDMS)、NAVSIM v2(90.3 EPDMS)和Navhard(35.0 EPDMS)上的实验展示了竞争驾驶性能。对navtest与内部高风险场景的定性比较进一步展示了风险意识推理和危害响应路径生成。本文代码将公开发布于:此 https URL
Distilling Visual Reasoning into Text Space
将视觉推理提炼进文本空间
- Authors: Wenhan Yang, Nilay Naharas, Ali Payani, Baharan Mirzasoleiman
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.34408
- Pdf link: https://arxiv.org/pdf/2609.34408
- Abstract
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
- 中文摘要
大型视觉语言模型(LVLM)在多模态推理方面展现出强有力潜力,但常常在需要超出输入图像直接可观察概念的任务上遇到困难。现有方法生成中间图像或潜在视觉符号以指导推理,但这些表示可能引入错误,并随着推理进程越来越干扰文本推理。我们提出了视觉到文本思维链提炼(V2T)框架,使LVLM能够内化视觉推理,而无需在推理时生成中间视觉表征。V2T首先通过交织的视觉和文本思维链培训教师LVLM,然后利用知识蒸馏技术,利用教师的logit和交叉熵监督,从真实文本推理中训练学生LVLM。当推理图像能够映射到原始图像时,V2T还能将教师的注意力提炼到对应区域,而地面真实边界框则可以进一步引导后续强化学习阶段。多样推理基准测试的实验表明,V2T始终优于教师和现有基线,在保留数据集上平均准确率提升14.3%,在更广泛的视觉评估套件中提升2.7%。此外,轻量化的SFT和大幅降低的强化学习使V2T的训练速度比最先进基线快42倍。
OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
OSPD:针对人格一致对话的政策自我提炼
- Authors: Rui Xu, Yikai Zhang, Aili Chen, Zicheng Zhao, Xu Yinghui, Libo Wu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34418
- Pdf link: https://arxiv.org/pdf/2609.34418
- Abstract
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
- 中文摘要
保持多回合对话中的人物形象一致性仍是角色扮演语言模型的核心挑战。外部教师的非策略提炼会导致分布不匹配,这种不匹配在对话回合间叠加,而强化学习则在主观角色忠实度中固有的奖励模糊性中挣扎。我们提出了OSPD,一种基于策略的自我蒸馏框架,同一模型在非对称信息下既作为教师又作为学生:教师获得完整的角色档案,而学生仅看到简短摘要,学生则从自身策略中生成轨迹。我们发现教师对角色扮演对话的信心表现出双峰结构---在角色批判标记处尖锐,但在泛指话语处则扩散---并引入角色意识分歧切换以匹配该结构。一种渐进式特质掩盖课程进一步推动了角色知识在语义维度上的分阶段内化。在CharacterBench、CharacterEval和SocialBench上的实验显示,OSPD相比监督式微调和多轮强化学习基线,显著提升了人格一致性,无需任何外部教师或奖励模型。
Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
编码代理记忆后训练:通过强化学习释放预训练文件操作的记忆潜力,用于长视野任务
- Authors: Lirui Luo, Kelong Mao, Heming Xia, Rongqing Li, Xinwei Yang, Luyu Chen, Kieran Wong, Yudong Guo, Xinrui Wang, Jiayin Zhu, Simiu Gu, Sulong Xu, Cong Fang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.34422
- Pdf link: https://arxiv.org/pdf/2609.34422
- Abstract
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
- 中文摘要
语言模型代理越来越多地处理交互历史超过模型活跃上下文的长视野任务。近期工作开始使用强化学习将内存控制纳入策略,通常依赖于相对短视野的领域特定训练环境中预定义的记忆工具。这种设置将学习到的记忆行为绑定到位于基础模型预训练之外的环境特定接口,必须从零学习,因此即使训练后,代理在长视野任务中仍难以使用内存。为解决这些局限性,我们引入了编码代理记忆馆(CAMG),这是一套涵盖Shop、Coding、DeepResearch和AutoResearch的长视野代理强化学习环境。除了每个环境的原生任务接口外,CAMG还提供可执行壳访问和持续化工作区,使代理能够在一集节目中创建、修改、搜索和重用文件作为内存。我们还引入了CAMG-RL,该系统在四个环境中联合训练单一策略,采用完全异步PPO,直接从下游任务奖励中学习基于文件的记忆行为,并从规模匹配的Qwen3.5模型中训练CAMG-RL-4B和CAMG-RL-9B。在SWE-bench Verified和MLE-bench Lite上,CAMG-RL-4B和CAMG-RL-9B分别与Qwen3.5-35B-A3B和Qwen3.5-122B-A10B竞争。
Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
Q-learning 因安全离线强化学习而受罚的变换器
- Authors: Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34426
- Pdf link: https://arxiv.org/pdf/2609.34426
- Abstract
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizing rewards, and adhering to the behavior regularization imposed by the offline dataset. To tackle this trilogy challenge, we propose Q-learning Penalized Transformer policy (QPT), a \emph{training--inference consistent} framework that bridges conditional sequence modeling with constraint-aware value estimation. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost, retaining strong behavior regularization. To inject explicit safety semantics during learning, we augment sequence-model training with a Q-shaped penalty using learned reward and cost Q-functions to favor high return under low constraint violation. At inference, the same Q-functions enforce the cost threshold and choose the highest-reward feasible action, closing the loop between training and deployment. We provide a principled analysis under stylized near-deterministic CMDPs, characterizing how Q-penalized conditional generation improve safety and performance. Empirically, QPT consistently outperforms strong safe offline RL baselines across 38 tasks on the DSRL benchmark, and exhibits robust zero-shot adaptation to different constraint thresholds.
- 中文摘要
本文探讨安全离线强化学习问题,涉及使用离线数据集训练策略以满足安全约束。该问题本质上具有挑战性,因为它需要平衡三个高度相互关联且相互竞争的目标:满足安全约束、最大化奖励,以及遵守离线数据集所施加的行为正则化。为应对这一三部曲挑战,我们提出了Q学习惩罚变换器策略(QPT),这是一个\emph{training--inference consistent}框架,连接条件序列建模与约束感知值估计。QPT训练一个基于轨迹上下文和目标回报/成本生成动作的Transformer策略,同时保持强行为正则化。为了在学习过程中注入显式安全语义,我们利用学习的奖励和成本Q函数,在序列模型训练中加入Q形惩罚,以有利于在低约束违规下获得高回报。在推断时,相同的Q函数执行成本阈值并选择最高奖励的可行动作,实现训练与部署之间的循环。我们在风格化近确定性CMDP下进行了原则性分析,描述了Q惩罚条件生成如何提升安全性和性能。实证上,QPT在DSRL基准测试中38项任务中持续优于强安全离线强韧强化学习基线,并展现出对不同约束阈值的鲁棒零射值适应能力。
LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
LLMs作为自适应元求解器:多策略强化学习用于工业规模优化
- Authors: Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen, Jianfeng Feng, Dongdong Ge, Yinyu Ye
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.34427
- Pdf link: https://arxiv.org/pdf/2609.34427
- Abstract
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
- 中文摘要
将基于LLM的优化从教科书规模实例扩展到现实世界的工业任务仍是一个关键的开放挑战。现有方法主要在小型、自包含的文本问题上评估,且常常承诺采用求解器集成范式,限制了其处理实际优化工作负载的规模和结构多样性的能力。本研究提出了一个实用框架,用于训练开源LLM以应对现实世界工业规模的优化。我们首先实证展示了求解器集成推理、精确组合算法和启发式搜索在不同问题结构和尺度上展现出互补优势。基于此,我们引入了策略多样强化学习(SDRL),将LLM训练为自适应优化元求解器。SDRL通过正确性门槛的层级多样性奖励,利用这种互补性,促进不同策略间及各策略内部的强有力探索,有效防止策略过早崩溃。我们还引入了混合格式训练方案,支持自包含的文本问题和基于文件的实例。在综合评估中,我们的框架在基准测试和工业规模优化任务中均优于现有的精细调优方法和前沿模型,包括DeepSeek-V4-Pro和GPT-5.5。
Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning
逃离局部视图:发现可解释多智能体强化学习的潜在概念
- Authors: Yijie Sun, Sanquan Sun, Yanda Zhu, Yuanyang Zhu, Yaohua Hu, Chunlin Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34459
- Pdf link: https://arxiv.org/pdf/2609.34459
- Abstract
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To address these challenges, we propose a novel interpretable framework, called escaping local views (ELV), which introduces semantically structured latent concepts to render policy decisions transparent. Specifically, each agent extracts low-dimensional semantic concepts from its local observation and action-observation trajectory. These concepts are jointly encoded into a contextual latent variable via a variational autoencoder (VAE), which builds a bridge between local views and global semantics. To explicitly model the decision of each agent, we employ a dual-path attention mechanism in which one module estimates the salience of individual concepts relative to the global context, while the other captures higher-order cooperative patterns with pairwise concept interactions. Furthermore, we incorporate a concept prediction module that derives an intrinsic reward from next-concept prediction errors, which incentivizes agents to explore regions of semantic novelty. Experiments in multiple environments verify that ELV not only achieves competitive performance but also explicitly provides how agents reason about their decisions.
- 中文摘要
由于多智能体强化学习中每个代理通常具有部分可观察性,高效合作具有挑战性。循环网络编码局部交互历史,但其隐藏的表示对个体决策背后信息的洞察有限。为应对这些挑战,我们提出了一种新的可解释框架,称为逃逸局部视图(ELV),引入语义结构化的潜在概念,使政策决策透明化。具体来说,每个代理从其局部观察和动作-观察轨迹中提取低维语义概念。这些概念通过变分自编码器(VAE)共同编码到上下文潜在变量中,搭建局部视角与全局语义之间的桥梁。为了显式建模每个代理的决策,我们采用了双径注意力机制,其中一个模块估算单个概念相对于全局上下文的显著性,另一个模块捕捉高阶合作模式,伴随着成对概念交互。此外,我们集成了一个概念预测模块,从下一概念预测错误中获得内在奖励,激励代理探索语义新颖区域。在多个环境中的实验验证了 ELV 不仅实现了竞争性能,还明确提供了代理如何推理决策。
Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
通过逐步预测实现的模型知情安全强化学习
- Authors: Victor Paredes, Ayonga Hereid
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.34486
- Pdf link: https://arxiv.org/pdf/2609.34486
- Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.
- 中文摘要
类人机器人承诺在杂乱、以人为中心的环境中具备多功能性,但实际部署需要原则性安全性。经典基于模型的步态生成器能产生可解释的动作,但通常缺乏现代强化学习(RL)方法的鲁棒性和适应性。我们提出了一个基于分析性角动量线性倒摆(ALIP)模板的模型驱动强化学习框架。我们通过离散指数控制障碍函数(DECBF)提供ALIP步进的分步安全证书,并将其用作(i)训练时间塑形信号和(ii)运行时动作过滤器,以最小化调整摆脚位置以满足模板级约束。全序安全通过MuJoCo中的Digit人形机器人和全体控制器栈进行实证评估。与无约束基线相比,我们的方法减少了报告的外部干扰试验中的安全违规事件,而较大的横向速度瞬变则显示出安全追踪的权衡。
FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL
FlexLoop:深度弹性循环策略用于深度强化学习中自适应测试时间计算
- Authors: Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34488
- Pdf link: https://arxiv.org/pdf/2609.34488
- Abstract
Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on $30$ online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to $\bf{43\%}$ and achieves up to $\bf{1.34\times}$ wall-clock speedup in a stress test.
- 中文摘要
循环架构通过在重复步骤中重复使用相同参数来扩展计算,最新研究显示它们在长视野任务中显著改善了深度强化学习策略。由于循环深度直接控制计算,可以预期循环策略自然支持跨循环深度的弹性推理。令人惊讶的是,我们发现预训练循环策略表现出严重的重复深度专用性:可靠决策集中在训练深度附近,即使计算量较少也可能将部署计算绑定在该深度。因此,实现深度弹性,即在部署时通过自适应计算实现跨循环深度的可靠决策,仍是一个关键挑战。为此,我们提出了FlexLoop,一种新型后训练框架,将预训练固定深度循环策略转换为深度弹性策略。FlexLoop 持续在原始 RL 目标基础上训练,保持全深度能力,同时进行邻近深度策略蒸馏,逐步将决策质量从更深层向浅层的重复步骤转移。最终策略支持跨循环深度的可靠推断,并通过重复深度一致性实现状态层次的自适应推理。在价值 $30 的在线和离线长视野目标条件环境中的实验显示,FlexLoop 在保持全深度性能的同时,使浅层更为有效。保持竞争力性能,FlexLoop 将平均重复深度降低高达 $\bf{43\%}$,并在压力测试中实现高达 $\bf{1.34\times}$ 的墙时钟加速。
Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
人工智能能在加密货币领域赚钱吗?衡量从回测到真实市场的差距
- Authors: Xingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei, Jiarui Liu, Chang Zhou, Fangzhou Ge, Chenyi Xu, Xikun Zhang, Renqiang Luo, Jie Zhang, Hong Cheng, Xinming Zhang, Hui Zhang, Yuan Fang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34510
- Pdf link: https://arxiv.org/pdf/2609.34510
- Abstract
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can generalize to unseen future markets or whether its backtested performance can be sustained in realistic trading frictions (e.g., latency, slippage, liquidity constraints, and market impact). We present a unified benchmark that evaluates representative machine learning, reinforcement learning, LLM-based, and agent-based trading methods in cryptocurrency markets through three progressively more realistic stages: historical backtesting, prospective exchange-based paper trading, and real-money live trading. These stages jointly increase temporal realism by moving from historical to unseen future markets, and execution realism by moving from offline simulation toward live trading. This protocol enables us to quantify the backtest-to-realization gap, identify when performance begins to deteriorate, and compare how this gap differs across major classes of AI trading methods. We further provide a unified open-source system supporting all three evaluation stages, together with a public platform that continuously updates benchmark results. Code is available at this https URL.
- 中文摘要
基于人工智能的交易方法已迅速从机器学习和强化学习发展到大型语言模型(LLM)和交易代理,但其性能仍主要通过历史回测来评估。此类评估有限地证明了方法是否能推广到未见的未来市场,或其回测表现能否在现实交易摩擦(如延迟、滑点、流动性约束和市场影响)下持续。我们提出了一个统一基准,通过三个逐步更现实的阶段——历史回测、未来交易所纸币交易和真钱真人交易——评估代表性机器学习、强化学习、基于LLM和基于代理的交易方法。这些阶段通过从历史市场向未见未来市场转变,增强了时间真实性,并通过从离线模拟转向实时交易提升执行现实性。该协议使我们能够量化回测到实现的差距,识别性能开始下降的时间,并比较不同主要AI交易方法类别之间的差异。我们还提供了一个统一的开源系统,支持所有三个评估阶段,并配备一个持续更新基准结果的公开平台。代码可在此 https 网址获取。
Verifying Neural Networks with Reinforcement Learning
通过强化学习验证神经网络
- Authors: Hai Duong, Thanh Le, ThanhVu Nguyen
- Subjects: Subjects:
Machine Learning (cs.LG); Software Engineering (cs.SE)
- Arxiv link: https://arxiv.org/abs/2609.34553
- Pdf link: https://arxiv.org/pdf/2609.34553
- Abstract
Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11% more instances while reducing branch exploration by 50%.
- 中文摘要
形式验证在确保深度神经网络(DNN)可靠性方面发挥关键作用,这些网络部署于安全关键系统中。现代DNN验证器采用分支限界框架,交替进行分支(分割成更小的子问题)和边界(剪枝子问题),以高效探索验证空间。然而,现有的分支启发式基于静态评分函数做出贪婪决策。它们不预期长期效率,也未利用验证数据日益增长的可用性来提升性能。这项工作引入了RSB,一种强化学习框架,学习如何精炼基线分支启发式。它训练演员-批判者架构,以最大化累计的未来奖励,而非即时得分。演员通过对原始神经元特征的观察和学习到的图嵌入生成注意力权重,这些方法重新调整基线启发式得分以指导神经元分支。对600个具有挑战性的实例的评估表明,RSB始终优于最先进的分支启发式,解决多11%的实例,同时减少分支探索50%。
Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
远景离线目标条件强化学习的扩散子目标规划
- Authors: Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34575
- Pdf link: https://arxiv.org/pdf/2609.34575
- Abstract
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
- 中文摘要
离线目标条件强化学习(GCRL)从无奖励数据中学习目标导向策略,但在长视野任务中,目标条件值函数常因奖励稀疏和折扣而提供不稳定的指导。分层方法通过子目标分解部分缓解了这一问题;然而,高层决策仍依赖噪声敏感的价值估计,导致复杂环境中行为不稳定。我们通过提出 \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}),这是一个基于扩散的高级子目标生成框架来解决这一限制。DSP 将高层规划视为对目标条件子目标的引导生成推理,学习条件和无条件流程,使无分类器的指导能够在推理时引入目标导向偏向。通过从高层规划中移除显式基于价值的指导,DSP通过生成模型生成可达且目标导向的子目标,同时保持层级执行。离线GCRL基准测试的实验表明,DSP在多种导航和操作任务中优于以往方法,尤其在需要多步子目标规划的迷宫环境中表现尤为出色。
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
从中间渲染中强化学习图像到代码生成
- Authors: Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34587
- Pdf link: https://arxiv.org/pdf/2609.34587
- Abstract
Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.
- 中文摘要
强化学习越来越多地用于对图像到代码生成的视觉语言模型进行后期训练,例如通过优化最终渲染输出计算的奖励,从参考图像生成SVG代码。然而,依赖单一终端奖励会带来稀疏反馈,与单个代币的贡献不匹配。生成程序可能包含准确重现目标图像部分的操作,同时也包含引入错误的部分,但所有代币均基于相同的最终结果训练。我们观察到许多中间代码前缀不仅可执行,而且已经产生有意义的部分渲染,反映向目标的进展。这一特性为生成过程中更密集的监督提供了自然来源。基于这一观察,我们引入了IR4RL,一个具有代币级渲染进度奖励的强化学习框架,将中间渲染之间的变化转化为生成序列的局部反馈。我们评估了图像到SVG和图像到TikZ生成的方法。在这两项任务中,我们的方法优于监督微调和标准GRPO,产生了新的先进开源模型。这表明中间渲染为强化学习后训练图像到代码模型提供了简单且有效的过程监督。
Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation
$Q^*$估计的深度加权贝尔曼残差最小化
- Authors: Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2609.34593
- Pdf link: https://arxiv.org/pdf/2609.34593
- Abstract
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
- 中文摘要
非策略评估是离线强化学习的基础组成部分,旨在利用预先收集的数据集评估和优化策略性能。然而,此类数据集常面临显著挑战,包括分布偏移、$Q$值高估和样本利用效率低。为解决这些问题,本文引入了加权Bellman残差最小化框架,通过有效整合专家演示与行为数据,整合了密度比加权。所提加权方案背离了深度强化学习理论分析中常见的完备性假设。我们建立了密度比估计的锐利收敛率,并推导出深度$Q^*$估计量超额风险的收敛率。大量实证评估表明,与现有方法相比,我们的方法在数值表现和政策泛化方面取得了显著提升,为合理利用专家演示提供了具体指导。
Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding
接地前总结:查询引导块凝聚,用于长视频时间接地
- Authors: Nanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu, Jinchao Zhang, Guoliang Kang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34598
- Pdf link: https://arxiv.org/pdf/2609.34598
- Abstract
Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a
summarize before grounding'' framework (namedSumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.
- 中文摘要
视频时间接地(VTG)旨在局部化对应语言查询的视频间隔。近期大型视觉语言模型(LVLM)在解决此类多模态推理任务方面展现出巨大潜力。然而,长视频通常包含大量冗余信息,扰乱LVLM挖掘查询相关证据。与导致大量训练记忆的密集帧抽样不同,以往可验证奖励强化学习(RLVR)通常采用稀疏采样,使训练变得可行,但可能遗漏关键证据。本文提出了一种“先总结再接地”框架(名为“SumGround”)用于长视频时间基础。SumGround的关键是执行查询引导的区块压缩,汇总和检索查询相关证据。具体来说,我们将视频拆分为若干块,并进行两级区块压缩。首先,我们引入了查询引导的潜在摘要,以查询引导提示的KV状态表示,将冗余的可视化标记压缩成紧凑且与查询相关的块摘要。此外,我们设计了一种关联摘要检索方案,用于排序和选择最有可能包含事件区间的块摘要。RLVR支持查询引导的潜在摘要和关联摘要检索方案。为减少内存消耗,我们提出了一个长度感知梯度门控模块,以选择性地阻止梯度反向传播到可视化标记。大量实验表明,SumGround在多个下游数据集中优于以往最先进方法表现优异,在长视频中表现显著。
The Low-Rank Structure of VLA Reinforcement Learning
VLA强化学习的低秩结构
- Authors: Minjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.34599
- Pdf link: https://arxiv.org/pdf/2609.34599
- Abstract
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $\pi_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-replacement experiments, we further show that these modules capture a disproportionate share of the performance gains from RL. We then characterize what is encoded in these Timestep Modules. First, we show that RL specializes them to the discrete denoising timesteps used during rollouts, and that this discrete-timestep training underlies the low-rank updates. Second, we find that among their outputs, the shift vector changes most distinctly under RL, and through probing, we show that shift update directions strongly predict task success (ROC-AUC up to $99.6\%$). Third, we find that the geometry of shift updates reflects task relationships, as their pairwise similarity correlates with cross-task transfer patterns. Building on these findings, we show that steering along shift update directions further improves RL-trained policies without additional RL training. Overall, we provide a systematic understanding of how RL reshapes VLA policies by studying how learned signals are encoded in parameter space, offering insights into more efficient and interpretable VLA post-training.
- 中文摘要
强化学习(RL)越来越多地被用于后期训练视觉-语言-行动(VLA)模型,但强化学习如何重塑这些策略仍不十分明了。我们发现,在广泛使用的基于流的VLA模型中,包括LIBERO、ManiSkill、MetaWorld和CALVIN上的$\pi_{0.5}$和GR00T~N1.5/N1.6,强化学习会诱导显著较低秩的参数更新,这些更新高度集中在动作专家的时间步模块中,这是一个小且此前被忽视的组成部分。通过系统的模块替换实验,我们进一步表明这些模块捕捉了强化学习性能提升的不成比例份额。随后,我们对这些时间步模块编码的内容进行了刻画。首先,我们展示了强化学习将其专门化为部署期间使用的离散去噪时间步长,而这种离散时间步训练是低秩更新的基础。其次,我们发现在它们的输出中,移位向量在强化学习下变化最为明显,通过探究,我们发现移位更新方向强烈预测任务成功率(ROC-AUC 最高可达99.6%)。第三,我们发现移位更新的几何形态反映了任务关系,因为它们的成对相似性与跨任务转移模式相关。基于这些发现,我们证明沿移位更新方向引导进一步提升了强化学习训练策略,无需额外强化学习训练。总体而言,我们通过研究学习信号如何在参数空间中编码,系统地理解强化学习如何重塑VLA策略,提供更高效、更易解释的训练后VLA的见解。
Nereus: Adaptive Parallelism for LLM Post-Training
Nereus:LLM 后期培训的自适应并行
- Authors: Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, Bo Zhao
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34645
- Pdf link: https://arxiv.org/pdf/2609.34645
- Abstract
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.
- 中文摘要
强化学习(RL)用于大型语言模型(LLM)的后期训练,协调多个模型在生成、推理和GPU集群上的训练过程。运行过程中可能发生多个变化因素,包括资源可用性、序列长度、内存压力和阶段瓶颈。因此,最初合适的执行计划随着时间推移可能变得缓慢甚至不可行。然而,适应模型共享GPU的作业存在重大挑战:判断新计划是否值得承担过渡成本,重用作业的分布式状态,以及协调跨模型和阶段的GPU传输。Nereus将这些挑战定位为一个成本感知型运行时,将RL后训练作业调整为高效的执行计划。其低开销控制器选择内存可行的全局计划,并通过基于运行任务校准的成本模型进行过渡。为了估计和执行一个转移,Nereus 将模型阶段的每个副本(一个模型在一个阶段)的分布状态表示为弹性模型单元。然后它使用全局过渡图来排序这些单元的转换和 GPU 传输。在基于真实数据构建的轨迹中,在线 TP/PP 适配相比初始固定 TP/PP 布局(带有 DP 缩放)降低了 27.7% 的平均步进延迟。在 1,000 步运行并达到 1,024 个 GPU 的过程中,六次转换占用总运行时间的 0.079%。Nereus 在不同集群中,端到端 8B PPO 吞吐量比 OpenRLHF 提升 2.14 倍——7.27 美元/时间点,比 Verl 提升 1.10 倍——1.47 美元/倍值。
Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization
统一轨迹匹配策略优化:多样化的T2I生成与VLA泛化
- Authors: Zhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu, Xuekai Zhu, Dingkang Liang, Kaiyan Zhang, Jianjun Li, Bowen Zhou, Xiang Bai
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.34688
- Pdf link: https://arxiv.org/pdf/2609.34688
- Abstract
Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.
- 中文摘要
奖励最大化强化学习(RL)被广泛用于文本到图像(T2I)生成的随机扩散和流策略的后期训练。然而,奖励最大化强化即使在参考 KL 或熵正则化下也会导致策略模式崩溃,使策略仅剩单一高奖励模式。在 T2I(T2I)中,这会产生类似的图像和奖励黑客。当扩展到视觉-语言-行动(VLA)模型时,同样的崩溃会移除其他成功策略,削弱任务和场景的泛化。为解决这一限制,我们引入了统一轨迹匹配策略优化(Uni-TMPO),这是一个统一的强化学习后训练框架,用于扩散和流策略。首先,Uni-TMPO 将标准化奖励转换为每个轨迹组内的目标分布,并从轨迹日志概率推导策略分布。然后,通过前向的Kullback-Leibler优化匹配两种分布,而不是最大化期望奖励。进度条件的粗到细调度调度器高效构建T2I轨迹。在统一框架内,反馈条件抽样利用更新的观测数据构建VLA轨迹。大量实验表明,Uni-TMPO在T2I奖励和VLA ID成功率上优于最强基线。更重要的是,它实现了最佳的T2I奖励多样性效率权衡和VLA对未完成任务和场景的推广,同时真实机器人评估展示了当高奖励目标被阻断时多重行动策略的价值。
Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control
稳定控制中零阶奖励塑造政策梯度的充分性
- Authors: Yisheng Zhang, Tao Wang, Sicun Gao
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.34695
- Pdf link: https://arxiv.org/pdf/2609.34695
- Abstract
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
- 中文摘要
奖励塑造是现代机器人控制的基础,采用深度强化学习(RL),但实践者仍大量依赖借鉴经典最优控制和轨迹优化的启发式原则。现有方法很少区分控制目标内在的奖励项与数值正则化,导致超参数调优脆弱。为确定奖励必须包含哪些量,我们研究稳定控制问题,重点关注零阶(配置)和一阶(速度)信息。我们理论和实证证明,策略梯度方法可以在不包含一阶奖励项的情况下成功解决稳定任务,添加此类项随着尺度增长会带来严重的敏感性。相反,我们的发现确认奖励函数必须在目标相关坐标上零阶完全,而在低耗散假设下,一阶状态在策略观察中仍为必要。总体而言,这些结果为机器人强化学习中的奖励设计提供了可操作且有原则的指导。
Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning
质量决定方向,长度决定形状大小:长度控制用于开放式强化学习
- Authors: Zijun Weng, Zhongan Bi, Xuanang Gao, Xiaohui Hu, Shuangyong Song, Yongxiang Li, Kaidong Yu, Xuanjing Huang
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.34718
- Pdf link: https://arxiv.org/pdf/2609.34718
- Abstract
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii) open-ended tasks lack a natural success boundary for deciding when efficiency should be prioritized, and (iii) dense, graded rewards often yield small within-group quality margins, making quality-induced advantages especially sensitive to reward-level length shaping, which can perturb their magnitudes and even reverse their signs. We therefore adopt an asymmetric principle: quality should determine the direction of reinforcement, while length should only shape its magnitude. We instantiate this principle with Quality-Gated Length Advantage Shaping (QGLAS), which first computes advantages from quality rewards alone, then adds bounded bonuses only to shorter positive-advantage responses, leaving all other advantages unchanged. The bonus strength is further adapted to within-group quality separation, allowing conciseness to matter more when quality-favored responses are similar and less when their quality differences are clear. Across different model families, open-ended benchmarks, and reward sources, QGLAS consistently achieves a stronger quality--length trade-off than representative baselines. At approximately 30% compression, QGLAS retains 98.4--102.0% of the macro-average quality gains achieved by quality-only RL over the base model, compared with 68.3--75.5% for these baselines at comparable compression.
- 中文摘要
强化学习(RL)不仅改变语言模型的说话内容,还改变它们的表达内容,常常增加反应长度,但代价是代币效率。在开放式强化学习中,控制这种长度增长尤其具有挑战性,因为(i)响应长度与质量纠缠,(ii)开放式任务缺乏自然的成功边界来决定何时应优先考虑效率,(iii)密集且分级的奖励往往产生小组内质量边缘较小,使得质量诱导的优势对奖励水平长度塑造特别敏感,这可能会扰动其幅度甚至反转符号。因此,我们采用不对称原则:质量应决定强化的方向,而长度应仅塑造其大小。我们通过质量门槛长度优势塑造(QGLAS)实现这一原则,该方法首先仅从质量奖励计算优势,然后仅对较短的正向优势反应添加有界奖励,其他优势保持不变。加成强度进一步适应组内质量分离,使得当质量偏好的回答相似时,简洁性更为重要;当质量差异明显时,简洁性更为重要。在不同模型家族、开放式基准和奖励来源中,QGLAS始终比代表性基线在质量-长度权衡上表现更强。在约30%压缩率下,QGLAS保留了仅质量强化学习相较基础模型的宏观平均质量提升的98.4--102.0%,而在相似压缩条件下,这些基线仅为68.3-75.5%。
DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations
DexWeave:通过人类演示学习灵巧的人形机车操控
- Authors: Naichuan Sun, Haotian Shen, Yizhang Zhang, Luying Feng, Haoze Wang, Yuanbo Xiangli, Yaochu Jin, Peidong Liu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.34724
- Pdf link: https://arxiv.org/pdf/2609.34724
- Abstract
Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (this https URL) for videos.
- 中文摘要
从人类演示中学习灵巧的人形机车操作不仅需要转移人类运动,还需要转移支撑该行为的协调互动结构。这具有挑战性,因为身体差异会扭曲身体动作、手腕位置、手指关节和物体交互之间的耦合,而在机器人动力学下,运动学上的精确参考仍难以实现。我们提出了DexWeave,一个统一框架,连接互动一致的动作重定向与解剖感知的全身政策学习。DexWeave首先采用两阶段重定向程序,通过专业求解器初始化身体和手部动作,随后在上半身交互链上进行耦合细化,同时保留下半身支撑。所得引用由一个具解剖感知的Transformer策略跟踪,该策略将解剖区域表示为结构化标记,并使用定向掩蔽注意力建模其依赖关系,对象信息选择性地条件上半身通路以实现灵巧互动。该策略联合输出身体和灵巧手的动作,直接通过强化学习训练,无需预训练的追踪策略、师生关系或后续残余细化。DexWeave提升了重定向的忠实度和交互一致性,同时实现了比MLP基线更高的操作性能和更快的策略收敛。我们还进一步将所学策略部署在配备Inspire灵巧手的物理Unitree G1人形上,演示了现实世界中灵巧的全身运动操作。视频请访问我们的项目页面(此 https URL)。
Keyword: diffusion policy
CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning
资本支出:通过体验自适应推理,高效将基础模型行为提炼为可部署机器人策略
- Authors: Shivam Aarya, Zhang Xi-Jia, Chengyue Huang, Junhyun Kim, Huishu Xue, Hrishit Leen, Roman Yakunin, Animesh Garg, Zsolt Kira
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.33007
- Pdf link: https://arxiv.org/pdf/2609.33007
- Abstract
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: this https URL
- 中文摘要
机器人学习主要依赖人类远程操作演示来获得有效的可学习行为。然而,人类操作的数据收集过程可能不直观、难以扩展且本质上是异步的。我们探讨另一种方案:通过使用基础模型本身作为自主演示器,将通用多模态基础模型中的物理行为提炼为可部署的机器人策略。虽然足够强大的模型可以生成成功的零样本操作轨迹,但在物理执行中反复调用这些轨迹缓慢且成本高昂,限制了其作为可扩展数据生成器的实用性。作为解决方案,我们引入了CAPEX——一种基于经验条件的演示收集框架,利用以往尝试的执行经验来调整基础模型必须观察、推理和重新规划的频率。我们在RoboCasa任务以及实体Franka和双手YAM臂平台上进行评估,衡量任务成功率、模型调用、令牌使用情况、收集时间和成本。我们还进一步训练Diffusion Policy和ACT,用于匹配的人类远程操作和基金会模型生成的演示集,以评估自主收集数据的下游学习价值。我们发现CAPEX使成功演示数量增加4.3倍,同时每次成功演示的成本降低80%。基于CAPEX生成数据训练的策略性能接近匹配人工演示训练策略;随着训练时间更长,这一差距在从零开始训练的策略中基本缩小。这些结果表明基础模型可以作为可扩展的可重用机器人体验来源。项目页面:此 https URL