生成时间: 2026-08-19 16:39:47 (UTC+8); Arxiv 发布时间: 2026-08-19 20:00 EDT (2026-08-20 08:00 UTC+8)
今天共有 30 篇相关文章
Keyword: reinforcement learning
Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning
通过LLM引导强化学习打造高效的个性化AI导师
- Authors: Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
- Subjects: Subjects:
Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16907
- Pdf link: https://arxiv.org/pdf/2608.16907
- Abstract
Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely focus on GenAI chatbot tutors that reactively answer student questions. We hypothesize that the efficacy of GenAI chatbot tutors can be substantially improved by proactively guiding student learning. To test this, we design a novel tutoring platform that tightly integrates a carefully-designed GenAI chatbot with a reinforcement learning algorithm for sequencing practice problems. Critically, this algorithm leverages rich signals from student-chatbot interactions to adaptively select practice problems of an appropriate difficulty level. In partnership with the Taipei City Government and American Institute in Taiwan, we deployed our tutoring platform in conjunction with a five-month course to teach Python to students across ten high schools. We randomized students between a fixed practice problem sequence and our adaptive sequencing algorithm. We find that adaptive sequencing increased unassisted final exam performance by 0.15 standard deviations (equivalent to 6-9 months of schooling by some estimates); mediation analysis suggests that gains were driven by increased engagement. Our work provides large-scale field evidence that student-chatbot interactions provide valuable signals for proactively optimizing and personalizing student learning.
- 中文摘要
生成式人工智能(GenAI)正在快速重塑教育,释放个性化辅导的潜力。然而,新兴平台主要专注于能够被动回答学生问题的生成式AI聊天机器人导师。我们假设,通过主动引导学生学习,生成式AI聊天机器人导师的效能可以显著提升。为此,我们设计了一个新颖的辅导平台,将精心设计的生成式人工智能聊天机器人与用于练习题排序的强化学习算法紧密集成。关键是,该算法利用学生与聊天机器人互动中的丰富信号,自适应地选择适合难度的练习题。我们与台北市政府和台湾美国学院合作,将辅导平台与为期五个月的课程结合起来,为十所高中的学生教授Python。我们将学生随机分配到固定练习题序列和自适应测序算法之间。我们发现,自适应测序使无辅助期末考试的表现提高了0.15个标准差(部分估计相当于6-9个月的学习时间);调解分析表明,提升的收益主要来自参与度的提升。我们的研究提供了大规模的实地证据,表明学生与聊天机器人的互动为主动优化和个性化学生学习提供了宝贵信号。
FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences
FetchMan:从模拟体验中学习可视化人形机车操作策略
- Authors: Omar Rayyan, Zhi Li, Max Argus, Yuxin Jiang, Chang Yu, Chenfanfu Jiang, Yuchen Cui
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.17027
- Pdf link: https://arxiv.org/pdf/2608.17027
- Abstract
Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robotics research. However, today's data-hungry algorithms make collecting sufficient demonstrations a struggle for tabletop manipulation, and even more so for humanoids that must also walk and balance. Learning from simulated data and transferring that behavior to the real world, as is commonly done in locomotion, sidesteps this struggle, so we replicate that recipe for loco-manipulation. In doing so, we find that cloning synthetic demonstrations results in a low performance ceiling no matter the amount of training data. Reinforcement learning breaks through it, and refining the cloned policy with Flow-GRPO on a single sparse reward yields performance that synthetic behavior cloning cannot match. Together, these stages form our end-to-end sim-to-real pipeline spanning more than 150,000 scenes, which we use to train FetchMan. We evaluate it on FetchMan-Bench, a simulation benchmark we release, and deploy it zero-shot on a real Unitree G1, where our single-object reach-and-pick policy walks to and grasps a target across unseen scenes at 73.3% success. Finally, we extend this recipe to multi-object training, a first step toward loco-manipulation generalist policies at this data scale.
- 中文摘要
能够推广到新场景和新物体的视觉操作策略长期以来一直是机器人研究的目标。然而,如今数据需求极高的算法使得收集足够演示成为桌面操作的难题,对于还必须行走和保持平衡的人形生物来说更是如此。通过从模拟数据中学习并将这种行为转移到现实世界,就像在机动中常见的做法一样,可以绕过这种难题,因此我们复制了这种机车操控的配方。在此过程中,我们发现无论训练数据多少,克隆合成演示的性能上限都很低。强化学习突破了这一点,用Flow-GRPO对单次稀疏奖励进行克隆策略的细化,能获得合成行为克隆无法匹敌的性能。这些阶段共同构成了我们从模拟到现实的端到端流程,涵盖超过15万个场景,我们用这些流程来训练FetchMan。我们在FetchMan-Bench(我们发布的模拟基准测试)上评估它,并在真实的Unitree G1上零射击部署,我们的单对象范围范围并选择策略能以73.3%的成功率穿越未见场景的目标。最后,我们将这一配方推广到多对象训练,这是迈向该数据尺度上机车操控通用策略的第一步。
Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation
Lambda保持控制:预测性肌肉骨骼模拟中,从最小任务奖励中产生类人运动
- Authors: Jun Hyuk Lee, Chihyeong Lee, Jooeun Ahn
- Subjects: Subjects:
Robotics (cs.RO); Graphics (cs.GR); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17030
- Pdf link: https://arxiv.org/pdf/2608.17030
- Abstract
The massive overactuation in the human musculoskeletal system makes it challenging to train musculoskeletal models to generate human-like motion via reinforcement learning, primarily because exploration in the resulting high-dimensional and redundant action space is extremely inefficient. To address this problem, we propose the $\lambda$-hold controller, inspired by the equilibrium-point (EP) hypothesis, which has been widely supported by extensive evidence from human motor control studies. The policy's control variable is the per-muscle EP threshold length $\lambda$, from which a stretch-reflex recruitment law computes the muscle excitations automatically. Holding each $\lambda$ over an interval of the gait phase also sharply reduces the frequency at which the policy must be queried. Consequently, the controller, to our knowledge for the first time, enables a muscle-actuated skeletal model to learn human-like sprinting using only a minimal reward within an hour of training. The efficient exploration through the proposed $\lambda$-hold controller is not merely an engineering trick but an approach grounded in physiology, bringing together the EP hypothesis, intermittent control, and optimal feedback control. Beyond encapsulating human-like behavior in predictive simulation, this achievement contributes to developing a learnable model of the human motor controller.
- 中文摘要
人体肌肉骨骼系统的大规模过度驱动使得通过强化学习训练肌肉骨骼模型产生类人运动具有挑战性,主要因为在由此产生的高维冗余动作空间中探索极其低效。为解决这一问题,我们提出了$\lambda$-保持控制器,灵感来自平衡点(EP)假说,这一假设已被大量人体运动控制研究支持。该策略的控制变量是每块肌肉的EP阈值长度$\lambda$,从中拉伸反射招募律自动计算肌肉兴奋量。在步态相位的一定区间内保持每个$\lambda$,也会大幅降低需要查询策略的频率。因此,据我们所知,控制器首次使肌肉驱动的骨骼模型能够在训练一小时内,仅用极少奖励学习类人冲刺。对所提议的$\lambda$-hold控制器进行高效探索不仅是工程技巧,更是一种基于生理学的方法,结合了EP假说、间歇性控制和最优反馈控制。除了在预测模拟中封装类人行为外,这一成就还有助于开发可学习的人体运动控制模型。
Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis
基于物理的强化学习用于随机距离-避免分析
- Authors: Hikaru Hoshino, Yorie Nakahira
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2608.17117
- Pdf link: https://arxiv.org/pdf/2608.17117
- Abstract
Stochastic reach-avoid analysis of controlled dynamical systems is an important tool for safety-critical control under uncertainty, in which the reach-avoid probability is characterized by a Hamilton-Jacobi partial differential equation (PDE). However, solving this PDE using conventional numerical methods becomes computationally intractable as the system dimension increases. Physics-informed neural networks (PINNs) may converge to inaccurate local minima when trained primarily through PDE-residual minimization. Reinforcement learning (RL) offers a scalable alternative, but its learned value functions may be inaccurate or inconsistent with the governing PDE. This paper proposes a physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis. We develop a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function. PDE-residual and boundary-condition losses are then introduced progressively to enforce consistency with the governing PDE and its boundary conditions. The proposed method mitigates the failure modes of conventional PINN techniques while achieving accuracy comparable to that of successfully trained PINNs. The effectiveness of the proposed framework is demonstrated through two case studies.
- 中文摘要
受控动力学系统的随机距离-避免分析是不确定性下安全关键控制的重要工具,其中达到-避免概率由哈密顿-雅可比偏微分方程(PDE)表征。然而,随着系统维度的增加,使用传统数值方法求解该偏微分方程会变得计算上难以处理。物理知情神经网络(PINN)在主要通过偏微分方程残差最小化训练时,可能会收敛到不准确的局部极小值。强化学习(RL)提供了一种可扩展的替代方案,但其学到的价值函数可能与主导的偏微分方程不准确或不一致。本文提出了一种基于物理的强化学习(PIRL)框架,结合了PINNs和强化学习的互补优势,用于随机距离-避免分析。我们开发了一种定时的PIRL算法,其中时间差分-演员-批评者学习首先引导批评者朝向有意义的接近距离-避免价值函数。随后逐步引入偏微分方程残差和边界条件损耗,以确保与支配偏微分方程及其边界条件的一致性。该方法在降低传统PINN技术的失效模式的同时,实现了与成功训练的PINN相当的精度。该框架的有效性通过两个案例研究得到了验证。
Q-Learning With World Models
Q-世界模型学习
- Authors: Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17163
- Pdf link: https://arxiv.org/pdf/2608.17163
- Abstract
Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised policy learning. Prior model-based RL methods often optimize the policy or value function directly on imagined rollouts, which is prone to compounding bias and struggles to scale to large, high-dimensional problems such as real-world robotics, a problem that worsens with task horizon and visual complexity. In this work, we instead ask whether we can leverage world models directly on top of standard Q-learning to improve performance, while remaining trained and grounded in the real, online setting. We propose QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation. Since the policy and value function are trained only on real transitions, QWM avoids compounding model bias while still gaining the sample-efficiency benefits of predictive search. On challenging manipulation benchmarks Robomimic and LIBERO, QWM significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performance.
- 中文摘要
非策略强化学习(RL)变得越来越高效,使得如对视觉-语言-行动模型进行强化微调等应用成为可靠且高性能的策略。世界模型为样本效率提供了进一步杠杆,因为它们预测状态变化而非仅仅是行动,但其成功主要局限于监督政策学习。以往基于模型的强化学习方法通常直接在想象中的推广上优化策略或价值函数,这容易产生叠加偏见,并且难以扩展到大型高维问题,如现实世界机器人,而这一问题随着任务视野和视觉复杂性加剧而加剧。在本研究中,我们探讨是否可以直接利用世界模型在标准Q学习之上提升表现,同时保持在真实在线环境中的训练和扎根。我们提出了QWM,这是一个利用世界模型在测试时基于想象轨迹进行测试时搜索的框架,基于Q学习,在在线推广和评估过程中选择高价值动作。由于策略函数和价值函数仅在实转移上训练,QWM避免了模型偏倚的复合,同时仍能获得预测搜索的样本效率优势。在挑战性的操作基准测试Robomimic和LIBERO上,QWM在样本效率和性能方面显著优于以往的先进方法。
Task Specialization Fine-Tuning for Contextual Reinforcement Learning
情境强化学习的任务专门化微调
- Authors: Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng, Jie Zhang, Roy Dong, Yining Ma, Cathy Wu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17180
- Pdf link: https://arxiv.org/pdf/2608.17180
- Abstract
Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by fine-tuning multiple policies for task specialization. This new paradigm, however, introduces unique challenges, such as heterogeneous marginal returns and sample inefficiency. This raises a critical research question: given a pretrained policy and a constrained budget, how much fine-tuning should each task region receive to enable sample-efficient CRL? To this end, we propose Task Specialization Fine-Tuning (TSFT), an online framework that predicts fine-tuning performance with a simple parametric model and exactly solves the resulting discrete budget allocation problem via integer linear programming. Extensive experiments across diverse decision domains, including combinatorial optimization, continuous control, and LLM fine-tuning, demonstrate that TSFT significantly outperforms baselines in task coverage and approaches oracle performance. Our work charts a new direction for model-based CRL, aligning with the modern pretrain-finetune era.
- 中文摘要
情境强化学习(CRL)旨在通过最大化相关任务上下文空间的任务覆盖,推广经典强化学习。虽然以往的工作通常从零开始训练,依赖单一策略的多任务学习或战略性训练多个策略,但我们主张统一的替代方案:先预训练一个初始表现良好的策略,然后微调多个策略以实现任务专精。然而,这一新范式带来了独特的挑战,如异质边际收益和样本效率低下。这引发了一个关键研究问题:在预训练政策和有限预算的情况下,每个任务区域应进行多少微调,以实现样本高效的CRL?为此,我们提出了任务专用微调(TSFT)的在线框架,通过简单的参数模型预测微调性能,并通过整数线性规划精确解决所得的离散预算分配问题。涵盖组合优化、连续控制和大型语言模型微调等多个决策领域的广泛实验表明,TSFT在任务覆盖率上显著优于基线,并接近oracle性能。我们的工作为基于模型的CRL开辟了新的方向,与现代预列车微调时代保持一致。
Reinforcement Learning as (Discrete) Potential Theory
作为(离散)势理论的强化学习
- Authors: Christopher Connolly
- Subjects: Subjects:
Machine Learning (cs.LG); Computer Science and Game Theory (cs.GT)
- Arxiv link: https://arxiv.org/abs/2608.17181
- Pdf link: https://arxiv.org/pdf/2608.17181
- Abstract
Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There is a deep connection between probability theory and potential theory. This paper reviews that connection and explores the potential-theoretic viewpoint for core reinforcement learning representations and algorithms under a fixed-policy assumption. This viewpoint may offer a path for improved sample efficiency and formal constraints that can be applied to RL. When the fixed-policy assumption is relaxed, the linear potential theory framework can be naturally extended to the nonlinear case.
- 中文摘要
强化学习(RL)理论从根本上依赖于通过马尔可夫链的概率论。概率论与势理论之间有着深厚的联系。本文回顾了这一联系,并探讨了在固定策略假设下核心强化学习表示和算法的潜在理论视角。这一观点可能为提升样本效率和形式约束提供一条可应用于强化学习的路径。当固定策略假设被放宽时,线性势理论框架可以自然地推广到非线性情形。
An O-RAN-Assisted MARL Approach for Dynamic Sidelink and Infrastructure Selection in V2X Communications
一种基于O-RAN辅助的MARL方法,用于V2X通信中的动态侧链和基础设施选择
- Authors: Maria Katarine Santana Barbosa, Kelvin Lopes Dias
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2608.17210
- Pdf link: https://arxiv.org/pdf/2608.17210
- Abstract
Future applications in the 6G-based Internet of Vehicles will leverage sidelink (SL) transmissions in Vehicle-to-Everything (V2X) scenarios. However, SL-based direct communication can significantly increase interference among vehicles and between vehicles and other entities of the Intelligent Transportation System. Thus, both Vehicle-to-Vehicle communications and Vulnerable Road Users (VRUs) uplink resources may be degraded or subject to starvation. Existing solutions primarily focus on improving resource allocation and pair selection. Nonetheless, they lack a comprehensive approach to tackle the communication modes and the entire network. To address these challenges, this paper leverages Open RAN to manage V2X communication and proposes a multi-agent reinforcement learning (MARL) resource-aware system. Open RAN provides control loops through a global view of the network and also an open interface-based framework for machine learning models applied to resource decision-making. Meanwhile, the MARL model aims to mitigate interference, optimize resource usage, and enhance quality of service by optimally selecting between sidelink and network transmissions. To reduce system complexity, this work employs a clustering strategy. Each agent manages a group of pairs, rather than assigning one agent to each pair. The solution supports this design by adopting a centralized training with decentralized execution approach, empowered by Open RAN. The strategy uses offline training and an off-policy approach, in which each agent stores experience for fine-tuning. Results indicate that the MARL approach reduces average loss by 21% and latency by 19% in Vehicle-only scenarios. In coexistence VRU scenarios, loss and latency drop by 18% and 30%, respectively, compared to the single-agent approach.
- 中文摘要
未来基于6G的车联网应用将利用侧链(SL)传输,应用于车对全(V2X)场景。然而,基于SL的直接通信会显著增加车辆之间以及车辆与智能交通系统其他实体之间的干扰。因此,车辆间通信和脆弱道路用户(VRU)的上行资源都可能被削弱或面临枯竭。现有解决方案主要侧重于改善资源分配和配对选择。然而,他们缺乏全面的方法来应对通信模式和整个网络。为应对这些挑战,本文利用Open RAN管理V2X通信,并提出了一个多智能体强化学习(MARL)资源感知系统。开放RAN通过全局视野提供控制环路,同时也为应用于资源决策的机器学习模型提供基于开放接口的框架。与此同时,MARL模型旨在通过在侧链和网络传输之间优化选择,减少干扰、优化资源使用并提升服务质量。为了降低系统复杂性,这项工作采用了聚类策略。每个代理管理一组配对,而不是为每对分配一个代理。该解决方案通过采用集中式训练和去中心化执行方法支持这一设计,并由开放RAN赋能。该策略采用离线训练和非策略方法,每个代理存储经验以便微调。结果显示,MARL方法在仅载具的情况下平均损耗降低21%,延迟降低19%。在共存VRU场景中,丢包和延迟分别下降18%和30%,相较于单代理方法。
Safe Deep Reinforcement Learning for Energy-Efficient HVAC Control in Multi-Zone Residential Buildings
多区住宅建筑中节能暖通空调控制的安全深度强化学习
- Authors: Oussama Ziadi, Abdelilah Rochd, Samir Idrissi Kaitouni, Mohamed Oualid Mghazli, Adnane Saoud
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2608.17235
- Pdf link: https://arxiv.org/pdf/2608.17235
- Abstract
HVAC systems represent a major share of building energy consumption. Traditional control strategies are limited in coordinating energy-comfort tradeoffs across multiple zones simultaneously. Reinforcement learning (RL) offers adaptive, data-driven control that optimizes performance over time. However, deploying learned neural network controllers in safety-critical building systems remains challenging due to lack of formal safety guarantees. We propose a safety-certified deep RL framework for multi-zone residential HVAC control. Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) agents are trained in an EnergyPlus/Sinergym simulation to minimize energy consumption while maintaining thermal comfort. Post-training safety certification is performed on the PPO policy using Lipschitz-based forward invariance analysis, building on existing tools for the computation of Lipschitz constants for neural networks, to guarantee constraint satisfaction. Both agents are evaluated over an annual simulation cycle in an eight-zone variable refrigerant flow (VRF) testbed. The PPO agent achieves 67\% comfort violation reduction compared to rule-based control, while the SAC agent achieves 27.6\% energy savings. The PPO policy satisfies formal safety certification with a margin of $2.003^\circ$C. These results demonstrate the feasibility of combining reinforcement learning with post-training safety verification for multi-zone building control.
- 中文摘要
暖通空调系统占建筑能源消耗的重要部分。传统控制策略在协调多个区域的能量与舒适权衡方面有限。强化学习(RL)提供自适应、数据驱动的控制,以优化长期表现。然而,由于缺乏正式的安全保障,在安全关键的建筑系统中部署学习过的神经网络控制器仍然具有挑战性。我们提出了一个经过安全认证的深度强化学习框架,用于多区住宅暖通空调控制。近端策略优化(PPO)和软演员-批判(SAC)代理在EnergyPlus/Sinergym模拟中接受训练,以最大限度地减少能耗同时保持热舒适性。培训后安全认证基于基于Lipschitz的前向不变性分析,基于现有的神经网络Lipschitz常数计算工具,确保约束满足。两款试剂均在八区可变制冷剂流量(VRF)测试平台中进行年度模拟周期评估。与基于规则的控制相比,PPO代理实现了67%的舒适度违规减少,而SAC代理则实现了27.6%的节能。PPO政策以2.003^\circ$C的保证金满足正式安全认证。这些结果表明,将强化学习与培训后安全验证相结合,实现多区建筑控制的可行性。
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Co-RL:多代理强化学习中多元群体中出现的无监督推理
- Authors: Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.17253
- Pdf link: https://arxiv.org/pdf/2608.17253
- Abstract
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text-only and multimodal domains, Co-RL consistently outperforms the base models and prior label-free approaches, while matching or surpassing supervised methods, without access to any ground-truth labels. Concretely, Co-RL yields average gains of 3.0-8.6% across seven text-only benchmarks for LLMs and 2.3-7.2% across four multimodal benchmarks for VLMs. Code is available at this https URL.
- 中文摘要
强化学习(RL)已成为提升语言和视觉语言模型推理的有力方法,但其最强的成功仍高度依赖于真实监督(例如可验证的奖励)。此类注释获取成本高昂,随着推理能力的提升,人类的可靠评估能力日益稀缺。自我奖励强化学习通过使模型能够从自身完成中推导出奖励信号,从而减少这种依赖。然而,仅依赖自我生成反馈进行训练可能会强化现有偏见和次优行为,降低反应多样性,最终导致反应趋同和训练崩溃。本研究展示了通过合作多代理训练可以实现无监督推理。我们介绍了Co-RL,这是一个框架,在该框架中,多个无参数的解耦模型通过RL同时通过其对等模型的奖励进行优化。我们还进一步表明,通过异质模型族、规模和重新表述的训练样本来增加队列多样性,可以减少驱动自我强化反馈循环的相关错误。这种多样性持续提升推理表现,维持行为多样性,并减少训练崩溃。在纯文本和多模态领域,Co-RL始终优于基础模型和之前无标签方法,同时在无需任何基础标签的情况下,甚至超越监督方法。具体来说,Co-RL在七个纯文本基准测试中获得3.0-8.6%的平均收益,在四个多模态基准测试中实现2.3-7.2%。代码可在此 https 网址获取。
A Hybrid End-to-End and Modular Control Architecture Toward Safe Vehicle Lateral Control: Combining Soft Actor-Critic with Model Predictive Control
一种混合端到端和模块化控制架构,实现车辆横向安全控制:软行为者-批评与模型预测控制的结合
- Authors: Farzaneh Tatari
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2608.17258
- Pdf link: https://arxiv.org/pdf/2608.17258
- Abstract
Connected and automated vehicles demand lateral controllers that are simultaneously accurate, low-effort, and safe under model error and sensor noise. Modular controllers such as model predictive control (MPC) are interpretable and constraint-aware but rely on accurate models and hand-tuned weights. End-to-end learned policies, in particular continuous-action deep reinforcement learning, are adaptable and require no hand-designed control law, but offer no intrinsic safety guarantees and limited interpretability. This paper presents a hybrid architecture that combines an end-to-end Soft Actor-Critic (SAC) policy with a constrained linear MPC into a single steering command, using the MPC's first-step optimum as the model-based anchor and a single monotone blending coefficient that interpolates between the two paradigms. The architecture is evaluated on a linearized lateral bicycle model against a PID baseline, a tuned linear MPC, and a stand-alone SAC policy, across nominal, single-axis robustness, and multi-initial-condition ensemble experiments. The hybrid retains the tracking quality of stand-alone SAC while remaining inside the MPC's actuator envelope and preserving a deterministic, model-based contribution to every steering command. The architecture provides an actuator-envelope guarantee by construction but does not establish recursive feasibility or terminal invariance, and the closed-form blend does not prevent all corner-case divergences at the boundary of the training distribution. A corner-case analysis shows that the blend attenuates but cannot prevent failure under distribution shift, motivating a connectivity-aware extension in which the blending coefficient is scheduled by vehicle-to-everything (V2X) signals to restore model-based authority. Limitations and a path toward a constrained-QP predictive safety filter are discussed.
- 中文摘要
互联和自动驾驶车辆需要横向控制器,这些控制器既准确、低成本,又能在模型误差和传感器噪声下安全。模块化控制器如模型预测控制(MPC)可解释且具约束感知能力,但依赖于准确的模型和手工调整的权重。端到端学习策略,特别是连续动作深度强化学习,具有适应性,无需手工设计的控制定律,但没有内在安全保障和有限的解释性。本文提出了一种混合架构,将端到端的软演员-批评者(SAC)策略与受限线性MPC结合为单一引导命令,使用MPC的第一步最优值作为基于模型的锚点,并以单一单调混合系数在两种范式之间插值。该架构基于线性化的横向自行车模型,结合PID基线、调优线性MPC和独立SAC策略进行评估,涵盖名义实验、单轴鲁棒性和多初始条件集合实验。混合动力保持了独立SAC的跟踪特性,同时保持在MPC执行器范围内,并保持对每个转向指令的确定性、基于模型的贡献。该架构通过构造提供了执行器-包络的保证,但并未确立递归可行性或终端不变性,封闭形式混合也无法防止训练分布边界处的所有角落情况发散。一个极端情况分析表明,混合在分配偏移下会衰减但无法防止失效,这促使采用一种连通性感知扩展,即通过车辆对所有信号(V2X)调度混合系数,以恢复基于模型的权威性。讨论了其局限性以及实现约束QP预测安全滤波器的路径。
The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs
少有人走的道路:FPGA的拥塞感知NoC布置与数据包路由
- Authors: Soheil Gholami Shahrouz, Vaughn Betz
- Subjects: Subjects:
Hardware Architecture (cs.AR)
- Arxiv link: https://arxiv.org/abs/2608.17266
- Pdf link: https://arxiv.org/pdf/2608.17266
- Abstract
To help scale to ever-larger and more complex designs, recent FPGA architectures now integrate network-on-chips (NoCs). NoCs help transfer high-bandwidth data over long distances within the chip without using scarce low-delay long routing wire segments. While NoC-enhanced FPGAs aid system integration and design reuse, they also complicate FPGA computer-aided design (CAD) flows by introducing new constraints and metrics. Placement and routing need to optimize NoC metrics like latency and bandwidth utilization and avoid link oversubscription (congestion), while simultaneously optimizing the programmable routing resource usage of the design modules attached to NoC routers. In this work, we develop several new approaches to reduce NoC congestion while minimizing the impact on other design metrics. First, we incorporate a NoC link congestion cost into the placement engine of the open-source CAD flow, versatile place & route (VPR). Second, we integrate turn model NoC routing algorithms into the placement engine to leverage path diversity to further reduce congestion. On average over a suite of 29 benchmarks, combining placement congestion modeling with turn model packet routing reduces NoC congestion by 90.7% at the cost of increasing aggregate bandwidth demand by 4%. In cases where the enhanced placement engine and NoC routing fail to fully resolve congestion, we formulate NoC routing as a Boolean satisfiability (SAT) problem. This approach yields significant additional improvements; the combined algorithm reduces congestion by 95.1% compared to the baseline placement. Finally, we enhance the reinforcement learning (RL) agent in VPR's placement engine by introducing a NoC-aware move type, resulting in an 8.8% reduction in wirelength on designs that make extensive use of the NoC.
- 中文摘要
为了帮助扩展到越来越大、更复杂的设计,最新的FPGA架构现在集成了片上网络(NoC)。NoC帮助在芯片内长距离传输高带宽数据,而无需使用稀缺的低延迟长路由线段。虽然NoC增强型FPGA有助于系统集成和设计重用,但它们也通过引入新的约束和指标,使FPGA计算机辅助设计(CAD)流程变得复杂。布置和路由需要优化NoC指标,如延迟和带宽利用率,避免链路超额订阅(拥塞),同时优化连接NoC路由器设计模块的可编程路由资源使用。在本研究中,我们开发了多种新方法,以减少NoC拥塞,同时最大限度减少对其他设计指标的影响。首先,我们将NoC链路拥堵成本整合到开源CAD流程、多功能布局与路由(VPR)的布置引擎中。其次,我们将转向模型NoC路由算法集成到布置引擎中,利用路径多样性进一步减少拥堵。在29个基准测试中,结合部署拥塞建模与转向模型分组路由,可将NoC拥塞降低90.7%,但总带宽需求增加4%。在增强型布置引擎和NoC路由未能完全解决拥塞的情况下,我们将NoC路由表述为布尔可满足性(SAT)问题。这种方法带来了显著的额外改进;该综合算法相比基线布局减少了95.1%的拥塞。最后,我们通过引入NoC感知的移动类型,增强了VPR布置引擎中的强化学习(RL)代理,使大量使用NoC的设计线长缩短了8.8%。
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
SignalReasoner:评估3B模型在信号数学推理中的上界
- Authors: Guozheng Sun
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17301
- Pdf link: https://arxiv.org/pdf/2608.17301
- Abstract
Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL, a comprehensive benchmark for mathematical reasoning in this domain. We examine two training paradigms: (i) direct reinforcement learning (RL) on WirelessMATHBench-XL with verifiable rewards; and (ii) supervised fine-tuning (SFT) on a distilled wireless-domain chain-of-thought corpus, followed by the same domain-specific RL stage. Across both paradigms, we benchmark Group Relative Policy Optimization (GRPO), Group Sequence Policy Optimization (GSPO), and Geometric-Mean Policy Optimization (GMPO). We aim to assess whether domain-aware CoT SFT serves as an effective initialization for subsequent RL, and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks. Our best model achieves an overall accuracy of 39.12\%, representing a more than threefold improvement over the untrained Base model (12.37\%).
- 中文摘要
通过监督式思维链的精细调优和可验证奖励的强化学习,后期训练显著提升了大型语言模型(LLM)的数学推理能力。然而,它们在信号处理问题中的应用仍然相对缺乏被充分探索。本报告探讨了将Qwen2.5-3B-Base应用于WirelessMATHBench-XL中研究生级信号数学问题的强化微调策略,该系统是该领域数学推理的综合基准。我们考察了两种训练范式:(i)在WirelessMATHBench-XL上的直接强化学习(RL),并提供可验证的奖励;以及(ii)在精简无线领域思维链语料库上的监督微调(SFT),随后进行相同的领域特定强化学习阶段。在这两种范式中,我们对组相对策略优化(GRPO)、组序策略优化(GSPO)和几何均值策略优化(GMPO)进行了基准测试。我们旨在评估域感知型CoT SFT是否能有效初始化后续强化学习,以及GSPO或GMPO在信号推理任务中是否在稳定性或准确性上优于GRPO。我们的最佳模型整体准确率为39.12%,比未训练的基础模型(12.37%)提升了三倍以上。
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
代理型ESOpt:以最小GPU需求微调长视野LLM代理
- Authors: Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17310
- Pdf link: https://arxiv.org/pdf/2608.17310
- Abstract
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale $\sigma$. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.
- 中文摘要
强化学习(RL)在单回合LLM微调方面表现有前景。然而,长视野代理推理引入了越来越多的分支交互和稀疏的奖励,暴露了强化学习的若干局限性:其基于反向传播的重训练栈使得对更大范围的大型语言模型进行微调变得不切实际,且更长的视野轨迹使得在强化学习中分配学分变得极为困难。本文论证进化策略(ES)可能是对长视野LLM代理进行微调的更好选择。与代理式强化学习相比,ES有三大优势:1)模型可扩展性:ES仅用极少的推理级GPU内存即可实现全参数优化,从而实现大型大型语言模型的微调。2)灵活性:其轻量级、黑盒反馈界面使ES易于通过提示空间演进(如技能优化和测试时计算)进行微调;3)长视野可扩展性:ES在不分解不同视野奖励的情况下执行轨迹级参数归因,随着视野长度增长,其可扩展性优于能动强化学习。基于这一见解,我们提出了Agentic ESOpt,一种全参数agentic微调框架,专为灵活参数-上下文共进设计。在每一步,代理型ESOpt会对当前LLM参数周围的扰动进行采样,评估带奖励的代理,并应用在线奖励加权更新。为了改善探索——适应权衡,Agentic ESOpt 进一步引入了扰动尺度 $\sigma$ 的余弦衰变计划。在WebArena-Lite上,Qwen-3.5-27B的全参数优化可提升无技能基线6.69%。在测试时自动启发式设计中,代理ESOpt在线执行提示-参数共演化,提升了36个设置中的28个匹配基线。
Robust Brachiation on a Life-Sized Dual-Arm Robot Using Waypoint-Guided Reinforcement Learning
利用航点引导强化学习对真人大小双臂机器人进行强健的断压
- Authors: Ayumu Iwata, Kento Kawaharazuka, Keita Yoneda, Takahiro Hattori, Kei Okada
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.17320
- Pdf link: https://arxiv.org/pdf/2608.17320
- Abstract
Brachiation is a form of locomotion in which primates move primarily using their arms, enabling traversal in environments without footholds. However, this motion requires highly coordinated whole-body movement and precise timing control for bar grasping and release. As a result, achieving robust behavior on life-sized robotic platforms remains challenging. In this study, we present a reinforcement learning-based method to realize brachiation on a life-sized dual-arm robot. The core of the proposed approach is Waypoint-Guided Reinforcement Learning (WGRL), a learning framework for inducing non-linear and complex motions. For high-difficulty tasks where imitation learning data are unavailable, WGRL guides behavior acquisition by sparsely specifying waypoints for the end-effector trajectory, while whole-body motion is generated through reinforcement learning. In addition, by integrating the waypoint-following guidance with rewards based on task success and mechanical energy, and training in an environment designed for Sim-to-Real transfer, the proposed method achieves both forward progression and motion stability. The acquired behavior is evaluated through Sim-to-Sim experiments under monkey-bar environments with geometric variations and hardware experiments, confirming robust brachiation including failure recovery behavior. This study provides effective learning design guidelines for realizing arm-based locomotion on life-sized robotic hardware and expanding the traversable workspace of robots.
- 中文摘要
弯曲是一种主要依靠手臂移动的运动形式,使灵长类动物能够在没有立足点的环境中进行移动。然而,这种动作需要高度协调的全身动作和精确的时机控制来抓握和释放杠铃。因此,在真人大小的机器人平台上实现稳健行为依然充满挑战。本研究提出了一种基于强化学习的方法,用于实现真人大小双臂机器人的肱骨分离。该方法的核心是航点引导强化学习(WGRL),这是一个用于诱导非线性和复杂运动的学习框架。对于缺乏模仿学习数据的高难度任务,WGRL通过为终点执行器轨迹设定稀疏路径点来指导行为习得,而全身运动则通过强化学习生成。此外,通过将航点跟踪引导与基于任务成功率和机械能的奖励相结合,并在为模拟到现实传递设计的环境中进行训练,所提方法实现了前进和运动稳定性。通过猴子架环境下的模拟对模拟实验和硬件实验评估获得的行为,确认了包括故障恢复行为在内的稳健的断裂。本研究为实现真人大小机器人硬件上的手臂驱动运动提供了有效的学习设计指导,并扩展机器人可移动的工作空间。
Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning
结合新奇与惊喜,以提升基于图像的强化学习中的体验优先级和探索
- Authors: Hoda Yamani, Henry Williams, Bruce A. MacDonald
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17373
- Pdf link: https://arxiv.org/pdf/2608.17373
- Abstract
Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. Improving efficiency requires mechanisms that prioritize informative experiences while also encouraging effective exploration. Prioritized Experience Replay (PER) addresses part of this challenge by reusing high-value transitions, while intrinsic rewards promote the exploration of novel or uncertain states. However, their integration has not been extensively studied. This paper introduces Novelty and Surprise Prioritized Experience Replay (NSPER), which uses novelty to capture underrepresented states and surprise to expose gaps in the agent's understanding of the environment. We further extend this with NSPER+R, integrating these signals as intrinsic rewards to jointly improve replay quality and exploration. Experiments on DeepMind Control Suite tasks show that NSPER and NSPER+R improve training efficiency and convergence speed compared to existing methods in image-based RL.
- 中文摘要
样本效率是强化学习(RL)中的核心挑战,尤其是在基于图像的领域,代理必须从高维视觉输入中学习。传统抽样常依赖随机或次优的体验选择,导致重复更新和学习缓慢。提高效率需要优先考虑信息体验的机制,同时鼓励有效探索。优先体验重玩(PER)通过重用高价值过渡来解决部分挑战,而内在奖励则促进对新颖或不确定状态的探索。然而,它们的整合尚未被广泛研究。本文介绍了新颖与惊喜优先体验回放(NSPER),利用新颖性捕捉代表性不足的状态,并通过惊喜揭示代理对环境理解中的空白。我们进一步扩展了 NSPER+R,将这些信号整合为内在奖励,共同提升重玩质量和探索性。DeepMind Control Suite 任务的实验显示,NSPER 和 NSPER+R 相比基于图像的强化学习方法,能提升训练效率和收敛速度。
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
LEGO-RL:编码代理的绑架-原生强化学习
- Authors: Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17393
- Pdf link: https://arxiv.org/pdf/2608.17393
- Abstract
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.
- 中文摘要
编码代理的强化学习越来越依赖长期运行的代理工具工具框架来管理工具集成、存储库上下文和执行反馈。然而,这些线束的原生执行环境本质上与策略梯度训练不一致:环境崩溃和奖励黑客会破坏结果信号,而列车推断差异则使推广行为与策略更新分离。为此,我们提出了LEGO-RL,一个框架,连接了原生编码代理与可扩展的策略梯度优化,而无需修改其内部控制流。LEGO-RL建立在三大支柱之上:(1)通过进程中的LLM代理进行忠实优化,捕捉原始生成流以实现令牌级对齐,并在训练器端实现稳健的日志概率重计算,即使在机束端的压缩或重序列化下也能实现;(2)通过可扩展的沙盒编排实现可靠执行,包括图像缓存和阶段防御以减轻奖励黑客攻击;以及(3)通过集成插件实现可观察的训练,该插件自动化验证和监控,配合实时用户界面进行细致轨迹诊断。我们通过在三种原生编码器线束上训练稀疏的MoE模型Qwen3.5-35B-A3B,使用GSPO来评估LEGO-RL。LEGO-RL在OpenHands SDK上提升了Qwen3.5-35B-A3B(64.0%至70.4%)、Claude Code(62.4%至68.2%)和OpenCode(57.2%至66.6%)在SWE-bench Verified上的数据,同时保持了推广-训练概率相关系数高于0.99。
REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models
REChart:大型推理模型的高效推理图表编辑
- Authors: Yuanbang Liu, Chenxi Ruan, Yihan Hou, Qiong Luo, Wei Zeng
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Programming Languages (cs.PL)
- Arxiv link: https://arxiv.org/abs/2608.17414
- Pdf link: https://arxiv.org/pdf/2608.17414
- Abstract
Chart editing requires inferring and modifying visualization code from a reference chart image based on an editing instruction, challenging fine-grained visual reasoning, instruction following, and executable code synthesis capabilities of MLLMs. Large reasoning models (LRMs) with extended Chain-of-Thought (CoT) reasoning are suitable for tackling such complex multimodal tasks. However, our preliminary study reveals an
inverted-U'' relationship between reasoning length and chart-editing performance: Excessive reasoning often leads tooverthinking,'' where models drift toward hallucinated visual details or get stuck in redundant reasoning loops. To address the gap, we introduce REChart, a two-stage training framework that provides process-level supervision over intermediate reasoning steps, improving both editing fidelity and reasoning efficiency. First, we synthesize 200k high-quality reasoning trajectories for supervised fine-tuning from a large image-instruction-code pool, using a role-specialized agentic Reason-Score-Refine workflow that iteratively refine the chart code toward higher quality. Second, we optimize the model via reinforcement learning with two complementary rewards: a \emph{fidelity} reward evaluating code correctness, visual fidelity, and structural consistency, and an \emph{efficiency} reward that assigns each rollout a random thinking budget, truncates the reasoning process, and credits the final reasoning segment according to its contribution to the output. On the ChartEdit and ChartMIMIC benchmarks, our model achieves state-of-the-art chart-editing performance among open-source models of comparable scale, while mitigating overthinking and reducing average reasoning token usage by 79.0\% under a maximum thinking budget of 16,384 tokens compared with the base model.
- 中文摘要
图表编辑需要根据编辑指令从参考图表图像推断和修改可视化代码,挑战MLLM的细粒度视觉推理、指令跟踪和可执行代码合成能力。具有扩展思维链(CoT)推理能力的大型推理模型(LRM)适合处理此类复杂的多模态任务。然而,我们的初步研究显示,推理长度与图表编辑表现之间存在“倒U”关系:过度推理常导致“过度思考”,模型会游离于幻觉般的视觉细节或陷入冗余的推理循环。为弥补这一空白,我们引入了REChart,一个两阶段培训框架,提供对中间推理步骤的过程级监督,提升编辑的忠实度和推理效率。首先,我们从一个庞大的图像-指令-代码池中综合20万条高质量推理轨迹进行监督微调,采用角色专用的智能推理-评分-精细工作流,迭代优化图表代码以提升质量。其次,我们通过强化学习优化模型,提供两个互补奖励:\emph{fidelity}奖励评估代码正确性、视觉忠实度和结构一致性,以及\emph{效率}奖励,为每个推广分配随机思维预算,截断推理过程,并根据推理对输出的贡献给最终推理部分。在 ChartEdit 和 ChartMIMIC 基准测试中,我们的模型在同等规模的开源模型中实现了最先进的图表编辑性能,同时在最大思考预算 16,384 个代币下,减少了 79.0% 的平均推理代币使用率,相比基础模型。
Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups
棱镜-GRPO:通过拆分相同结果组实现更快的VLA策略优化
- Authors: Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu, Qing Ping, Han Ding, Guande Wu, Panpan Xu, Jun Huan
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17423
- Pdf link: https://arxiv.org/pdf/2608.17423
- Abstract
GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This simplification comes with a sampling cost: group-relative advantages require multiple rollouts from each scene. Under binary success rewards, groups whose rollouts all succeed or all fail have zero advantage and are discarded by dynamic sampling. These groups are especially common early in training, when most rollouts fail, wasting much of the expensive robotic rollout budget. We introduce Prism-GRPO, which augments binary outcome reward with a weighted trajectory-level execution-quality score. By splitting same-outcome groups into a quality spectrum, Prism-GRPO recovers training signal while ensuring that every success still outranks every failure. Quality scores can be derived from simulator contacts, executed actions, or visual observations, avoiding task-specific progress rewards. We prove that Prism-GRPO never increases the probability that a sampled group is discarded for having zero advantages, and derive a gradient-alignment condition under which its combined update remains a local ascent direction for task success. Across four RoboTwin tasks spanning different horizons and coordination patterns, Prism-GRPO improves success and quality at matched rollout budgets and reaches target success rates with up to 56% fewer rollouts. It also suppresses a reward-hacking shortcut, with the cleaner behavior transferring under direct deployment to a real robot. Through ablations, we show consistent gains across contact-, smoothness-, and VLM-derived quality signals.
- 中文摘要
GRPO越来越多地被用于视觉-语言-行动(VLA)策略的强化学习,因为与PPO不同,GRPO不需要培训批评者。这种简化带来了采样成本:群组相对优势需要从每个场景多次展开。在二元成功奖励下,所有推展成功或全部失败的组没有优势,且通过动态抽样被丢弃。这类小组在培训初期尤为常见,大多数推广失败,浪费了大量昂贵的机器人推广预算。我们引入了棱镜-GRPO,它通过加权轨迹级执行质量评分来增强二元结果奖励。通过将相同结果组划分为质量谱系,Prism-GRPO恢复了训练信号,同时确保每一次成功都优于所有失败。质量评分可以通过模拟器接触、执行动作或视觉观察得出,避免任务特定的进度奖励。我们证明了棱镜-GRPO从不增加因无优势而被丢弃的概率,并推导出梯度对齐条件,使其合并更新仍为任务成功的局部上升方向。在四项跨越不同视野和协调模式的机器人双生任务中,Prism-GRPO 提升了匹配部署预算的成功率和质量,并以多达 56% 的部署次数实现目标成功率。它还抑制了奖励黑客的捷径,更干净的行为会直接转移到真实机器人身上。通过消融,我们在接触、平滑和VLM来源的质量信号上均有持续提升。
Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context
迈向更优秀的多回合用户交互代理:下一次用户回合不仅仅是上下文
- Authors: Yiwen Zhao, Zhihao Wen, Yuchen Mao, Mingxuan Jiang, Yihao Hu, Pan Wang, Xin Zhang, Wei Wu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17499
- Pdf link: https://arxiv.org/pdf/2608.17499
- Abstract
User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repair. The next user turn is more than context: it also provides noisy, temporally local evidence about the preceding user-to-user segment. We introduce \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}), which aligns each reaction with that segment, derives a locally normalized reaction advantage, and adds it to verified terminal outcome advantage without an extra critic or rollout. Against an outcome-only Interactive GRPO control matched in simulator, visible dialogue, initialization, rollout, and optimization, \textsc{FACA} improves the nine-domain $\tau$-family average across three independently trained runs by 5.91 and 10.22 percentage points at 8B and 14B, respectively. Gains concentrate in Telecom; at 8B, randomizing reaction polarity removes the Telecom gain. The same ordering holds zero-shot on Pare-Bench and Co-Gym. These results demonstrate that next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents.
- 中文摘要
面向用户的工具代理必须协调对话和工具使用,随着用户目标在多个回合中展开。然而,交互式强化学习通常将每次推广简化为终端奖励,将同样的功劳归功于有效的诱导、错误和后续修复。下一个用户回合不仅仅是上下文:它还提供了关于前一个用户对用户段的噪声、时间上的局部证据。我们引入了 \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}),它将每个反应与该段对齐,推导出局部归一化的反应优势,并将其添加到经过验证的终极结果优势中,无需额外批评或推广。在模拟器中匹配结果的交互式GRPO对照、可见对话、初始化、展开和优化时,\textsc{FACA}在8B和14B的三个独立训练运行中,分别提升了9个领域$\tau$家族平均5.91个百分点和10.22个百分点。收益集中在电信领域;在8B处,随机化反应极性会去除电信增益。同样的顺序在Pare-Bench和Co-Gym中也适用零中率。这些结果表明,下一回合用户反应为提升多回合用户交互代理提供了可操作的本地信用。
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
评估强化学习可解释性方法对修复代理漏洞的帮助程度
- Authors: Ram Rachum, Yotam Amitai, Bálint Gyevnár, Reuth Mirsky, Cameron Allen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17524
- Pdf link: https://arxiv.org/pdf/2608.17524
- Abstract
This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded proxies like subjective ratings or prediction accuracy. We suggest evaluating XRL methods by how effectively their generated explanations help to diagnose and fix malfunctioning reinforcement learning (RL) agents. We propose EvalXRL, a benchmark in which a Large Language Model (LLM) coding agent uses different XRL methods to diagnose a held-out malfunction in an RL agent, and then repair it. Our proposed benchmark iterates across (environment $\times$ malfunction $\times$ XRL method) tuples and uses the reward signal of the RL agents to form a final score for each XRL method. The coding agent may use the method interactively: invoke the XRL method, process its output, form new hypotheses on what is broken, and invoke the method again with parameters adjusted for testing these hypotheses. This closed-loop structure may be described as a simplified version of the scientific method. Some XRL methods provide self-evaluations that follow this pattern; we propose the first head-to-head comparison of multiple XRL methods in closed-loop usage.
- 中文摘要
本初步论文概述了可解释强化学习(XRL)方法的计划评估基准。当前的评估依赖于功能基础的指标,如忠实度和紧凑性,以及基于人为基础的代理指标,如主观评分或预测准确性。我们建议评估XRL方法,通过其生成的解释对诊断和修复故障强化学习(RL)代理的有效程度来评估。我们提出了EvalXRL,这是一个基准测试,其中大型语言模型(LLM)编码代理使用不同的XRL方法诊断强化学习代理中长期存在的故障,并进行修复。我们提出的基准测试跨越(环境 $\times$ 故障 $\times$ XRL 方法)元组,并利用强化学习代理的奖励信号为每个 XRL 方法生成最终评分。编码代理可以交互式使用该方法:调用XRL方法,处理其输出,对破损内容构建新假设,然后再次调用该方法,参数调整以检验这些假设。这种闭环结构可以被描述为科学方法的简化版本。一些XRL方法提供遵循此模式的自我评估;我们提出了闭环应用中多种XRL方法的首次正面比较。
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
tinyDSM:资源受限毫米机器人的技能建模与开发框架
- Authors: Markus D. Kobelrausch, Michael Miedler, Axel Jantsch
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.17596
- Pdf link: https://arxiv.org/pdf/2608.17596
- Abstract
In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.
- 中文摘要
本研究探讨了使得小型、资源有限的系统(如厘米大小的毫体机器人)能够自主探索、学习和适应其能力的发展机制。强化学习算法通过我们提出的tinyDSM的相互作用,引导代理的技能习得和适应,该系统整合了内在动机和基于适应度的评估。我们追求极简、固定的技能,同时鼓励开放式新技能的发展。我们方法的一个关键强调是编码最少的先验通用知识,这作为系统从初始知识中进一步学习系统特定依赖关系的基础起点。因此,我们的方法设计上试图涵盖非常通用的应用领域。该方法论基于(a)具有内在动机的发展机制,以及(b)认知架构(知识、推理、学习),同时(c)利用最小的资源。它利用层级知识图谱和运动学推理器来建模和评估简单和高级的运动相关技能。在我们的实验中,我们使用体积为36 cm^3的资源受限 millirobot,配备 Raspberry Pi Pico 32 位微控制器(RP2040),该控制器集成了除相机系统外的所有描述功能和能力,且为 9 kB。从学习最基础的运动技能开始,millirobot 在 15 分钟内自主地从简单的线性和角度运动进步到复杂的几何图案。为补充物理实验,我们进行了基于仿真的分析,实现了学习算法和内在动机参数之间的系统比较。
Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision
迭代抓握姿态精炼:二维视觉的深度强化学习方法
- Authors: Amir Arsalan Nematollahi, Shayan Ahmadi, Mehdi Tale Masouleh, Ahmad Kalhor
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
- Arxiv link: https://arxiv.org/abs/2608.17628
- Pdf link: https://arxiv.org/pdf/2608.17628
- Abstract
Developing robots capable of understanding and manipulating objects requires compact, interpretable, and generalizable representations. This work proposes a reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN). Using 2D overhead images captured in a simulated environment, a geometric-based algorithm generates initial grasp candidates, which are iteratively refined by the proposed framework, transforming failed grasps into successful ones. Experiments conducted on 300 objects from the Dex-Net dataset using a UR5 manipulator demonstrate the framework's effectiveness, achieving a 100% success rate on objects previously deemed ungraspable by geometrical methods. The framework's sim-to-real transferability is further validated through physical experiments on a Delta parallel robot, where a refined grasp successfully manipulates an object that was previously ungraspable. The findings underscore the effectiveness of reinforcement learning in addressing challenges in robotic grasping, offering a scalable and adaptable solution for contact-rich manipulation tasks.
- 中文摘要
开发能够理解和操作物体的机器人需要紧凑、可解释且可推广的表征。本研究提出了一个基于强化学习的机器人抓取精细框架,将基于关键点的对象表示与深度Q网络(DQN)集成。利用在模拟环境中捕捉的二维顶置图像,基于几何的算法生成初始抓取候选,这些候选抓取经由所提出的框架迭代完善,将失败的抓取转化为成功抓取。使用UR5操作器对Dex-Net数据集中的300个对象进行的实验证明了该框架的有效性,对此前几何方法认为无法抓握的对象实现了100%的成功率。该框架的模拟到现实的可转移性通过在Delta平行机器人上的物理实验进一步验证,在该实验中,精细的抓握成功操作了此前无法抓握的物体。研究结果强调了强化学习在解决机器人抓取挑战中的有效性,为接触丰富操作任务提供了可扩展且适应性的解决方案。
rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment
rl-triton:用于强化学习学分作业的高性能Triton GPU内核
- Authors: Lars Simon Zehnder
- Subjects: Subjects:
Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF)
- Arxiv link: https://arxiv.org/abs/2608.17641
- Pdf link: https://arxiv.org/pdf/2608.17641
- Abstract
We present rl-triton, an open-source library of high-performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. The core contribution is a unified associative scan framework that recasts seven distinct RL estimation algorithms - Generalized Advantage Estimation (GAE), V-Trace, Retrace($\lambda$), TD($\lambda$) returns, discounted returns, eligibility traces, and episodic prefix sums - as instances of a single first-order linear recurrence solved in $O(\log T)$ parallel steps. All algorithms share the same associative scan operator, with algorithm-specific fused Triton kernels constructing their recurrence coefficients on-chip. We verify the associative operator algebraically and define the treatment of terminated and truncated episodes explicitly. Benchmarks show a 1.6-5.70$\times$ full-call speedup over a vectorized this http URL baseline in the massively parallel simulation regime (thousands of environments, short rollouts). The reported range covers all seven algorithms on both GPUs, both with and without per-step truncation handling. For most algorithms, speedups increase at longer sequence lengths, as the baseline requires more scan stages as $\log T$ grows, each adding an intermediate HBM round-trip. The library is available at this https URL.
- 中文摘要
我们介绍rl-triton,一个开源的高性能GPU内核库,用于强化学习学分分配,采用Triton实现。核心贡献是一个统一的联想扫描框架,将七种不同的强化学习估计算法——广义优势估计(GAE)、V-trace、Retrace($\lambda$)、TD($\lambda$)返回、贴现返回、资格追踪和情节前缀和——作为单一一阶线性递归的实例,在$O(\log T)$平行步骤中解决。所有算法共享相同的结合扫描算子,算法专用的融合特里同核在芯片上构建其复发系数。我们通过代数验证了结合算子,并明确定义了终止和截断插曲的处理方式。基准测试显示,在大规模并行仿真(数千个环境,短暂部署)中,整个调用速度比向量化的HTTP URL基线提升了1.6到5.70美元\时间。报告范围涵盖了两款GPU上所有七种算法,包括带步截断处理和不支持的。对于大多数算法,加速在序列长度越长时越大,因为基线需要更多扫描阶段,随着 $\log T$ 的增长,每个阶段都会增加中间的 HBM 往返。该库可通过此 https URL 访问。
Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control
基于物理知情的世界模型进行合作混合交通控制的离线多智能体强化学习
- Authors: Lu Liu, Chi Xie, Xi Xiong
- Subjects: Subjects:
Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2608.17739
- Pdf link: https://arxiv.org/pdf/2608.17739
- Abstract
This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We propose a physics-informed world model-based offline multi-agent reinforcement learning framework that reconstructs a physically interpretable global traffic state from local CAV observation-action histories, with coupled macroscopic-microscopic traffic dynamics providing physics-based supervision. A probabilistic ensemble world model learns traffic-state transitions and system rewards, while model disagreement quantifies epistemic uncertainty. Multi-step imagined rollouts with pessimistic rewards and uncertainty-driven truncation are then used for offline policy learning. Experiments in a SUMO-based on-ramp bottleneck using approximately $1\times10^6$ offline transitions show that physics supervision improves state reconstruction and world-model prediction accuracy.
- 中文摘要
本研究探讨了在混合交通中部分可观测高速公路瓶颈处协同控制联网与自动驾驶车辆(CAV),旨在缓解拥堵,而无需依赖完整的全球交通状态或在线试错。我们提出了一个基于物理数据的世界模型、离线多代理强化学习框架,能够从本地CAV观测-动作历史重建物理可解读的全局交通状态,并结合宏观-微观交通动态,提供基于物理的监督。概率性集合世界模型学习交通状态转变和系统奖励,而模型不一致则量化认识论不确定性。多步想象式推广,奖励悲观且不确定性驱动的截断,随后用于离线政策学习。基于SUMO的匝道瓶颈实验,使用约$1\times10^6$的离线转换,表明物理监督能提升状态重建和世界模型预测准确性。
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
用低资源语言思考:SFT构建了什么,强化学习修复了什么,准确性看不到什么
- Authors: Ayoub Kirouane, Christos Petrocheilos
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG); Robotics (cs.RO); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2608.17744
- Pdf link: https://arxiv.org/pdf/2608.17744
- Abstract
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
- 中文摘要
拿三个前沿专家混合模型(阿里巴巴、OpenAI、NVIDIA;每个3.6-4.0B活跃参数)进行微调,用低资源语言推理。在准确度基准测试中几乎没有变化,基准本身在这个尺度上就是噪声:只更换随机种子,分数会提升7.7个百分点,超过我们测量的所有数据和配方效应。那个零点是我们的第一个结果。真正的变化发生在无法准确看见的地方。基础模型从不用希腊语思考:即使问题是希腊语,1000条推理痕迹中也只有0个,因此模型在用用户无法阅读、审计或纠正的推理形式中正确回答。经过监督微调(SFT),每个检查点都以题目语言发布,涉及~98%的题目,一个族以少3倍的词,四个模型的语法判断性提升,整体能力在每个基础基础的几个点内有所提升:没有遗漏,流利度得以提升。我们提出了六个行为维度使这些变化可测量,每个维度都被限制为拒绝任何与输出长度相关的指标,并报告了我们自身工具的谎言:六个失败,每个失败都被对照组捕获。SFT无法解决的是自身缺陷:四分之一的答案跳过了请求的格式,答案会渗入推理通道,明确的“用英语思考”被遵守的时间不到一半。带有可验证奖励、训练前预注册的强化学习,能彻底修复前两项(回退24%至2.5%,泄漏率3.5%至0.0%,均针对固定随机奖励对照),并移动第三项(+9.1pp),而希腊式推理习惯则能在仅准确度的梯度中保持完整。我们会释放五个检查点。仪器、控制和预注册可传输至任何低资源语言;希腊的案例让我们来衡量它们。
Debate Training Reduces Reward Hacking in RLAIF
辩论培训减少RLAIF中的奖励黑客行为
- Authors: Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17776
- Pdf link: https://arxiv.org/pdf/2608.17776
- Abstract
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its AI judge, degrading task performance, a problem that worsens precisely when the judge is weaker than the policy, the setting most relevant to overseeing increasingly capable AI systems. We study mathematics tasks, where final-answer correctness is verifiable, allowing us to measure reward hacking dynamics. We train a Gemini~2.5 Flash-class policy with a frozen, weaker Gemini~2.5 Flash Lite judge, comparing a single-player RLAIF baseline against debate. While the baseline quickly hacks the judge, debate maintains judge performance throughout training, leading to a higher peak validation accuracy (45\% performance gap recovered) that persists through many RL steps. Additional experiments show that: 1) further weakening the judge leads to faster hacking, but this can be compensated by adding an additional debate round; 2) debate incentives override prompted misalignment; 3) RL using an LLM judge has a smaller train/validation reward gap than RL from verifiable rewards; 4) learning to critique to convince the judge using ground truth labels is possible but slow. Taken together, our results are a positive update on the feasibility of debate, while highlighting that balancing multi-agent training is critical: without player constraints, adversarial training risks defaulting to critic judge-hacking. We show that critique word limits (effective up to 150 words) successfully balance the game and avoid judge hacking, though this introduces a trade-off by restricting critic expressive clarity.
- 中文摘要
我们证明,强化学习通过辩论微调LLM(一场由较弱的LLM评审者裁决的两人对抗游戏,由生成器与批评者对抗)进行,相比于基于AI反馈强化学习(RLAIF)基线,减少了奖励黑客行为。奖励黑客是RLAIF的核心障碍:随着训练的推进,政策学会利用AI评判的系统性错误,降低任务表现,这一问题在评审比策略弱时加剧,而策略正是监管日益强大AI系统的背景。我们研究数学任务,其中最终答案的正确性是可验证的,从而测量奖励黑客的动态。我们用一个冻结、较弱的Gemini~2.5 Flash Lite评判来训练Gemini~2.5 Flash类政策,将单人RLAIF基线与辩论进行比较。虽然基线很快就能破解评委,但辩论在整个训练过程中都能维持评委的表现,从而实现更高的峰值验证准确率(恢复了45%的性能差距),这种差距在多个强化学习阶段都能持续。其他实验显示:1)进一步削弱裁判会导致黑客更快,但可以通过增加一轮辩论来弥补;2)辩论激励覆盖导致的不一致;3)使用LLM评审的强化学习在可验证奖励方面比强化学习的训练/验证奖励差距更小;4)学会用真实标签来说服评委是可能的,但进展缓慢。综合来看,我们的结果是对辩论可行性的积极更新,同时强调了平衡多智能体训练至关重要:没有玩家约束,对抗训练有可能陷入批评裁判黑客攻击。我们展示了批评字数限制(有效于150字以内)成功平衡了游戏,避免了评委的黑客行为,尽管这会带来权衡,限制了评论家的表达清晰度。
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
通过图结构在线难度估算实现高效的RLVR调度
- Authors: Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.17941
- Pdf link: https://arxiv.org/pdf/2608.17941
- Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.
- 中文摘要
带有可验证奖励的强化学习(RLVR)提升了大型语言模型的推理能力,但依赖于昂贵的推广探索。将相同的探索预算分配给不同难度的样本效率低下:简单的样本可能会被重复推出,而困难但可学习的样本可能探索不足。现有的自适应调度器通过基于课程的样本选择或基于估计样本难度的非统一推广分配来解决这一不匹配问题。然而,获得可靠的在线难度估算仍然具有挑战性:专门的探测会增加大量生成开销,而基于历史的估计器则面临冷启动,没有初始观测和陈旧反馈,且通常忽略样本间的关系。为解决这些限制,我们提出了一种即插即用的基于图表的在线难度估算器,能够在相关样本间共享推广反馈,并持续更新难度估算,从而在无需专门探测的情况下缓解冷启动和停滞。具体来说,我们首先基于语义和推理相似性构建一个难度感知的样本图。基于该图,我们引入潜在困难状态,并使用Potts先置法鼓励邻近样本共享相同状态。随后,我们采用州级Beta-二项模型,汇总每个州的推广结果。最后,我们使用在线均值场变分算法,随着新反馈的到来不断更新潜态赋值和状态级难度。我们的框架可集成到样本选择和推广分配调度器中,实现难度自适应探索,无需专门探查。跨多个基模型、强化调度器和基准测试的实验表明,我们的框架实现了更好的性能。
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
迈向零样子任务转移,采用神经符号世界模型
- Authors: Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.17959
- Pdf link: https://arxiv.org/pdf/2608.17959
- Abstract
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.
- 中文摘要
最先进的基于模型的强化学习方法学习神经世界模型,这些模型允许在潜在空间中规划政策,而无需假设底层环境的结构。虽然具有表现力,这些模型通常依赖任务:它们学习与训练任务相关的不可解释潜在表征,因此难以推广到新任务。在本研究中,我们提出了一种新颖的世界模型表述,其中奖励预测仅依赖于整个潜伏状态中结构化、符号化成分的子集。解耦观察重建和奖励预测使我们能够学习能够将零射值(即无需环境相互作用)适应同一符号状态空间上定义的新奖励函数的世界模型。我们讨论了学习这些神经符号世界模型的主要优势和挑战,并展示了我们方法相较于纯神经方法的强大泛化特性。
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
基于LLM反馈的策略不变奖励塑造:混合强化学习代理框架
- Authors: Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.18008
- Pdf link: https://arxiv.org/pdf/2608.18008
- Abstract
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.
- 中文摘要
将大型语言模型与强化学习结合的做法越来越多,但LLM衍生的奖励信号的理论地位往往被隐含。我们将混合的LLM规划器和RL控制器架构形式化为目标增强马尔可夫决策过程,并证明当LLM的状态进展分数作为有界势函数使用时,所得的整形项即使在LLM分数不准确时仍保持最优策略集。这种保障比一般的LLM即奖励方法更强。我们在一个小型MDP上以四种潜在配置(包括一个对抗配置,其比例为基础奖励幅度的二十倍)进行数值验证。
Keyword: diffusion policy
There is no result