生成时间: 2026-07-22 18:27:13 (UTC+8); Arxiv 发布时间: 2026-07-22 20:00 EDT (2026-07-23 08:00 UTC+8)
今天共有 38 篇相关文章
Keyword: reinforcement learning
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
S2T-RLHF:稳定优先RLHF的分层学分分配
- Authors: Wei Chen, Guanghui Zhu, Yafei Li, Limin Wang, Yihua Huang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18258
- Pdf link: https://arxiv.org/pdf/2607.18258
- Abstract
Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.
- 中文摘要
基于偏好的奖励模型的人类反馈强化学习(RLHF)常常表现出不稳定的训练动态。一个关键因素是标准RLHF依赖单一序列层级标量奖励,该奖励传递至代币级策略更新,使响应中的信用分配本质上存在模糊性。近期研究试图通过将奖励细化为更密集的代币级监督来解决这一问题,通常依赖于更细粒度的信用分配能提升优化的隐含假设。我们认为这一假设不完整:当偏好信号噪声大且仅在反应水平定义时,过于细粒度的奖励细化会放大奖励不确定性并破坏学习稳定性。为解决此问题,我们提出了一种细度感知的层级学分分配原则,强调以稳定性为导向的奖励设计,而非最大化分配的精确度。根据这一原则,句子作为自然的中间粒度,平衡语义连贯性与对令牌级噪声的鲁棒性。基于这一观点,我们介绍了S2T-RLHF。该句子到代币的奖励分解框架首先在句子间分配序列层级的偏好奖励,然后在每个句子内应用有界的代币级细化,无需奖励模型重训练或代币级监督。跨多个数据集和优化环境的实验表明,S2T-RLHF在保持竞争偏好对齐的同时,提升了训练的稳定性和鲁棒性。
Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks
多时间尺度的潜在作用持续DLL用于边缘云网络中的联合优化
- Authors: Vo Phi Son, Van-Dinh Nguyen, Ngoc Hung Nguyen, Trinh Van Chien, Symeon Chatzinotas
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.18288
- Pdf link: https://arxiv.org/pdf/2607.18288
- Abstract
Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization. To address this challenge, we study a joint service placement, computational delegation, and power control (JSCP) problem to minimize the average end-to-end (e2e) latency. The resulting JSCP problem is a mixed-integer nonconvex and NP-hard optimization problem due to the strong coupling between discrete and continuous variables. To enable tractable optimization and stable system adaptation, we exploit the inherent difference in decision dynamics and decompose the problem into long-term system configuration and short-term resource allocation subproblems. Based on this formulation, we propose a two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) to jointly optimize service placement, user association, computational delegation, task offloading, and user transmit power. A latent action representation based on a variational autoencoder is introduced to efficiently compress the high-dimensional combinatorial action space. Simulation results demonstrate that the proposed framework effectively adapts to dynamic network conditions and achieves near-optimal performance compared to branch-and-bound solutions. It achieves up to a 20.8% reduction in average e2e latency and a 13% improvement in resource utilization over the scheme without the computational delegation, while converging approximately 50% faster than conventional proximal policy optimization.
- 中文摘要
在动态任务到达和资源异构的情况下,跨层和云层的负载不平衡会降低分层边缘云计算(HECC)系统的延迟性能,导致严重的排队延迟和资源利用效率低下。为应对这一挑战,我们研究了联合服务部署、计算委派与电源控制(JSCP)问题,以最小化平均端到端(e2e)延迟。由此产生的JSCP问题是一个混合整数非凸和NP难优化问题,这源于离散变量与连续变量之间的强耦合。为了实现可处理的优化和稳定的系统适应,我们利用决策动力学的固有差异,将问题分解为长期系统配置和短期资源分配子问题。基于该表述,我们提出了一个具有潜在动作空间(2T-MDRL-LA)的两时间尺度多层深度强化学习框架,以联合优化服务配置、用户关联、计算委托、任务卸载和用户传输能力。引入了基于变分自编码器的潜在作用表示,以高效压缩高维组合作用空间。模拟结果表明,所提出的框架能够有效适应动态网络条件,并且相较于分支定界解实现了接近最佳的性能。在不使用计算委派的情况下,该方法平均e2e延迟可降低多达20.8%,资源利用率提升13%,收敛速度约比传统近端策略优化快50%。
Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal
深度强化学习以掌握巴格查尔的非对称策略
- Authors: Ranjit Raut, Aarav Subedi, Sagun Rai, Aaryan Shakya, Manoj Shakya
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT)
- Arxiv link: https://arxiv.org/abs/2607.18296
- Pdf link: https://arxiv.org/pdf/2607.18296
- Abstract
Baghchal is a two-player asymmetric board game with Nepali origins where four tigers are to capture goats and twenty goats desire to keep tigers in immobility. Although Baghchal has a complex structure which is strategic, has perfect information structure, and has cultural meaning, it has not been adequately covered in deep reinforcement learning (RL) literature. This paper gives a systematic exploration of four deep RL solutions Deep Q-Network (DQN), REINFORCE, Proximal Policy Optimization (PPO) and MuZero that are trained on one side of the asymmetric gameplay of Baghchal and then evaluated on the other side. The algorithms are rated based on win rate, draw rate, average captures, training convergence and computational cost. It is experimentally found that MuZero generates the best performance in both tasks, achieving 86 percent win over these Tiger and 62 percent win over these Goat and the ability to do so is due to the model-based planning machine through the Monte Carlo Tree Search. PPO is the most realistic algorithm and is provided to be competitive over both asymmetric tasks with significantly reduced computational costs compared to MuZero. Emergent strategic behavior analysis shows that model-based strategies are optimal over long-horizon planning, whereas value-based counterparts like DQN are more biased up towards the Tiger role owing to the more substantial reward signal.
- 中文摘要
Baghchal 是一种起源于尼泊尔的双人非对称棋盘游戏,四只老虎用来捕捉山羊,二十只山羊则想让老虎保持行动不动。尽管巴格查尔拥有复杂的战略结构、完美的信息结构和文化意义,但在深度强化学习(RL)文献中尚未充分涉及。本文系统地探讨了四种深度强化学习解决方案:深度Q网络(DQN)、REINFORCE、近端策略优化(PPO)和MuZero,这些方案在Baghchal非对称玩法一侧训练,另一侧进行评估。算法的评分基于胜率、抽球率、平均捕获次数、训练收敛性和计算成本。实验发现,MuZero在这两项任务中都产生了最佳性能,对这些“老虎”的胜率为86%,对这些“山羊”的胜率为62%,而这种能力得益于基于模型的规划机通过蒙特卡洛树搜索实现。PPO 是最现实的算法,旨在在两种非对称任务中具备竞争力,计算成本显著低于 MuZero。新兴战略行为分析表明,基于模型的策略在长期规划中最优,而价值导向的对应方法如DQN则更倾向老虎角色,因为奖励信号更为显著。
Decentralized Multi-agent Reinforcement Learning for Resilient Critical Infrastructures
针对韧性关键基础设施的去中心化多智能体强化学习
- Authors: Minghui Ding, Evangelos Pournaras
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.18359
- Pdf link: https://arxiv.org/pdf/2607.18359
- Abstract
Critical infrastructures are increasingly distributed, interdependent, and exposed to evolving disruptions, making resilience a central requirement for their operation and control. This paper argues that decentralized multi-agent reinforcement learning (MARL) should be understood not merely as a distributed alternative to centralized training with decentralized execution but as a paradigm structurally aligned with the requirements of resilient critical infrastructures. This perspective is grounded in an analysis of the properties of decentralized MARL and the requirements of critical infrastructures, including scalability to large numbers of agents, support for privacy and local autonomy, robustness to failures, and interaction-driven adaptation among interdependent components. However, structural alignment alone is insufficient for practical deployment. This paper identifies credit assignment and communication as two central conditions for its practical feasibility. Credit assignment determines whether local learning remains aligned with system-level objectives, while communication determines whether coordination can be learned and maintained under realistic operational constraints. Building on these challenges, this paper proposes a research agenda focused on structure-aware, causality-aware, and resilience-aware credit assignment; communication for both coordination and credit assignment; and safe, timely, and recoverable decentralized learning under deployment constraints. Overall, this paper reframes decentralized MARL as a promising but conditional foundation for resilient critical infrastructures.
- 中文摘要
关键基础设施日益分散、相互依赖,并面临不断变化的干扰,因此韧性成为其运行和控制的核心要求。本文主张,去中心化多智能体强化学习(MARL)不应仅仅被视为去中心化训练的分布式替代方案,而应被视为一种结构上与韧性关键基础设施需求相契合的范式。这一观点基于对去中心化MARL属性及关键基础设施需求的分析,包括对大量代理的可扩展性、隐私和本地自治的支持、对故障的鲁棒性,以及相互依赖组件间的交互驱动适应。然而,仅靠结构对齐不足以实现实际部署。本文指出学分分配和沟通是其可行性的两个核心条件。学分分配决定本地学习是否与系统目标保持一致,而沟通决定协调是否能在现实操作约束下学习和维持。基于这些挑战,本文提出了一个以结构意识、因果关系意识和韧性意识为核心的研究议程;协调和学分分配的沟通;以及在部署约束下安全、及时且可恢复的去中心化学习。总体而言,本文将去中心化MARL重新定位为一个有前景但有条件的韧性关键基础设施基础。
FARO: Feasibility-Aware Robot Motion Optimization
FARO:可行性感知机器人运动优化
- Authors: Michal Ciebielski, Shafeef Omar, Aaron Johnson, Majid Khadiv
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.18362
- Pdf link: https://arxiv.org/pdf/2607.18362
- Abstract
Fast planning of novel behaviors in unseen scenarios remains a fundamental challenge in robotics. The high-dimensional, hybrid, and underactuated nature of humanoid loco-manipulation continues to hinder the realization of this goal. In this paper, we address this challenge by proposing a nested kino-dynamic framework for rapid feasibility checking and dynamically consistent trajectory generation given a candidate contact sequence. By integrating this module with a feasibility-guided tree search and a Large Language Model (LLM)-based contact plan sampling strategy, we demonstrate that the proposed framework can substantially improve the search process. Furthermore, we show that the generated trajectories can be tracked using a reinforcement learning (RL)-based controller and show that the resulting trajectories are of sufficiently high quality for execution in real-world loco-manipulation scenarios. A supplementary video is available at: this https URL.
- 中文摘要
在未见场景中快速规划新颖行为仍是机器人学中的根本挑战。人形机车操控的高维、混合且驱动不足的特性,持续阻碍这一目标的实现。本文通过提出一个嵌套的运动动力学框架,用于快速可行性检查和在候选接触序列下动态一致的轨迹生成,来应对这一挑战。通过将该模块与可行性引导树搜索和基于大型语言模型(LLM)的联络计划抽样策略集成,我们证明所提出的框架能够显著改善搜索过程。此外,我们证明了生成的轨迹可以通过基于强化学习(RL)的控制器跟踪,并证明这些轨迹在真实机车操作场景中具有足够高的质量。补充视频可在此处观看:https 网址。
Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion
迈向四足行走的扭矩驱动强化学习
- Authors: Jordan Dowdy, Jean Chagas Vaz
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2607.18365
- Pdf link: https://arxiv.org/pdf/2607.18365
- Abstract
Reinforcement learning (RL) for legged robots is advancing locomotion, demonstrating its ability to adapt to new and challenging terrain. Traditionally, these RL locomotion frameworks are position-based, making the policy less adaptable to terrain types and requiring state estimation techniques in the observation space, i.e., linear velocity. Moreover, these RL frameworks often use small, lightweight quadrupeds that are limited in their viability for high-complexity tasks due to hardware constraints. This work explores an RL torque control framework for heavyweight high-torque quadrupeds. The RL framework in this paper can traverse rough terrain and effectively track a desired linear velocity without requiring knowledge of the agent's current velocity. Using Nvidia's Isaac Sim and Isaac Lab, simulation results of the RL torque control policy are shown on the Unitree B1 quadruped, achieving speeds of 3.5 m/s and 1.5 rad/s. In addition, the quadruped can walk up and down stairs without the aid of an exteroceptive sensor.
- 中文摘要
腿部机器人的强化学习(RL)正在推动运动技术的发展,展示了其适应新环境和挑战性环境的能力。传统上,这些强化学习的移动框架基于位置,使得该策略对地形类型的适应性较低,且需要在观测空间中进行状态估计技术,即线速度。此外,这些强化学习框架通常使用小型轻量化的四足动物,但由于硬件限制,其在高复杂度任务中的实用性有限。本研究探讨了重型高扭矩四足动物的强化扭矩控制框架。本文中的强化学习框架可以穿越崎岖地形,并有效追踪所需的线速度,而无需了解代理当前的速度。利用英伟达的Isaac Sim和Isaac Lab,RL扭矩控制策略的模拟结果在Unitree B1四足机上展示,实现了3.5 m/s和1.5 rad/s的速度。此外,四足动物可以在没有外感传感器辅助的情况下上下楼梯。
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
部分可观察性下时间知识图记忆的神经符号元政策
- Authors: Taewoon Kim, Vincent François-Lavet, Michael Cochez
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18368
- Pdf link: https://arxiv.org/pdf/2607.18368
- Abstract
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.
- 中文摘要
部分可观察的强化学习需要决定随着时间推移要保留、取回和遗忘的内容。我们引入了一种神经符号元策略,学习在每个决策点应用哪种符号记忆启发式,同时保持执行符号化。我们的环境在RoomKG中使用时间知识图记忆,隐藏状态和观察以资源描述框架(RDF)图表示,记忆则通过时间RDF三重注释进行补充。该模型结合了记忆内容的知识图谱编码与用于问题解答、探索和遗忘的值头,从而实现了一个既自适应又可检查的控制器。这为该工作提供了直接的语义网基础,通过基于RDF的表示、与注释兼容的图语义语义以及基于图的符号操作,对显式内存状态进行操作。在列车/测试室的长期记忆容量为512时,符合限定符的StarE-GNN配置在符号、神经和神经符号系统中实现了最佳的保留性能,同时保持了记忆管理决策的步级可追溯性。
RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
RRPO:参考相对策略优化与分层条件推广
- Authors: Yuxin Xiong, Xunyi Jiang, Rohan Surana, Xintong Li, Sheldon Yu, Nikki Lijing Kuang, Ryan A. Rossi, Jingbo Shang, Tong Yu, Julian McAuley, Junda Wu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18470
- Pdf link: https://arxiv.org/pdf/2607.18470
- Abstract
Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled rollouts can be compared within a group using task-provided correctness signals. However, extending group-relative optimization beyond verifiable settings is challenging because success in many tasks is not captured by a single correctness criterion. We propose \textbf{Reference-Relative Policy Optimization (RRPO)}, which generalizes GRPO by replacing direct correctness-based advantage construction with reference-relative contrastive comparisons. RRPO first uses \emph{stratified conditional rollouts} to construct positive and negative anchor sets, and then trains a metric projection head with a set-contrastive objective to compare candidate rollouts against these anchors. The resulting alignment scores directly define contrastive advantages: during policy optimization, the projection head is frozen, and the scores are centered within each rollout group in a standard group-relative objective. We evaluate RRPO using anchor-based contrastive advantages throughout policy optimization, without relying on task ground-truth verifiers. Across verifiable reasoning, open-ended generation, and post-SFT settings, RRPO remains competitive with verifier-based optimization, improves over weakly supervised baselines, and provides additional gains after supervised fine-tuning.
- 中文摘要
群体相对策略优化(GRPO)在基于可验证反馈的强化学习中表现出强效,该方法可利用任务提供的正确性信号在群体内比较抽样的推广。然而,将群相对优化扩展到可验证的设置之外具有挑战性,因为许多任务的成功并非仅由单一的正确性标准所涵盖。我们提出了 \textbf{参考相对策略优化(RRPO)},通过用参考相对对比比较替代直接基于正确性的优势构建来推广 GRPO。RRPO首先使用\emph{分层条件滚动}构建正负锚点集合,然后训练一个带有集合对比目标的度量投影头,以比较候选推演与这些锚点。由此产生的对齐分数直接定义了对比优势:在策略优化过程中,预测头被冻结,分数集中在每个推广组内,并以标准的组相对目标计算。我们在策略优化过程中采用基于锚点的对比优势来评估RRPO,而不依赖任务真实验证器。在可验证推理、开放式生成和后SFT环境中,RRPO仍能与基于验证者的优化竞争,优于弱监督基线,并在监督微调后提供额外收益。
Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
Search-on-Graph-R1:用强化学习训练大型语言模型搜索知识图谱
- Authors: Jia Ao Sun, Hao Yu, Fengran Mo, Zhan Su, Yuchen Hui, Bang Liu, Jian-Yun Nie
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.18481
- Pdf link: https://arxiv.org/pdf/2607.18481
- Abstract
Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontier-scale inference makes them costly to deploy. We present Search-on-Graph-R1 (\sogrone{}), which internalizes this navigation into a compact 8B model through supervised fine-tuning (SFT) followed by reinforcement learning (RL). Our central idea is to scaffold a frontier teacher with each question's gold SPARQL query, so the teacher traverses a known answer-bearing path with a live \texttt{Search} tool rather than having to discover the path itself. Since every call executes against a live Freebase server, the resulting trajectories are grounded in the knowledge graph by construction. On WebQSP, CWQ, and GrailQA, \sogrone{} at 8B surpasses every frozen frontier-LLM system in our comparison and posts the strongest results on CWQ of any system we compare against. It does so using no auxiliary module at inference and no LLM judge during training. Isolating each training stage shows that SFT and RL contribute complementary gains, our approach transfers across model families, and RL learns to reach answers in fewer \texttt{Search} calls than its SFT initialization.
- 中文摘要
知识图谱问题解答(KGQA)需要从主题实体导航到几个关系之外的答案。近期方法促使前沿大型语言模型通过检索工具探索该图,但其依赖前沿尺度推断使得部署成本较高。我们介绍了Search-on-Graph-R1(\sogrone{}),它通过监督微调(SFT)和强化学习(RL)将该导航内化为紧凑的8B模型。我们的核心想法是用每题的黄金SPARQL查询来支撑前沿教师,让教师通过实时的\texttt{Search}工具走过已知的答案路径,而不必自己发现路径。由于每次调用都是在实时的Freebase服务器上执行的,因此由此产生的轨迹通过构造基于知识图谱。在WebQSP、CWQ和GrailQA上,\sogrone{}在8B的对比中超过了所有冻结的前沿大型语言模型系统,并且在CWQ上发布了我们对比的所有系统中最强的结果。它在推理时不使用辅助模块,培训期间也没有LLM评审。隔离每个训练阶段表明,SFT和RL互补地贡献收益,我们的方法跨模型族迁移,且RL学会在比SFT初始化更少的\texttt{Search}调用中获得答案。
The Open Ant: A Robot Platform for Reinforcement Learning Research
开放蚂蚁:强化学习研究机器人平台
- Authors: Elena Sorina Lupu, Patrick Spieler, Khurram Javed, Kris De Asis, John D. Martin, Martha Steenstrup, Joseph Modayil
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2607.18488
- Pdf link: https://arxiv.org/pdf/2607.18488
- Abstract
Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations. The predominance of simulations makes translating research to physical reality uncertain for both algorithms and researchers. We propose a physical platform that is designed to simplify the transition. In this paper, we present the Open Ant: a physical variant of the commonly used Gymnasium Ant environment, along with a simulation. We demonstrate that competent walking policies can be learned from scratch in approximately one hour directly from the physical robot's experience for two substantially different RL algorithms: SARSA($\lambda$) and Soft Actor-Critic (SAC). Separately, we show policies that were learned in simulation transfer to reality. We also examine how well the platform supports a nimble experimental ecosystem. Specifically, we observe the speed with which new users from diverse backgrounds achieve their first success with the platform, and how easily the platform can be repaired and updated when hardware issues arise. Both the hardware design and software are available as open-source on GitHub for ease of customization. In summary, we advocate for the use of the Open Ant for RL researchers who frequently use simulated environments, so they can more easily include robot experiments in their evaluations.
- 中文摘要
强化学习(RL)研究在物理和模拟领域均取得了成功;然而,主流方法仍根植于模拟。模拟的主流使得算法和研究人员都难以将研究转化为物理现实。我们提出一个实体平台,旨在简化过渡过程。本文介绍了开放蚂蚁:一种常用的体育馆蚂蚁环境的物理变体,并附带一个模拟。我们展示了,在两种截然不同的强化学习算法(SARSA($\lambda$)和软演员-批判者(SAC)中,大约一小时内可以直接从零学习出具能力的步行策略。另外,我们展示了在模拟中学到的政策如何转化为现实。我们还考察了该平台对灵活实验生态系统的支持程度。具体来说,我们观察到来自不同背景的新用户首次使用该平台取得成功的速度,以及当硬件出现问题时,平台修复和更新的速度。硬件设计和软件均可在GitHub上开源,便于自定义。总之,我们主张为经常使用模拟环境的强化学习研究者使用Open Ant,以便他们更容易地在评估中加入机器人实验。
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling
自动化数据工程与特征选择,用于熔融沉积建模中翘曲检测案例研究
- Authors: Saleh Valizadeh Sotubadi, Nazanin Mahjourian, Vinh Nguyen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18515
- Pdf link: https://arxiv.org/pdf/2607.18515
- Abstract
This study contributes toward development of an Automated Data Processing (ADP) framework designed to evaluate and reinforce optimal machine learning model-feature combinations for predictive tasks in fused deposition modeling (FDM) process datasets. The methodology is centered around a reinforcement learning-inspired policy updating mechanism, where multiple machine learning models are trained on both full feature sets and feature subsets selected through Shapley-based Explainable AI (SHAP XAI) across 217 datasets. At each episode, the framework assesses the predictive accuracy and F1-scores of each model-feature pair, computes a scalar reward, and updates $Q$ values to guide future model selection. SHAP XAI feature importance was employed to generate reduced yet informative feature subsets to enable the framework to explore performance with dimensionality. The policy was shown to evolve over multiple episodes, with reward distributions used to visualize performance stability. Overall, results indicate that leveraging the ADP framework through XAI algorithms successfully converges toward optimal model-feature configurations with improved accuracy and stability. Specifically, the proposed framework improves the test-set AUC from 0.9248 to 0.9731 and increases the mean reward value by more than fifty percent compared with the baseline full-feature configuration.
- 中文摘要
本研究有助于开发自动化数据处理(ADP)框架,旨在评估和强化融合沉积建模(FDM)过程数据集中预测任务的最佳机器学习模型-特征组合。该方法论围绕一种基于强化学习的策略更新机制展开,多个机器学习模型通过基于Shapley的可解释人工智能(SHAP XAI)在217个数据集中选择的完整特征集和特征子集进行训练。每期节目中,框架评估每对模型-特征的预测准确率和F1分数,计算标量奖励,并更新$Q$值以指导未来模型选择。采用SHAP XAI特征重要性来生成简化但信息丰富的特征子集,使框架能够以维度探索性能。该政策经过多次演变,奖励分布用于可视化表现稳定性。总体来看,结果表明,利用 XAI 算法利用 ADP 框架,能够以更高的准确性和稳定性,成功地趋向最优的模型-特征配置。具体来说,该框架将测试集AUC从0.9248提升至0.9731,并使平均奖励值相较基线全功能配置提高了50%以上。
Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces
连续状态-动作空间的网络多智能体强化学习可扩展策略优化
- Authors: Dongming Wang, Pengcheng Dai, Wenwu Yu, Wei Ren
- Subjects: Subjects:
Multiagent Systems (cs.MA); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.18554
- Pdf link: https://arxiv.org/pdf/2607.18554
- Abstract
We develop the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm for cooperative reinforcement learning in networked Markov decision processes with continuous state and action spaces. Each agent maintains a local actor over a bounded graph neighborhood, and a localized least-squares temporal-difference critic evaluates a truncated action-value function through a spectral random-feature representation of the local transition kernel. The analysis makes four contributions. First, the truncated action-value function is constructed as a conditional expectation over the neighborhood, yielding a well-posed localized Bellman theory that removes the continuation-kernel mismatch of naive truncation arguments. Second, we expose a dimensional obstruction to temporal-difference stability for normalized random features and prove an unconditional excitation bound that reduces stability to a symmetric persistence-of-excitation condition, monitorable through an online matrix-concentration certificate. Third, under exponential spatial decay of agent interactions, the excitation condition, and smoothness of the objective, CDCPG drives an averaged per-agent stationarity measure to within any excess $\epsilon$ of an explicitly characterized approximation floor using $\widetilde{\mathcal{O}}(\epsilon^{-2})$ shared-oracle samples, and the excess dependence matches the smooth nonconvex first-order rate; per-agent computation and communication are governed by the neighborhood size rather than the network size. Fourth, an adaptive-locality rule selects the radius that balances truncation and graph-decay residuals against the target accuracy. Experiments on a networked linear-quadratic benchmark corroborate the locality and feature-dimension predictions.
- 中文摘要
我们开发了连续分布耦合策略梯度(CDCPG)算法,用于网络化马尔可夫决策过程中具有连续状态和动作空间的协作强化学习。每个代理在有界图邻域上维护一个局部演员,局部最小二乘时间差分批评者通过局部转移核的谱随机特征表示来评估截断动作值函数。该分析提出了四个贡献。首先,截断作用值函数被构造为邻域上的条件期望,从而得到一个良态局部贝尔曼理论,消除了朴素截断论证的延核不匹配。其次,我们暴露了归一化随机特征时间差稳定性的维障碍,并证明了无条件激发束缚,将稳定性降为对称的激发持久状态,可通过在线矩阵浓度证书监测。第三,在代理相互作用的指数空间衰减、激发条件和目标物的平滑性条件下,CDCPG将平均的每位代理平稳度度量驱动到使用$\widetilde{\mathcal{O}}(\epsilon^{-2})$共享预言机样本明确刻画的近似底的任意多余$\epsilon}以内,且超额依赖性与平滑非凸一阶速率相匹配;每个代理的计算和通信由邻域大小而非网络大小控制。第四,自适应局部性规则选择半径,以平衡截断和图衰减残差与目标精度。网络化线性二次基准测试的实验证实了局域性和特征维度的预测。
Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States
在强化学习中以关系隐性状态为规划作为涌现行为
- Authors: Armin Sommer
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18589
- Pdf link: https://arxiv.org/pdf/2607.18589
- Abstract
Reinforcement learning is conventionally divided into model-based and model-free methods. In this taxonomy, model-based methods perform lookahead planning over a learned world model, whereas model-free methods learn a reactive state-action mapping. Recent work, however, has shown that planning can emerge from model-free reinforcement learning alone. The conditions under which this behavior emerges from a pure reward-maximization objective have so far remained unclear. In this paper, we present evidence that, in the observed cases, the hidden-state structure of the neural architecture is the deciding factor. We find that a network of relational hidden states, each anchored to an environment state and exchanging messages along learned relations, acquires a planning mechanism. These hidden states recover the environment's transition structure in their learned relations, and improve the policy at decision time by planning over the learned graph. In a matched control agent that must additionally discover which cells represent which states, no such binding arises, and no planning follows from it. We argue that this explains the observed phenomenon of emergent planning in model-free reinforcement learning and raises the question of how common such emergent planning might be more generally. Finally, we hypothesize that the discovered mechanism could describe how planning emerges from pure reward maximization in the human brain through a neural architectural prior.
- 中文摘要
强化学习通常分为基于模型的方法和无模型的方法。在该分类法中,基于模型的方法对已学习的世界模型进行前瞻性规划,而无模型方法则学习反应式状态-动作映射。然而,近期研究表明,规划可以仅靠无模型强化学习实现。这种行为源自纯粹奖励最大化目标的条件至今仍不明确。本文提出证据表明,在观察到的案例中,神经结构的隐性状态结构是决定性因素。我们发现,一个由关系隐形态组成的网络,每个状态锚定于环境状态,并沿学习关系交换消息,能够获得一种规划机制。这些隐藏状态在其学习关系中恢复了环境的过渡结构,并通过在学习的图上进行规划,在决策时改进策略。在匹配控制剂中,如果还必须确定哪些细胞代表哪些状态,不会产生这种结合,也不会因此进行任何规划。我们认为这解释了无模型强化学习中观察到的涌现规划现象,并引发了这种涌现规划在更广泛范围内可能有多普遍的问题。最后,我们假设发现的机制可以描述规划如何通过神经结构先验从人脑纯粹的奖励最大化中产生。
A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space
具有连续动作空间的协作任务的自我演化默认动作
- Authors: Shuangyao Huang
- Subjects: Subjects:
Machine Learning (cs.LG); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2607.18597
- Pdf link: https://arxiv.org/pdf/2607.18597
- Abstract
Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Carlo sampling often introduce bias into policy gradients and fail to guarantee convergence to local optima, as the sampled actions may not have been sufficiently trained. To address these limitations, we propose SAFE, a novel MARL framework that employs a counterfactual baseline conditioned on a self-evolving default action sampled from each agent's experience buffer. This design naturally extends to continuous action spaces without relying on additional simulations, reward models, or environment-specific prior knowledge. The baseline accurately quantifies each agent's contribution, and introduces no bias into the deterministic policy gradient, ensuring convergence to local optima. Extensive experiments on cooperative vehicular tasks demonstrate that SAFE consistently outperforms state-of-the-art models.
- 中文摘要
反事实功劳赋值在多智能体强化学习(MARL)中已被证明有效,适用于离散行动空间,但其扩展到连续动作协作任务仍具挑战性。现有通过蒙特卡洛抽样近似反事实基线的方法,常常在政策梯度中引入偏见,且无法保证收敛到局部最优,因为采样的动作可能未经过充分训练。为解决这些限制,我们提出了SAFE,这是一种新的MARL框架,采用一个反事实基线,条件是从每个代理的经验缓冲区中抽样的自我演化默认动作。这种设计自然延伸到连续动作空间,无需依赖额外的模拟、奖励模型或环境特定的先验知识。基线准确量化了每个代理的贡献,并未引入确定性策略梯度的偏见,确保趋向局部最优。对协作车辆任务的广泛实验表明,SAFE始终优于最先进模型。
Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach
ITNTN中的智能多无人机导航:一种层级大型语言模型方法
- Authors: Zijiang Yan, Hao Zhou, Wael Jaafar, Jianhua Pei, Ping Wang, Halim Yanikomeroglu, Hina Tabassum
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2607.18604
- Pdf link: https://arxiv.org/pdf/2607.18604
- Abstract
The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kinematics and multi-tier network handovers. While Deep Reinforcement Learning (DRL) offers rapid tactical control, it lacks the zero-shot strategic reasoning required to quickly adapt to dynamic Integrated Terrestrial and Non-Terrestrial Networks (ITNTNs). Conversely, Large Language Models (LLMs) excel at semantic reasoning but suffer from high inference latency, rendering them unsuitable for real-time aerodynamic control. To bridge this gap, we propose a novel Hierarchical LLM-driven control framework. A massive cloud-based LLM deployed on a High-Altitude Platform Station (HAPS) manages slow-timescale global load balancing, while lightweight edge-LLMs on individual UAVs translate local observations into tactical sub-goals. These sub-goals guide a fast-timescale physical DRL controller to execute collision-free, handover-aware trajectories. Simulation results demonstrate that our agentic architecture significantly reduces collision rates and improves aggregate system throughput compared to existing baselines.
- 中文摘要
高速无人机(UAV)在三维空中公路上的部署,需要对物理飞行运动学和多层网络切换进行强有力的协调。虽然深度强化学习(DRL)提供了快速的战术控制,但它缺乏快速适应动态综合地面与非地面网络(ITNTN)所需的零单次战略推理。相反,大型语言模型(LLMs)在语义推理方面表现出色,但由于推理延迟较高,不适合实时空气动力学控制。为弥合这一差距,我们提出了一种新的层级大型语言模型驱动控制框架。部署在高空平台站(HAPS)上的大型云大型语言模型(LLM)负责缓慢的全球负载均衡,而单架无人机上的轻量级边缘大型语言模型则将局部观测数据转化为战术子目标。这些子目标引导快速时间尺度的物理DRL控制器执行无碰撞、知交接的轨迹。模拟结果表明,我们的代理架构显著降低了碰撞率,并提升了整体系统吞吐量,相较于现有基线。
Exposure-Based Reinforcement Learning to Rank
基于暴露的强化学习以排名
- Authors: Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2607.18689
- Pdf link: https://arxiv.org/pdf/2607.18689
- Abstract
Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e.g., from precision or discounted cumulative gain to fairness-of-exposure or ranking distillation. However, standard RL is ineffective and computationally costly due to the enormous action space in LTR settings. Existing methods reach computational efficiency through custom gradient computation algorithms, but they are very complex to implement and often clash with auto-differentiation. Consequently, existing RL for LTR is not attractive to many practitioners. We reconsider RL for LTR while actively avoiding reliance on custom gradients. Contrary to the existing approaches, we focus on variance reduction and GPU computation. In doing so, we discover that high sample-efficiency can be reached through baseline corrections and partial marginalization. Furthermore, we propose an abstraction that places gradient estimation behind a document-exposure distribution, this enables seamless plug-and-play integration with auto-differentiation. Thereby, one only has to implement a loss as a differentiable function of exposure and RL for LTR can optimize it using auto-differentiation. Our experimental results reveal that our new exposure-based RL for LTR approach converges considerably faster and at significantly higher ranking performance than existing custom gradients, with no additional costs in computation time when using GPUs. In contrast, existing custom gradients result in severe stability issues when converging over many epochs, which never occur for our methods. Thus, we considerably improve RL for LTR methodology by increasing its effectiveness, efficiency, and ease of application.
- 中文摘要
强化学习(RL)方法用于学习排名(LTR)可以优化(几乎)任何排名目标,例如从精度或贴现累计增益到暴露公平性或排名提炼。然而,由于LTR环境中作用空间巨大,标准强化学习效率低下且计算成本较高。现有方法通过定制梯度计算算法实现计算效率,但实现起来非常复杂,且常与自微分发生冲突。因此,现有的长期性强化学习对许多从业者来说并不具吸引力。我们重新考虑长期关系的强化生活,同时积极避免依赖自定义梯度。与现有方法相反,我们专注于方差缩小和GPU计算。在此过程中,我们发现通过基线修正和部分边缘化可以实现高样本效率。此外,我们提出了一种抽象,将梯度估计置于文档暴露分布后面,实现了与自动微分的无缝即插即用集成。因此,只需将损失作为暴露的可微函数实现,长期反复学习(RL)可以通过自微分来优化损失。我们的实验结果显示,基于曝光的新型LTR强化学习收敛速度显著快,排名性能显著高于现有自定义梯度,且在使用GPU时计算时间不增加成本。相比之下,现有的自定义梯度在多个纪元收敛时会带来严重的稳定性问题,而我们的方法从未发生过这种情况。因此,我们通过提高LTR方法论的有效性、效率和应用简便性,显著提升了其性能。
Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents
策略跟随多智能体深度强化学习,考虑提供给其他智能体的控制策略
- Authors: Yamato Takahagi, Gentoku Nakasone, Yoshinari Motokawa, Toshiharu Sugawara
- Subjects: Subjects:
Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18719
- Pdf link: https://arxiv.org/pdf/2607.18719
- Abstract
This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after learning and enables uninstructed agents to implicitly complement the overall work based on the actions of other agents. Multi-agent applications using deep learning have shown potential; thus, to achieve extensive social applications, humans should be able to control learned agents using simple methods to respond to environmental and social changes. Even without such changes, learned coordination often does not match the expectations of human managers, making it preferable to control coordination structures to match human intentions. Some studies have aimed to control agent behavior using simple instructions. However, they assumed that instructions are provided to all agents, which is time-consuming and not evident when designing a better cooperation regime. Ideally, specific agents should receive key action instructions, while others should automatically complete the remaining tasks. The proposed method, which extends previous work on controllability in multi-agent deep reinforcement learning, enables uninstructed agents to adaptively complement overlooked tasks and areas. The experimental results show that agents using the proposed method can shift to another cooperative structure and achieve better performance than those using conventional methods.
- 中文摘要
本研究提出了一种多智能体系统的学习方法,使智能体在学习后通过人工管理指令进行控制,并使未受指示的智能体能够基于其他智能体的行为隐式补充整体工作。利用深度学习的多智能体应用已展现出潜力;因此,为了实现广泛的社会应用,人类应能够通过简单方法控制学习中的智能体,以应对环境和社会变化。即使没有这些变化,习得的协调往往也无法满足人类管理者的期望,因此更倾向于控制协调结构以符合人类意图。一些研究旨在通过简单的指令控制代理行为。然而,他们假设指令会向所有代理提供,这很耗时,在设计更好的合作机制时并不明显。理想情况下,特定客服应获得关键操作指令,而其他客服则应自动完成剩余任务。该方法扩展了此前关于多智能体深度强化学习可控性的研究,使未受指导的智能体能够自适应地补充被忽视的任务和领域。实验结果表明,采用该方法的代理可以转向另一种协同结构,并获得比传统方法更好的表现。
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
陈旧但稳定:稳定异步强化学习的陈旧自适应信任区域
- Authors: Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.18722
- Pdf link: https://arxiv.org/pdf/2607.18722
- Abstract
Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.
- 中文摘要
异步强化学习通过将推展生成与优化解耦来提升吞吐量,但由于策略滞后、引擎延迟和专家混合路由,陈旧是不可避免的副产品。从信任区域的角度来看,这种不匹配至关重要:训练-推断发散控制有限视界内的近似误差,而PPO裁断仅对采样的外部更新进行门控,作为采样代理而非完整策略约束。因此,在异步模式下,高陈旧更新的控制仍然很弱,而陈旧的发布最为重要。我们引入了陈旧-自适应信任区(SAT),该区域使用分离抽样对数比作为实用的陈旧代理,通过基于陈旧度的核尺度识别每批中高不匹配尾部,并仅收缩名义PPO区间的符号选择端点。这在新截获的外频带上保持了基准行为,同时对新截获的外向频段执行更保守的更新。我们证明了局部区间包含性和点状悲观性相对于PPO,展示了自适应规则如何在异构陈旧性下重塑更新几何。我们在基于Qwen3-30B-A3B-Base构建的解耦异步强化学习环境中评估SAT,使用SGLang作为推理引擎,并以Megatron进行训练。在此设定下,SAT-GSPO配备R3达到最佳观测值的AIME24 avg@8,延迟1时达到35.83,延迟8时达到34.79,而SAT-GSPO延迟1时达到34.17。自适应裁剪和路由重放分别作为互补稳定器,针对不匹配尾部和路由不一致性。总体而言,将剪辑间隔与陈旧异质性对齐,有效稳定异步强化学习。
From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning
从轨迹到指令:语言条件元强化学习
- Authors: Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18830
- Pdf link: https://arxiv.org/pdf/2607.18830
- Abstract
Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks. MAML training proceeds in two loops: an inner loop where the global parameters are adapted to task-specific parameters, and an outer loop where these task-specific parameters are evaluated and losses are back-propagated to improve the global parameters. Traditionally, the inner loop adaptation is performed by collecting trajectories from the task environment and applying gradient updates on the empirical expected return, which can be a costly operation. We note that it is the outer loop that drives the actual learning of global parameters, and therefore the inner loop adaptation mechanism need not be restricted to be gradient-based. This observation leads us to ask: Can we replace the inner loop trajectory collection and gradient update with a simpler, task-specific signal? In many practical settings, tasks are naturally accompanied by language instructions. Leveraging these instructions as a direct task-specific signal, we propose LA-MAML (Language Adapted MAML), which modifies the inner loop by adapting the global policy parameters in a single step through a learned embedding of the task instruction, replacing the inner loop trajectory collection and gradient-based updates. Experiments on the BabyAI benchmark demonstrate that LA-MAML achieves competitive or improved performance compared to baselines at a significantly lower per-iteration wall-clock training time. These results demonstrate that language instructions are an effective and efficient substitute for trajectory-based inner loop adaptation in meta RL.
- 中文摘要
模型无关元学习(MAML)是一种广泛使用的强化学习(RL)框架,通过学习能够快速适应新任务的全局策略参数实现高效的迁移。MAML训练分为两个循环:内环,全局参数适配任务特定参数;外环,评估这些任务特定参数并反向传播损失以改进全局参数。传统上,内环自适应通过收集任务环境中的轨迹并对经验期望回报进行梯度更新来实现,这可能是一项代价高昂的操作。我们注意到,实际学习全局参数的是外环,因此内环的适应机制无需限制为基于梯度。这一观察促使我们思考:我们能否用更简单、针对特定任务的信号替代内环轨迹收集和梯度更新?在许多实际环境中,任务自然会伴随着语言指令。利用这些指令作为直接的任务特定信号,我们提出了LA-MAML(语言适应MAML),通过学习嵌入任务指令,在单步内调整全局策略参数,取代了内环轨迹收集和基于梯度的更新。BabyAI基准测试的实验表明,LA-MAML在显著较低的每次迭代壁钟训练时间下,能够达到与基线相比具有竞争力或提升的性能。这些结果表明,语言指令是元强化学习中基于轨迹的内环适应的有效且高效的替代品。
Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments
动态环境中基于无人机的参与感测传递强化学习
- Authors: Xin Ouyang, Songxin Lei, Xusen Guo, Yutian Jiang, Sijie Ruan, Yuxuan Liang
- Subjects: Subjects:
Machine Learning (cs.LG); Computers and Society (cs.CY)
- Arxiv link: https://arxiv.org/abs/2607.18874
- Pdf link: https://arxiv.org/pdf/2607.18874
- Abstract
Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separately encoding distinct task features and sequentially evaluating UAV suitability before task selection. At the micro level, a wind-aware velocity controller learns fine-grained velocity scheduling to adapt to dynamic environmental variations. Extensive experiments on real-world datasets demonstrate that TSRL significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.
- 中文摘要
利用无人机(UAV)进行城市感知已成为一种强大的范式,可以通过灵活的空中众包监测城市状况,例如空气质量和噪音水平。尽管有这样的潜力,现有的无人机感测方法忽视了风等环境干扰,这些因素极大地影响了无人机的速度和能源效率。因此,直接将现有方法应用于动态环境中的联合交付与感知范式面临两个严峻挑战:(1)随着车队规模扩大,可扩展性瓶颈;以及(2)宏观任务调度与微观速度控制之间的多时间尺度决策异质性。为解决这些问题,我们将该问题形式化为SensUAV,并提出了一个双时间表强化学习框架(TSRL)。具体来说,TSRL将决策分为两个合作层面。在宏观层面,任务嵌入感测调度器通过分别编码不同的任务特征,并在任务选择前顺序评估无人机的适用性来处理可扩展性。在微观层面,风感测速度控制器学习细粒度的速度调度,以适应动态环境变化。对真实数据集的大量实验表明,TSRL显著优于基线,杭州实现了20.1%和上海46.6%的平均系统利润提升。
Circuit Claims Depend on What Is Extracted and How It Is Compared
电路权利要求取决于提取的材料及其比较方式
- Authors: Yang Sheng, Jie Fu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18921
- Pdf link: https://arxiv.org/pdf/2607.18921
- Abstract
Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior. We argue that this reading is under-determined: preserving behavior does not single out one circuit, because the claim it supports depends on which circuit is reported and how two circuits are compared. We make this concrete in a synthetic Lean tactic-prediction benchmark -- predicting the next step of a proof -- where fixed proof rules with randomized surface form let differences between extracted circuits be attributed to these choices rather than to the task. Across dense and weight-sparse checkpoints (most weights constrained to zero) of the same transformer, evaluated on atomic (single-rule) and compositional (multi-rule) proofs, we vary which extracted object is reported (a compact prediction-preserving circuit, a broader graph that also keeps surrounding read, write, and routing structure, or the smallest subgraph meeting a post-ablation loss threshold), and whether each attention head's query and key are represented jointly or separately. Exact component-to-component edge overlap is low and sensitive to these choices, at times dropping to a random baseline, while two coarser summaries stay stable: the set of selected attention heads, and the circuit-size ranking of conditions that differ in which supervised checkpoint initializes reinforcement learning (RL). The largest accuracy gains from RL on compositional proofs come with the most structure beyond the atomic circuits. A circuit-level claim is therefore well defined only once one states which circuit is reported, the pruning threshold used to extract it, and the level at which circuits are compared. We distill these requirements into a reporting practice for circuit-extraction studies.
- 中文摘要
电路提取识别出一小部分模型组件,这些元件的存在在消融下保持目标行为,所得电路通常被视为该行为背后的机制。我们认为这种解读是未充分确定的:保持行为并不只针对某一个电路,因为它支持的主张取决于报告的电路以及两个电路的比较方式。我们在合成精益战术预测基准测试中具体化这一点——预测证明的下一步——其中固定的证明规则和随机曲面形式允许提取电路之间的差异归因于这些选择,而非任务本身。在同一变换器中,在密度高且权重稀疏的检查点(大多数权重受限为零)上,基于原子(单规则)和复合(多规则)证明,我们会根据报告的对象(一个紧凑的预测保持电路、更宽的图(同时保留读写和路由结构,或最小的子图达到消融后丢失阈值),以及每个注意力中心的查询和键是联合表示还是单独表示。组件间的精确边缘重叠较低且对这些选择敏感,有时会降至随机基线,而两种较粗略的总结则保持稳定:所选注意力头的集合,以及监督检查点初始化强化学习(RL)不同条件的电路大小排序。强化逻辑在合成证明上最大的精度提升来自于原子电路之外的结构。因此,只有在说明报告的电路、提取该电路所用的修剪阈值以及比较电路的级别后,电路层级权利要求才被明确定义。我们将这些要求提炼为电路提取研究的报告实践。
H$^2$SD: Hybrid Hindsight Self-Distillation
H$^2$SD:混合后见之明自我蒸馏
- Authors: Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Tao Gui, Xiaocheng Feng, Bing Qin
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2607.18955
- Pdf link: https://arxiv.org/pdf/2607.18955
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse supervision and limited token-level credit assignment. On-policy distillation (OPD) provides denser supervision by distilling token-level distributions from a stronger teacher model, but requires an additional teacher and typically assumes a shared vocabulary. On-policy self-distillation (OPSD) removes this dependency by conditioning the same model on privileged information to construct a teacher policy. However, directly matching the teacher distribution may cause information leakage and unstable optimization. RLSD avoids direct matching by using the teacher signal only to modulate update magnitudes, but it cannot provide an explicit correction direction when the sampled reasoning fails. To address this tradeoff, we introduce $\mathrm{H}^{2}\mathrm{SD}$, a hybrid hindsight self distillation framework that uses the teacher differently according to trajectory correctness. For successful trajectories, the teacher receives the student response confirmed as correct together with a rephrasing instruction, and its probabilities on the original response tokens are used to modulate update magnitudes without changing the direction determined by the reward. For failed trajectories, we condition the teacher on a reference hint containing key reasoning steps and a verified answer, and minimize the reverse KL divergence from the student to the teacher. Experiments on multiple challenging reasoning benchmarks show that H$^2$SD consistently outperforms representative RLVR, OPSD, and RLSD baselines while maintaining stable optimization and favorable generation efficiency.
- 中文摘要
带有可验证奖励的强化学习(RLVR)显著提升了大型语言模型在数学推理和代码生成等任务中的推理能力。然而,大多数RLVR方法为整个轨迹分配标量结果奖励,导致监督稀疏且代币级的积分分配有限。策略提炼(OPD)通过从更强的教师模型中提炼代币级分布,提供更密集的监督,但需要额外的教师,且通常假设共享词汇。政策自蒸馏(OPSD)通过将同一模型置于特权信息上,构建教师策略,从而消除这种依赖。然而,直接匹配教师分布可能导致信息泄漏和优化不稳定。RLSD通过仅使用教师信号调制更新幅度来避免直接匹配,但当采样推理失败时,它无法提供明确的纠正方向。为了解决这一权衡,我们引入了$\mathrm{H}^{2}\mathrm{SD}$,一种混合式的事后诸葛亮自我提炼框架,根据轨迹正确性不同地使用教师。对于成功的轨迹,教师会收到确认正确的学生回答和重新表述的指令,并利用其在原始反应标记上的概率来调节更新幅度,而不改变奖励决定的方向。对于失败的轨迹,我们以包含关键推理步骤和验证答案的参考提示为条件,最小化学生与教师之间的反向KL偏差。对多个具有挑战性的推理基准测试的实验表明,H$^2$SD在保持稳定优化和有利生成效率的同时,始终优于代表性的RLVR、OPSD和RLSD基线。
Measuring Reward-Seeking via Contrastive Belief Updates
通过对比性信念更新测量追求奖励
- Authors: Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya, Felix Hofstätter, Jason Wolfe, Theodore Ehrenborg, Bronson Schoen, Alexander Meinke
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.18966
- Pdf link: https://arxiv.org/pdf/2607.18966
- Abstract
Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective. This "reward-seeking" is difficult to measure because a model that pursues the grader's judgment and one that pursues the intended objective behave identically whenever the grader rewards the intended behavior. We measure reward-seeking using Contrastive Synthetic Document Finetuning to change a model's beliefs about what the grader rewards, putting those beliefs in conflict with what users or developers want, and measuring the rate at which the model adopts each party's preferred behavior. Applied to intermediate checkpoints of a capabilities-focused OpenAI o3 RL run, without safety training, we find that these checkpoints often side with grader preferences over those of users or developers on coding and alignment tasks. This tendency to side with the grader trends upward throughout RL training. For example, in an environment that forces a choice between keeping a promise to a supervisor and breaking it to complete the task, a late capabilities-focused o3 checkpoint breaks the promise 87% of the time when SDF documents say the grader rewards task completion, versus 9% when they say it rewards honesty (a choice its chain-of-thought often makes explicit). An earlier checkpoint is far less sensitive (40% vs. 24%). Our method also generalizes to reward-hacking models. A model organism trained to reward-hack (gpt-oss-120b) is more than twice as sensitive to grader preferences as the unmodified model, with the mean behavioral shift in favor of the grader rising from 33% to 86%. These results indicate that RL can increase reward-seeking over the course of training, producing models that may act against their developers' intentions when they believe that doing so leads to higher reward.
- 中文摘要
经过强化学习训练的语言模型可能会学习优化评分者的判断,而非预期目标。这种“追求奖励”很难测量,因为一个追求评分者判断的模型和追求预期目标的模型,每当评分者奖励预期行为时,行为都是相同的。我们使用对比合成文档微调来衡量奖励寻求,以改变模型对评分者奖励对象的信念,使这些信念与用户或开发者的需求发生冲突,并测量模型采纳双方偏好行为的速度。应用到以能力为核心的OpenAI o3 RL运行的中间检查点,且未经过安全培训,我们发现这些检查点往往更偏向评分者偏好,而非用户或开发者在编码和对齐任务上的偏好。这种倾向于支持评分者的倾向在强化学习中逐渐加剧。例如,在一个强制在遵守对主管承诺和违背承诺以完成任务之间做出选择的环境中,当SDF文档称评分者奖励任务完成时,晚期能力导向的o3检查点有87%的概率违背承诺,而当他们说奖励诚实时(这一选择在其思维链中常被明确表达)则为9%。早期的检查点灵敏度要低得多(40%对24%)。我们的方法也推广到奖励黑客模型。经过训练以奖励黑客(gpt-oss-120b)的模型生物对评分者偏好的敏感度是未修改模型的两倍多,平均行为偏向评分者从33%升至86%。这些结果表明,强化学习在训练过程中可以增加追求奖励的行为,产生一些模型在开发者认为这样做能带来更高奖励时,可能会违背其意图。
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning
捞出搭便车者:基于Shapley的奖励归因,通过强化学习实现平行推理
- Authors: Wentao Zhang, Haoyu Zhang, Xinke Jiang, Yuxuan Cheng, Yuhan Pan, Miao Li, Zhipeng Qiao, Tao Feng, Zhen Tao, Dengji Zhao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18979
- Pdf link: https://arxiv.org/pdf/2607.18979
- Abstract
Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.
- 中文摘要
大型语言模型(LLMs)擅长多步推理,但当前的并行推理方法往往未能区分各个推理路径的贡献。许多路径可能冗余、误导甚至有害,但结果级奖励分配的是统一的奖励,导致学习信号模糊且训练不稳定。我们提出了平行Shapley,一种强化学习框架,赋予多路径推理中细粒度路径级贡献。我们将每条路径视为合作博弈中的玩家,利用夏普利值量化边际贡献,使用生成奖励模型评估路径效用和蒙特卡洛采样以实现高效近似。数学推理基准测试的实验表明,平行夏普利在提供更稳定和可解释的训练时,优于现有基线。我们的框架有效地“筛选搭便车者”,按比例分配奖励,并提升LLM中的多路径推理能力。
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio
雅典娜-大脑技术报告:一款用于通用智能和具身互动的高效机器人大脑
- Authors: Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo, Jiaming Song, Xutao Wang, Yi Zhao, Jiangpin Liu, Jie Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.18985
- Pdf link: https://arxiv.org/pdf/2607.18985
- Abstract
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-device brain, preserving the broad general intelligence of LLMs while enabling effective high-level interaction with embodied environments. Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model. We present \textbf{Athena-Brain-8B}, an 8B LLM designed to serve as an on-device brain for embodied intelligence for embodied intelligence. Through a multi-stage post-training pipeline consisting of General Supervised Fine-Tuning, General Reinforcement Learning, Embodied Expert training, and Model Merge, Athena-Brain-8B maintains strong general capabilities while acquiring strong high-level embodied interaction capabilities and generating concise responses for efficient embodied interaction. Experimental results demonstrate the effectiveness of Athena across both general and embodied evaluations. Compared with the corresponding Qwen3-8B thinking model, Athena-Brain-8B achieves comparable performance on general language and reasoning benchmarks while generating substantially shorter responses. On in-domain embodied benchmarks, Athena-Brain-8B consistently outperforms models of similar scale and surpasses several substantially larger frontier models evaluated zero-shot, demonstrating that compact language models can effectively integrate strong general intelligence with embodied capabilities.
- 中文摘要
大型语言模型(LLMs)在语言理解、推理和世界知识方面展现出了卓越的能力。随着具身代理能力的提升,对能够作为设备内大脑的紧凑型模型的需求日益增长,既能保留LLM的广泛通用智能,又能与具身环境实现高效高层交互。然而,现有方法往往优先考虑通用智能或专门的具象能力,这使得在单一模型中满足这两种需求变得具有挑战性。我们介绍 \textbf{Athena-Brain-8B},一款 8B 大型语言模型,设计用于具身智能的设备内大脑。通过多阶段的培训后流程,包括通用监督微调、通用强化学习、具身专家培训和模型合并,Athena-Brain-8B 保持了强大的通用能力,同时获得强大的高层次具身互动能力,并生成简洁的响应以实现高效的具身互动。实验结果显示雅典娜在一般和具体评估中都表现出显著效果。与相应的Qwen3-8B思维模型相比,Athena-Brain-8B在一般语言和推理基准测试上表现相当,但生成的响应时间明显更短。在领域内具象基准测试中,Athena-Brain-8B持续优于同等规模模型,并超过多个评估过的大型前沿模型,证明紧凑语言模型能够有效整合强的通用智能与具身能力。
DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
DobicVLM:通过群体相对政策优化,使胸部X光报告生成与临床基础的项目化奖励保持一致
- Authors: Thanni Adewuyi, Angelica Obayi, Andem Aniekan, Samuel Okoko, Angel Ezendu, Ephraim Usani, Ademide Animasaun, Philip Chibundu, Christian Maurice, Mary Donald Essien, Oluwaseun Odunsi, Oluwasegun Oguntuase, Abiodun Adereni
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.18988
- Pdf link: https://arxiv.org/pdf/2607.18988
- Abstract
Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completeness, and semantic faithfulness. We introduce DobicVLM, a vision-language model combining supervised fine-tuning on MedGemma-4B with Group Relative Policy Optimization (GRPO) and clinically-grounded programmatic rewards. Our approach uses interpretable, rule-based reward components; structural verification, anatomical checklist, semantic similarity, and length constraints to enforce clinical standards without neural reward models. Trained on 1,000 de-identified image-report pairs from a private clinical dataset (with ethics approval and compliance to local regulations), DobicVLM is evaluated via blinded expert review on 69 held-out cases. DobicVLM outperforms Gemini 2.5 Flash across the majority of criteria, achieving the highest impression accuracy (27.2%) and medical terminology (86.5%) compared to both Gemini 2.5 Flash and MedGemma 4B baselines, with minor trade-offs in completeness and referrals. This demonstrates GRPO's value for transparent alignment in resource-limited settings. Keywords: Vision-Language Models, Radiology Report Generation, Reinforcement Learning, Medical AI, GRPO
- 中文摘要
医学影像是诊断的基石,但自动化的胸部X光报告生成在结构一致性、解剖完整性和语义忠实性方面存在困难。我们介绍了DobicVLM,一种结合了MedGemma-4B监督微调、群体相对政策优化(GRPO)及临床基础的项目奖励的视觉语言模型。我们的方法采用可解释的、基于规则的奖励组件;结构验证、解剖检查表、语义相似性和长度限制,以在不依赖神经奖励模型的情况下执行临床标准。DobicVLM基于来自私人临床数据集的1000对去标识图像报告对进行训练(已获得伦理批准并符合当地法规),通过盲法专家评审对69个未完成案例进行评估。DobicVLM在大多数标准上优于Gemini 2.5 Flash,在展示准确率(27.2%)和医学术语(86.5%)上均优于Gemini 2.5 Flash和MedGemma 4B基线,但在完整性和转诊方面存在轻微权衡。这体现了GRPO在资源有限环境中透明对齐的价值。关键词:视觉语言模型、放射报告生成、强化学习、医疗人工智能、GRPO
Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation
采用具有可验证奖励的强化学习以促进分子生成
- Authors: Mingxuan Ouyang, Hao Lan, Wanyu Lin
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.19044
- Pdf link: https://arxiv.org/pdf/2607.19044
- Abstract
Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design. Current methods primarily rely on supervised training or fine-tuning with limited datasets, which are insufficient to capture complex molecular design objectives. While some approaches attempt to guide generation toward specific goals, they often lack direct optimization mechanisms, making it difficult to align generated molecules with desired properties. To tackle these challenges, we propose \textbf{LLMol}, a principled reinforcement learning framework that directly incorporates verifiable rewards for targeted molecule generation. The key insight is to formulate molecular design as a goal-conditioned sequence prediction task, where verifiable rewards serve as explicit supervision to drive generation toward desired objectives. LLMol follows a two-stage training paradigm combining supervised learning and reinforcement learning. In the first stage, large language models are supervised fine-tuned to capture chemical syntax and molecular distributions. In the second stage, we introduce Reinforcement Learning with Verifiable Rewards (RLVR), which directly integrates property-based reward signals to guide molecular generation toward task-specific objectives. To address the high variance and instability common in discrete sequence optimization, we adopt Group Relative Policy Optimization (GRPO), a stable on-policy algorithm that smooths reward signals and improves training robustness. This framework enables LLMol to effectively handle a range of molecular design tasks, including single-property targeting (e.g., penalized logP, QED) and structure-constrained optimization. Experimental results demonstrate that LLMol consistently outperforms existing methods, achieving higher success rates and improved efficiency across diverse molecular benchmarks.
- 中文摘要
利用大型语言模型(LLMs)进行分子生成,在化学和药物设计中展现出显著潜力。目前的方法主要依赖监督训练或有限数据集的微调,这些数据不足以实现复杂的分子设计目标。虽然一些方法试图引导生成朝向特定目标,但往往缺乏直接的优化机制,使得生成的分子难以与预期性质对齐。为应对这些挑战,我们提出了 \textbf{LLMol},一种原则性的强化学习框架,直接包含可验证的目标分子生成奖励。关键见解是将分子设计表述为一个目标条件序列预测任务,其中可验证的奖励作为明确监督,推动生成朝着期望目标发展。LLMol采用两阶段训练范式,结合监督学习和强化学习。第一阶段,大型语言模型被监督并微调以捕捉化学句法和分子分布。第二阶段引入了可验证奖励强化学习(RLVR),直接整合基于属性的奖励信号,引导分子生成实现任务特定目标。为解决离散序列优化中常见的高方差和不稳定性,我们采用了群相对策略优化(Group Relative Policy Optimization,GRPO),这是一种稳定的策略上算法,能够平滑奖励信号并提升训练鲁棒性。该框架使LLMol能够有效处理多种分子设计任务,包括单一属性的定向(如惩罚logP、QED)和结构约束优化。实验结果表明,LLMol 持续优于现有方法,在多种分子基准测试中实现更高的成功率和更高的效率。
Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning
参数化动作强化学习中多智能体演员-批评算法的比较研究
- Authors: Ubayd Ali Bapoo, Clement N Nyirenda
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.19117
- Pdf link: https://arxiv.org/pdf/2607.19117
- Abstract
Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization. Prior work established the effectiveness of single-agent actor-critic algorithms - Greedy Actor-Critic (GAC), Soft Actor-Critic (SAC), and Truncated Quantile Critics (TQC) - on benchmark parameterized action tasks, but their extension to multi-agent settings remains largely unexplored. This paper presents a comparative study of shared-experience multi-agent extensions of these algorithms: Multi-Agent Greedy Actor-Critic (MAGAC), Multi-Agent Soft Actor-Critic (MASAC), and Multi-Agent Truncated Quantile Critics (MATQC). Rather than following the centralized training, decentralized execution (CTDE) paradigm, the proposed framework uses multiple independent actor-critic agents that share a replay buffer while maintaining separate policy and value networks. We evaluate the algorithms on the Platform-v0 and Goal-v0 benchmarks against their single-agent counterparts, using three-, five-, and ten-agent configurations to assess scalability. Performance is measured by average evaluation return and training time across ten independent runs, with one-way ANOVA and Tukey HSD post-hoc tests used to assess statistical significance. Results show that the multi-agent framework consistently improves Greedy Actor-Critic performance, while MASAC and MATQC show comparatively modest gains over their single-agent versions. Increasing the number of agents beyond five yields limited additional performance while substantially raising computational cost, particularly for MAGAC. These results highlight a trade-off between learning performance and computational efficiency, offering insight into the scalability of shared-experience multi-agent actor-critic methods for parameterized action reinforcement learning.
- 中文摘要
参数化动作强化学习在需要离散动作选择和连续参数化的环境中表现出优异的性能。此前已有研究证明单代理演员-批评算法——贪婪演员-批评者(GAC)、软演员-批评者(SAC)和截肢分位批评者(TQC)——在基准参数化动作任务中的有效性,但其在多代理环境中的推广仍然鲜有深入探讨。本文对这些算法的共享经验多代理扩展进行了比较研究:多智能体贪婪演员-批评者(MAGAC)、多智能体软性演员-批评者(MASAC)和多智能体截断分位批评者(MATQC)。该框架不遵循集中式训练、去中心化执行(CTDE)范式,而是使用多个独立的actor-critic代理共享重放缓冲区,同时维护独立的策略和价值网络。我们通过三代理、五代理和十代理配置,将Platform-v0和Goal-v0基准测试的算法与单代理对应对应指标进行评估。表现通过十次独立运行的平均评估回报和训练时间衡量,采用单向方差分析和Tukey HSD事后检验评估统计显著性。结果显示,多智能体框架持续提升贪婪的演员-批评者表现,而MASAC和MATQC相较于单智能体版本的提升则相对有限。将代理数量增加到超过五个,会限制额外的性能,同时显著增加计算成本,尤其是对MAGAC而言。这些结果凸显了学习性能与计算效率之间的权衡,为共享经验多代理演员-批评方法在参数化动作强化学习中的可扩展性提供了见解。
Coherence in Control: Bridging Many-Core Mapping and Routing through Cost Unification
控制中的一致性:通过成本统一桥接多核映射与路由
- Authors: Guochu Xiong, Xiangzhong Luo, Weichen Liu
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2607.19158
- Pdf link: https://arxiv.org/pdf/2607.19158
- Abstract
The rapid growth of data-intensive applications increases communication demands in many-core systems, where cache coherence, while essential for correct communication and data consistency, introduces substantial overhead due to frequent data sharing and coherence activities. As system scale and workload complexity grow, the resulting coherence traffic intensifies communication pressure, making the co-optimization of task mapping and routing essential for improving system performance. However, most existing approaches overlook cache coherence, leaving a substantial portion of coherence-induced communication unaccounted for and creating a mismatch between optimization objectives and actual communication patterns. Furthermore, by employing separate cost evaluators for mapping and routing, these approaches complicate objective coordination, may lead to conflicting decisions, and fail to capture the coherence-induced coupling between the two stages. To address these challenges, we propose CoCo, a coherence-aware co-optimization framework that jointly integrates task mapping and routing under a unified cost model for realistic scenarios. This unified model integrates communication cost, coherence overhead, and load imbalance into a single objective, enabling coherence-aware decision-making and effective trade-offs among optimization goals. Guided by this model, CoCo combines coherence-guided task mapping with reinforcement learning-based routing, where directional link weights are adjusted according to communication behavior to improve traffic distribution, enabling coherence-aware co-optimization for many-core systems. Experimental results show that CoCo reduces link utilization by 88.46%, packet delay by 17.40%, and execution time by 17.58% compared with existing approaches, highlighting the importance of cache coherence in co-optimization design.
- 中文摘要
数据密集型应用的快速增长增加了多核系统的通信需求,缓存一致性虽然对正确通信和数据一致性至关重要,但由于频繁的数据共享和一致性活动,缓存一致性会带来大量开销。随着系统规模和工作负载复杂度的增加,产生的相干流量加剧了通信压力,使得任务映射和路由的协同优化对于提升系统性能至关重要。然而,大多数现有方法忽视了缓存一致性,导致大量由一致性诱导的通信未被考虑,导致优化目标与实际通信模式之间存在不匹配。此外,通过使用不同的成本评估器进行映射和路由,这些方法使客观协调变得复杂,可能导致决策冲突,并未能捕捉两阶段之间由相干性诱导的耦合。为应对这些挑战,我们提出了CoCo,这是一个具相干性意识的协同优化框架,将任务映射与路由整合于统一的成本模型下,以实现现实情景。该统一模型将通信成本、一致性开销和负载不平衡整合为单一目标,实现一致性感知决策和优化目标间的有效权衡。在该模型的指导下,CoCo 结合了相干引导任务映射与基于强化学习的路由,后者根据通信行为调整方向链路权重,以改善流量分布,从而实现多核系统的相干感知协同优化。实验结果显示,与现有方法相比,CoCo 将链路利用率降低了 88.46%,数据包延迟降低了 17.40%,执行时间减少了 17.58%,凸显了缓存一致性在协同优化设计中的重要性。
Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning
翻译前的推理:用结构化推理增强法律机器翻译
- Authors: Aixiu An, Michael Jungo, Eloi Eynard, Mark Drenhaus, Andreas Fischer, Jean Hennebert, Sébastien Rumley
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.19181
- Pdf link: https://arxiv.org/pdf/2607.19181
- Abstract
Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires. The recent emergence of reasoning-capable language models opens new possibilities for tackling such challenges. They add to a set of other previously proposed techniques to enhance the translation quality, which includes supervised fine-tuning and reinforcement learning. In this work, we perform a comparison between these various approaches. More particularly, we evaluate small language models such as Qwen3.5 4B, Qwen3.5 9B, and Gemma 3 12B enhanced with various re-training paradigms and compare their performances against frontier reasoning models. We focus on the Swiss legal system, which -- with its unique multilingual statutes -- offers a particularly challenging testbed for reasoning-augmented models. Our results show that the quality of small ``base'' models can be greatly enhanced, and that reinforcement learning with verifiable rewards can be applied to NMT in the legal domain and surpasses the translation quality of supervised fine-tuning. The performance of enhanced small models is close to the one of state-of-the-art reasoning models yet remains inferior. We also note that re-training paradigms yield diminishing returns as model size increase. The code and models are publicly available at this https URL.
- 中文摘要
法律领域的神经机器翻译(NMT)是一项语言和概念上极具挑战性的任务,主要由于法律语言的复杂性和高度的精确度要求。最近出现的具备推理能力的语言模型为应对此类挑战打开了新可能。它们补充了之前提出的其他技术,以提升翻译质量,包括监督微调和强化学习。在本研究中,我们对这些不同方法进行了比较。更具体地,我们评估了通过多种重训练范式增强的小型语言模型,如Qwen3.5 4B、Qwen3.5 9B和Gemma 3 12B,并比较它们与前沿推理模型的表现。我们重点关注瑞士法律体系,其独特的多语言法规为推理增强模型提供了极具挑战性的试验场。我们的结果表明,小型“基础”模型的质量可以被大幅提升,且带有可验证奖励的强化学习可以应用于法律领域的NMT,并且超越监督微调的翻译质量。增强型小模型的性能接近最先进的推理模型,但仍较为逊色。我们还注意到,随着模型规模的增加,重新训练范式会产生收益递减。代码和模型在此HTTPS网址公开。
Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation
在不确定性估计下离线强化学习的保守查询与自适应正则化
- Authors: Li-Rong Zhou, Qin-Wen Luo, Sheng-Jun Huang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.19199
- Pdf link: https://arxiv.org/pdf/2607.19199
- Abstract
Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset, but its performance is fundamentally limited by dataset coverage. Action preference queries leverage expert feedback without additional environment interaction, enabling policy improvement during offline training. However, existing methods still face two key challenges: selecting informative preference queries and effectively exploiting the collected feedback. Current approaches typically rely only on the distance between policy actions and dataset actions for query selection, while enforcing fixed constraints that keep the policy close to queried preferences. Such strategies often lead to unstable policy updates and integrate poorly with value regularization. To address these limitations, we propose Conservative Query and Adaptive Regularization under Uncertainty Estimation, a lightweight framework that jointly improves preference querying and preference exploitation. Specifically, we employ a Morse network to estimate the uncertainty of policy actions with respect to the offline dataset. Based on this uncertainty, we introduce a conservative query strategy that selectively queries actions near the dataset to preserve Bellman-update stability, together with an uncertainty-aware adaptive regularization scheme that dynamically adjusts data-level constraints during policy optimization. We integrate our framework with CQL and evaluate it extensively on the D4RL benchmark. Experimental results demonstrate superior or competitive performance across a wide range of tasks.
- 中文摘要
离线强化学习(RL)旨在从静态数据集中学习有效的策略,但其性能在根本上受限于数据集覆盖范围。行动偏好查询利用专家反馈,无需额外环境互动,从而在离线培训期间实现政策改进。然而,现有方法仍面临两个关键挑战:选择有信息量的偏好查询和有效利用收集到的反馈。当前方法通常仅依赖策略动作与数据集动作之间的距离来进行查询选择,同时强制执行固定约束,使策略接近查询偏好。此类策略常导致政策更新不稳定,且与价值正则化整合不佳。为解决这些局限性,我们提出了保守查询和不确定性估计下的自适应正则化,这是一个轻量级框架,共同提升偏好查询和偏好利用。具体来说,我们采用莫尔斯网络来估计离线数据集中政策行动的不确定性。基于这种不确定性,我们引入了一种保守的查询策略,选择性查询数据集附近的动作以保持Bellman更新稳定性,同时采用一种不确定性意识的自适应正则化方案,在策略优化过程中动态调整数据级约束。我们将框架与CQL集成,并在D4RL基准测试上进行了广泛评估。实验结果显示,在广泛的任务中表现出优异或具有竞争力的性能。
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
超越评分预测:基于LLM的论文评分与通过强化学习与评分标准奖励的反馈生成
- Authors: Xuefeng Jin, Jiashuo Zhang, Teng Cao, Bin Yang
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.19219
- Pdf link: https://arxiv.org/pdf/2607.19219
- Abstract
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited. We propose RLAES, a unified LLM framework that jointly optimizes essay scoring and feedback generation through RL. To make feedback quality measurable, interpretable, and usable for training, we introduce Rubric-based Feedback Evaluation (RFE), an essay-grounded feedback evaluation framework comprising 166 fine-grained binary rubric items and an LLM-as-judge. Building on RFE, we propose Adaptive Gated Feedback Optimization (AGFO), which activates rubric-based feedback rewards on demand during RL, reducing evaluation overhead while improving feedback quality. We also propose Adjacent Contrastive Reasoning (ACR) to improve ordinal score calibration by explicitly contrasting adjacent score levels. Experimental results show that the RFE framework captures essay-feedback consistency, exhibits strong pairwise discriminative power, and closely aligns with expert preferences. On the ASAP benchmark, RLAES-AGFO achieves the best scoring performance among LLM-based methods (QWK = 0.803), while maintaining feedback quality comparable to GPT-5.5 and avoiding the feedback degradation observed under score-only RL. Code and datasets are publicly available at this https URL.
- 中文摘要
大型语言模型(LLMs)已被广泛应用于自动论文评分(AES)和自动反馈生成(AFG)。然而,现有研究主要依赖即时工程或监督微调,而关于强化学习(RL)训练后和反馈质量自动评估的系统性研究仍然有限。我们提出了RLAES,一个统一的LLM框架,通过强化学习共同优化论文评分和反馈生成。为了使反馈质量可衡量、可解释且适合培训,我们引入了基于评分标准的反馈评估(RFE),这是一个基于论文的反馈评估框架,包含166个细致的二元评分标准条目和一名LLM评审。基于RFE,我们提出了自适应门控反馈优化(AGFO),在强化学习期间按需激活基于评分标准的反馈奖励,降低评估开销同时提升反馈质量。我们还提出了相邻对比推理(ACR)方法,通过显式对比相邻分数水平来改善序数评分校准。实验结果显示,RFE框架能够捕捉论文反馈的一致性,展现出强大的两两辨别能力,并且与专家偏好高度一致。在ASAP基准测试中,RLAES-AGFO在基于LLM的方法中获得最佳评分(QWK = 0.803),同时保持与GPT-5.5相当的反馈质量,避免了仅评分强化学习下观察到的反馈退化。代码和数据集可在此 https URL 公开获取。
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
推理的代价:神经机器翻译强化学习中的成本与质量权衡
- Authors: Michael Jungo, Aixiu An
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.19226
- Pdf link: https://arxiv.org/pdf/2607.19226
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest research indicating that RLVR could be the preferred training method for translating legal documents due to the induced reasoning capabilities, it raises the question whether it is really attributed to the reasoning or more generally to the training paradigm. We investigate the importance of including the model's reasoning trace in the generated responses during both training and inference by systematically omitting it from one of the phases. Our experiments show that including the reasoning, specifically during inference, has a positive effect on the overall translation quality. Furthermore, we recognise that the reasoning leads to an increase in output tokens, hence we study the cost-quality tradeoff between the increased computational demands and the improved translation quality.
- 中文摘要
带可验证奖励的强化学习(RLVR)已被确立为大型语言模型(LLMs)后期训练的可行范式,包括下游任务,如神经机器翻译(NMT)。最新研究表明,由于诱导推理能力,RLVR可能成为翻译法律文件的首选训练方法,这也引发了一个问题:这究竟是归因于推理本身,还是更广泛地归因于训练范式。我们通过系统地在其中一个阶段中省略模型的推理痕迹,探讨在生成的响应中包含模型推理痕迹的重要性。我们的实验表明,在推理过程中包含推理,对整体翻译质量有积极影响。此外,我们认识到这种推理导致输出代币数量的增加,因此我们研究了计算需求增加与翻译质量提升之间的成本与质量权衡。
S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning
S3:通过限制粗动态学的不确定性,在层级强化学习中实现稳定子目标选择
- Authors: Kshitij Kumar Srivastava, Kshitij Jerath
- Subjects: Subjects:
Machine Learning (cs.LG); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2607.19232
- Pdf link: https://arxiv.org/pdf/2607.19232
- Abstract
Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution. It has been widely successful in solving long-horizon and complex tasks, where flat-RL algorithms have difficulty in learning. However, while the low-level agent in HRL benefits from dense feedback and abundant trial opportunities, the high-level agent receives sparse, delayed feedback from the environment and its performance depends on the low-level execution capability. In this paper, we study whether subgoal selection by the high-level agent can be performed more strategically, by providing it with dynamics-aware intrinsic motivation. Since motivation based on primitive transition dynamics would require broad coverage of the state-action space, we propose to use coarse dynamics, i.e., environment transitions aggregated over multiple steps at the temporal scale at which the high-level agent operates. This approach stabilizes the high-level policy by learning to minimize the predictive uncertainty associated with the coarse dynamics, and provides a guided structure for navigation. We model the predictive uncertainty by evaluating different dispersion metrics as approximated by a Mixture Density Network (MDN). Empirically, we observe that a dense, dynamics-aware intrinsic reward leads to risk-averse subgoal selection, enabling it to outperform state-of-the-art HRL methods in non-stationary long-horizon environments.
- 中文摘要
分层强化学习(HRL)旨在将战略规划与原始执行区分开来。它在解决长期且复杂的任务方面取得了广泛成功,而这些任务在平面强化学习中较为困难。然而,虽然HRL中的低级代理受益于密集的反馈和丰富的试验机会,而高级代理则从环境中接收到稀疏且延迟的反馈,其性能依赖于低级别执行能力。本文探讨了高级代理是否可以通过赋予其动态感知的内在动机,更策略性地执行子目标选择。由于基于原始转移动力学的动机需要对状态-动作空间的广泛覆盖,我们建议使用粗态动力学,即在高层次代理工作的时间尺度上,环境转变分多个步骤聚合。这种方法通过学习最小化与粗动态相关的预测不确定性,稳定了高层策略,并为导航提供了指导结构。我们通过评估不同的扩散度指标,并以混合密度网络(MDN)近似来建模预测不确定性。实证上,我们观察到密集且动态感知的内在回报导致风险厌恶的子目标选择,使其在非固定的长视野环境中优于最先进的HRL方法。
A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors
一种增强学习增强型液体燃料反应堆网络模型,用于预测燃气轮机的稀薄爆出
- Authors: Philip John, Eloghosa Ikponmwoba, Pinaki Pal, Opeoluwa Owoyele
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2607.19281
- Pdf link: https://arxiv.org/pdf/2607.19281
- Abstract
This study introduces a reinforcement learning (RL) framework for generating optimal liquid-fueled reactors to improve lean blowout (LBO) predictions in gas turbine combustors. Existing approaches for determining cluster boundaries rely on manual heuristics or distance-based metrics in the input space. In contrast, the proposed method is goal-oriented, explicitly accounting for the target metric (e.g., LBO prediction accuracy) during cluster formation. The framework employs a multi-stage clustering--classification strategy: an initial clustering step (e.g., $k$-means clustering) generates a large set of homogeneous micro-clusters, followed by an actor-critic RL agent that merges them into optimal reactor zones. The validation study, performed using a Jet-A mechanism (119 species, 841 reactions), shows the RL framework offers improved predictive fidelity compared to $k$-means and captures the correct LBO trends, while achieving substantial speedups relative to the high-fidelity computational model. Overall, the RL-driven approach demonstrates strong potential as a computationally efficient reduced-order modeling technique that can complement high-fidelity simulations for rapid design-space exploration.
- 中文摘要
本研究引入了一种强化学习(RL)框架,用于生成最优液体燃料反应堆,以提升燃气轮机燃烧器中的脱脂爆出(LBO)预测。现有确定簇边界的方法依赖于输入空间中的手动启发式或基于距离的度量。相比之下,所提方法是目标导向的,明确考虑了集群形成过程中的目标指标(例如LBO预测准确率)。该框架采用多阶段聚类分类策略:初始聚类步骤(例如$k$-mean聚类)生成大量均质微聚类,随后由actor-critic RL代理合并为最优反应堆区。该验证研究采用Jet-A机制(119个物种,841个反应)进行,显示RL框架相比$k$均值提供了更好的预测精度,并捕捉了正确的LBO趋势,同时相较于高保真度计算模型实现了显著的加速。总体而言,强化学习驱动的方法展现出作为一种计算高效、低阶建模技术的强大潜力,可以补充高精度仿真,实现快速设计空间探索。
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
脱离上下文的GRPO:利用特权信息学习推理困难问题
- Authors: Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.19313
- Pdf link: https://arxiv.org/pdf/2607.19313
- Abstract
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by the original prompt without that guidance.} {We introduce} Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training. Empirically, our algorithm achieves a 3.9\% absolute improvement (13.8\% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks with negligible additional cost.
- 中文摘要
带有可验证奖励的强化学习(RLVR)提升了大型语言模型中的推理能力。然而,典型的RLVR方法在困难问题上失败:当模型无法生成任何正确解时,它接收到\textit{zero}学习信号。在培训中提供特权指导,如解决方案前缀,可以帮助克服这一学习悬崖,引导模型朝向{非零奖励的正确解}方向。{我们称这些推出为 \textit{off-context}:它们由包含特权指导的训练提示生成,而目标目标由原始提示定义,没有该指导。}{我们介绍}脱离上下文GRPO(OC-GRPO)是GRPO的最小修改变体,采用引导式推送,但应用重要性修正目标,将更新引导回原始无引导目标,避免导致未校正引导训练不稳定的不匹配。从经验来看,我们的算法在标准数学推理基准测试中平均比普通GRPO实现了3.9%的绝对提升(相对增益13.8%),且额外成本极小。
ISO: An RLVR-Native Optimization Stack
ISO:RLVR 原生优化栈
- Authors: Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2607.19331
- Pdf link: https://arxiv.org/pdf/2607.19331
- Abstract
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.
- 中文摘要
带有可验证奖励的强化学习(RLVR)正在快速提升语言模型的推理能力,但将奖励反馈转化为权重空间更新的优化层仍被理解不足。基于我们之前的分析(Zhu 等,2025),我们通过模型权重的奇异结构研究了这一缺失层,并识别了谱继承:RLVR可以重用基础模型的权重谱,同时通过对应的输入和输出奇异帧的变化获得新的行为。我们将谱继承操作化为等谱优化(ISO),这是一个RLVR原生的固定频谱优化框架,具有互补的离线和在线实例。离线时,ISO-Merger将共享基专家的帧变更整合为单一固定频谱模型,无需合并后数据、推广、梯度更新或策略内提炼(OPD)。它恢复了互补的专业能力,并在无数据合并方法中实现了最强的综合性能。在线时,ISO-Optimizer 将选定的基优化器(包括 AdamW 和 Muon)应用于帧变量,同时保持基谱固定。在参数从1.5B到8B的推理和编码任务中,ISO-Optimizer提升了报告运行的准确性,并以显著减少的训练步骤达到匹配分数。在Qwen3-8B基地,AdamW经过270个训练步骤后,总准确率达到0.495。ISO-AdamW在仅进行100次训练后即可达到相同精度,并在210次训练后进一步提升至0.509。综合而言,ISO为RLVR缺失的优化层提供了具体答案:与其全面继承预训练优化,不如围绕奖励驱动适应结构设计后训练:继承谱系,优化帧。
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
OmniReasoner:通过原生工具使用长音视频思考
- Authors: Yu Chen, Caorui Li, Ziyu Xiong, Yidong Wang, Mingqi Gao, Shuman Liu, Biao Liu, Chunfeng Yang, Anxiang Zeng, Haibo Zhang, Chaofan Chen
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2607.19339
- Pdf link: https://arxiv.org/pdf/2607.19339
- Abstract
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervised fine-tuning and reinforcement learning, to decide whether and where to call a zoom-in tool before answering. OmniReasoner first builds a low-cost global preview of the full stream and then, when needed, calls the zoom-in tool with a requested temporal interval for higher-fidelity visual and audio inspection before answering. Because the model observes different sampling granularities before and after this call -- a sparse global preview and a denser local clip -- we introduce TimeAnchor, which keeps the tool's temporal argument valid and round-trip-consistent across these granularities, rather than tied to frame indices from a particular sampling rate. To make this tool-use behavior trainable without expensive manual interval annotation, we build a Temporal Augmented Data Engine that synthesizes tool-use post-training trajectories by video editing and composition. Experiments across omnimodal and video benchmarks show that OmniReasoner improves both answer accuracy and temporal grounding while concentrating high-fidelity computation on informative regions. Code is available at this https URL.
- 中文摘要
对于全模态大型语言模型来说,长时间的音视频推理困难,因为决定性证据往往稀疏、跨模态且在统一高保真输入下保存成本过高。我们介绍OmniReasoner,一个用于长音视频思考的后期培训工具使用框架:全模态大型语言模型通过监督微调和强化学习,决定是否以及在哪里调用放大工具,然后再回答。OmniReasoner 首先构建一个低成本的全局预览,然后在需要时调用放大工具,提供所需的时间间隔以实现更高保真度的视觉和音频检查,然后才接听。由于模型在调用前后观察到不同的采样粒度——一个稀疏的全局预览和更密集的局部剪辑——我们引入了TimeAnchor,它使工具的时序参数在这些粒度之间保持有效且往返一致,而不是绑定于特定采样率的帧索引。为了使这种工具使用行为可训练,无需昂贵的人工间隔注释,我们构建了一个时间增强数据引擎,通过视频编辑和合成合成工具使用后的训练轨迹。跨全模态和视频基准测试的实验表明,OmniReasoner 不仅提升了答案的准确性,还能提升时间基础,同时将高保真计算集中在信息区。代码可在此 https URL 访问。
Keyword: diffusion policy
There is no result