生成时间: 2026-10-01 23:14:35 (UTC+8); Arxiv 发布时间: 2026-10-01 20:00 EDT (2026-10-02 08:00 UTC+8)
今天共有 72 篇相关文章
Keyword: reinforcement learning
A Moving-Horizon Approximate Branch-and-Reduce Method for Deep Classification Trees
一种用于深分类树的移动视界近似分支与缩减方法
- Authors: Chenxuanyin Zou, Jiayang Ren, Qiangqiang Mao, Jing Liu, Marcus Lai, Yankai Cao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2609.38194
- Pdf link: https://arxiv.org/pdf/2609.38194
- Abstract
Despite the importance for interpretability, decision trees face severe scalability challenges. Existing global optimal methods are often limited by binary feature selection and shallow tree depths, whereas traditional heuristic approaches frequently sacrifice predictive accuracy. To overcome these limitations, this paper proposes a moving-horizon approximate branch-and-reduce method to train near-optimal deep classification trees on large-scale datasets with continuous features. Built on a hierarchical root-subtree optimization framework, the method solves the root-level problem via branch-and-reduce while approximating the induced subtree problem using greedy heuristics. Although the underlying framework is capable of guaranteeing global optimality, the approximation, which functions as a lookahead rollout in a reinforcement learning context, significantly boosts efficiency for deeper structures. A low-cost moving-horizon strategy is then employed to iteratively refine model accuracy. Extensive numerical results demonstrate that our method exceeds the testing accuracy of existing heuristic baselines while offering significantly greater scalability, in terms of both dataset size and tree depth, than global optimal solvers.
- 中文摘要
尽管可解释性重要,决策树仍面临严重的可扩展性挑战。现有的全局最优方法常受二元特征选择和浅树深度限制,而传统的启发式方法则常牺牲预测准确性。为克服这些限制,本文提出一种移动视界近似分支缩减方法,用于在具有连续特征的大规模数据集上训练近优的深度分类树。该方法基于层级根子树优化框架,通过分支和约简解决根级问题,同时利用贪婪启发式近似诱导子树问题。尽管底层框架能够保证全局最优性,但该近似作为强化学习背景下的前瞻性推广,显著提升了深层结构的效率。随后采用低成本的移动视界策略迭代优化模型精度。大量数值结果表明,我们的方法超越了现有启发式基线的测试精度,同时在数据集规模和树深度方面提供了显著更高的可扩展性,优于全局最优求解器。
SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets
SynIL:利用协同效应从不完美演示数据集中进行离线模仿学习
- Authors: Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38225
- Pdf link: https://arxiv.org/pdf/2609.38225
- Abstract
Imitation learning enables robots to acquire complex skills directly from massive demonstration datasets, but its performance degrades severely when datasets are contaminated with suboptimal or noisy demonstrations. While prior quality-assessment methods attempt to filter or reweight data, they typically rely on manual pre-selection of expert reference data or task-specific heuristics, limiting scalability. To address this challenge, we introduce SynIL (Synergy-based Imitation Learning), a novel framework for automated, label-free demonstration quality assessment in offline reinforcement learning. Grounded in neuroscientific evidence that motor synergy, a low-dimensional coordinated structure in movement, correlates directly with motor proficiency, SynIL algorithmically quantifies synergy manifestation to generate dense, transition-level reward signals via self-supervised reward regression. Comprehensive evaluations on D4RL locomotion benchmarks and multi-human Robomimic manipulation datasets demonstrate that synergy-derived rewards correlate strongly with ground-truth rewards. Furthermore, SynIL substantially outperforms Behavior Cloning (BC) and achieves performance comparable to, and in sparse-reward human teleoperation scenarios, superior to, offline reinforcement learning trained on true environment rewards.
- 中文摘要
模仿学习使机器人能够直接从海量演示数据集中习得复杂技能,但当数据集被次优或噪声干扰时,其性能会严重下降。以往的质量评估方法尝试过滤或重权数据,但通常依赖于手动预选专家参考数据或任务特定的启发式方法,限制了可扩展性。为应对这一挑战,我们引入了SynIL(基于协同的模仿学习),这是一个用于离线强化学习中自动化、无标签的演示质量评估的新框架。基于神经科学证据,表明运动协同是运动中低维协调结构,与运动熟练度直接相关,SynIL通过算法量化协同表现,通过自监督奖励回归生成密集的过渡级奖励信号。对D4RL运动基准和多人机器人模拟操控数据集的综合评估表明,协同衍生的奖励与真实奖励高度相关。此外,SynIL的表现远超行为克隆(BC),在稀疏奖励人类远程操作场景下,表现优于基于真实环境奖励训练的离线强化学习。
On the Off-Policy Teacher in On-Policy Distillation
论非政策教师在非政策提炼中
- Authors: Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Blöbaum, Purak Jain
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.38360
- Pdf link: https://arxiv.org/pdf/2609.38360
- Abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
- 中文摘要
On-policy 蒸馏(OPD)最近成为一种有前景的培训后范式,学生在教师密集监督下,从自身策略生成的轨迹中学习。然而,OPD引入了根本性的不对称性:虽然抽样轨迹对学生来说是政策内,但对教师却是非策略。教师通常优化为继续使用自身策略生成的前缀,但在OPD中必须监督学生生成的前词。实证上,我们发现随着前缀的变长,其延续表现会下降。为解决这一问题,我们提出了学生-教师更新(SCOUT),这是一种共同培训框架,使教师适应学生生成的前缀。除了标准的OPD更新外,SCOUT还定期通过带有可验证奖励的强化学习优化教师的条件能力,教师从学生前缀生成延续,从结果奖励中学习。受控实验表明,SCOUT提升教师从学生生成前缀继续的能力,支持学生条件反射教师适应的预期机制。在多种师生配置、模型量表和推理领域中,SCOUT也持续提升政策提炼的有效性。
From Codebase to Culprit (C2C): Reducing the Search Space for Bugs with Semantic Retrieval and Hierarchical Reinforcement Learning
从代码库到罪魁祸首(C2C):通过语义检索和层级强化学习减少错误搜索空间
- Authors: Ankur Garg, Corey Yang-Smith, Rishav Rishav, Ahmad Abdellatif, Samira Ebrahimi Kahou
- Subjects: Subjects:
Software Engineering (cs.SE); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38402
- Pdf link: https://arxiv.org/pdf/2609.38402
- Abstract
We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across multiple levels of granularity: files, functions, and lines of code. To mirror developer's natural top-down debugging workflows, C2C integrates semantic retrieval and Hierarchical Reinforcement Learning (HRL) in a two-stage process. First, it performs recall-oriented retrieval of buggy candidates via semantic vector similarity search using bug-report text, including available stack-trace information, against a database of embeddings, where the embeddings are fine-tuned via contrastive learning with CodeBERT. Building on this reduced search space, the HRL framework incrementally localizes bugs, reasoning from files to functions and ultimately to individual lines of code. Unlike prior approaches which operate at a single granularity, C2C enables multi-resolution localization while maintaining contextual consistency across decisions. Experiments on real-world Java and Python datasets demonstrate that C2C improves retrieval precision and localization accuracy. Ablation studies further highlight the contributions of hierarchical decomposition, structured learning signals, and reward shaping in advancing multi-level bug localization.
- 中文摘要
我们引入了C2C(从代码库到罪犯),这是一个精确的错误定位框架,逐步缩小多个粒度层级的调试搜索空间:文件、函数和代码行。为了镜像开发者自然的自上而下调试工作流程,C2C将语义检索和层级强化学习(HRL)整合为两阶段过程。首先,它通过语义向量相似性搜索,利用错误报告文本(包括可用的栈追踪信息)对有缺陷候选的回忆检索,基于嵌入数据库,通过CodeBERT的对比学习微调嵌入。基于这一缩小的搜索空间,HRL框架逐步本地化错误,从文件推理到函数,最终推理到单个代码行。与以往单粒度的方法不同,C2C支持多分辨率本地化,同时保持决策间的上下文一致性。在现实世界Java和Python数据集上的实验表明,C2C提高了检索精度和本地化准确性。消融研究进一步强调了层级分解、结构化学习信号和奖励塑造在推动多层次错误本地化方面的贡献。
Draft: A Parametric Tool for Robot Design Exploration
草图:机器人设计探索的参数化工具
- Authors: David Nguyen, Marcelo Coelho, Sangbae Kim
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.38405
- Pdf link: https://arxiv.org/pdf/2609.38405
- Abstract
Robot performance is often limited by the cost of iterating on morphology and control together, since every computer-aided design (CAD) change has to be carried into a simulation-ready model before control work begins. Co-design methods attempt to close this gap, but each uses a model generator written for a single platform or lack the use of real-world data to suggest that designs are plausible. We present Draft, a parametric generation tool whose generalized engine compiles any parametric tree of serial chains into a simulation-ready MJCF model, without CAD. It allows engineers to explore design tradeoffs through easily adjustable models and evaluate how changes influence controller performance. Draft grounds the free parameters of each design using trends fitted to a survey of $114$ actuators and $49$ published robot descriptions, so that a generated robot is anchored to real-world hardware. We validate those trends wholistically by building twins of four off-the-shelf robots, whose masses agree to $1.10\times$ geometric mean fold error. Finally, we demonstrate how Draft exposes design tradeoffs by evaluating three quadrupeds through a two-stage reinforcement learning curriculum.
- 中文摘要
机器人性能通常受限于同时迭代形态和控制的成本,因为每次计算机辅助设计(CAD)变更都必须先导入仿真准备模型,然后才能开始控制工作。协同设计方法试图弥合这一差距,但每种方法都使用为单一平台编写的模型生成器,或缺乏实际数据来表明设计是否可行。我们介绍Draft,一款参数生成工具,其通用引擎可将任意串行链的参数树编译成无需CAD的MJCF模型。它允许工程师通过易于调整的模型探索设计权衡,并评估变更如何影响控制器性能。Draft通过对114美元执行器和49美元发布机器人描述的趋势进行基础,确定每个设计的自由参数,使生成的机器人能够锚定于真实硬件。我们通过构建四台现成机器人的双胞胎,这些机器人的质量同意几何平均折叠误差为1.10美元乘以,从而全面验证了这些趋势。最后,我们展示了Draft如何通过两阶段强化学习课程评估三只四足动物,揭示设计权衡。
ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning
ArgGYM:结构化可可行推理的程序化、引擎验证基准
- Authors: İbrahim Ethem Deveci, Funda Tan Çalık, Barış Deniz Sağlam, Duygu Ataman
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.38409
- Pdf link: https://arxiv.org/pdf/2609.38409
- Abstract
Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduce ArgGYM, a procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning. ArgGYM decomposes this reasoning into twelve tasks and grounds task-specific scoring in a symbolic argumentation engine that computes the formal states used to evaluate model outputs. It includes a frozen benchmark of 1,440 verified instances across fifteen curriculum configurations, two argument preference orderings (weakest-link and last-link), and two set orderings (elitist and democratic), while the same generators and verifiers can produce fresh instances for evaluation that reduces dependence on static test sets and for verifiable-reward training. On the frozen benchmark, frontier and open-weight models show sharply different reasoning profiles: they can recover substantial parts of structured answers without solving the complete task, and performance declines in later curriculum configurations with longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers for reproducible evaluation and RLVR training.
- 中文摘要
大型语言模型推理的最新进展主要得益于基准测试和具有自动验证奖励的强化学习环境,尤其是在数学、代码和形式逻辑领域。这些设置使模型准确性更容易评估和优化,但在固定问题规范和稳定评估标准下的成功在多大程度上转化为非此类领域推理仍不明确。现实世界的推理常常在不完整且可修订的信息下进行:结论可能被临时支持,被反证推翻,通过进一步论证恢复,或在有更有力理由出现时修正。此类推理通常称为可败推理。我们介绍ArgGYM,一个程序基准和兼容RLVR的结构化可败推理训练环境。ArgGYM将推理分解为十二个任务,并在符号论证引擎中进行任务特定评分,计算用于评估模型输出的形式状态。它包括一个固定基准,涵盖1440个已验证实例,跨越15种课程配置,两个论元偏好排序(最弱环节和最后环节),以及两种集合排序(精英和民主),同时相同的生成器和验证器可以生成新的评估实例,减少对静态测试集和可验证奖励训练的依赖。在冻结基准中,前沿模型和开放权重模型显示出明显不同的推理轮廓:它们可以在不解决完整任务的情况下恢复大量结构化答案,且在依赖更长、结构更互动的后续课程配置中表现下降。我们发布了基准测试、生成器和验证器,用于可复现的评估和RLVR训练。
MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary
MOBA-VL:实时MOBA解说的事件定位多回合强化学习
- Authors: Shengyun Zhong, Xinkang Zhao, Ziyuan Chu, Linchao Zhu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.38428
- Pdf link: https://arxiv.org/pdf/2609.38428
- Abstract
Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at this https URL.
- 中文摘要
多人在线竞技场(MOBA)电竞的实时解说需要视觉语言模型(VLM)逐秒流畅且准确地解说现场比赛。现有的流媒体VLM听起来自然,但常常遗漏关键事件,如击杀和目标。为解决这一限制,我们使用游戏遥测,准确记录每个事件发生时间,作为监督信号。我们引入了MOBA-VL,一个基于该信号训练的9B参数模型,采用事件局部多回合强化学习,奖励描述每个事件的回合。我们还收集了MOBACast,包含3场MOBA游戏的860场职业比赛(约460小时),配有单词级时间戳解说,以及MOBACast-Bench,作为举办赛事的基准。在MOBACast-Bench上,MOBA-VL在完整比赛(63.25分对55.12分)和剪辑(63.45分对比DeepSeek-V4.1-Flash的56.22分)上取得了最高总体得分。事件本地化的积分还将事件回忆率从34.5提升至42.1,通过监督微调。代码和数据将会发布,演示可在匿名项目页面访问,链接为https。
PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration
PrivMeSA:通过本地远程大型语言模型协作实现的隐私感知自演进多智能体医学系统
- Authors: Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38458
- Pdf link: https://arxiv.org/pdf/2609.38458
- Abstract
Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.
- 中文摘要
部署在本地的临床大型语言模型(LLM)代理可以咨询更强大的远程模型,但这样做有风险暴露患者信息。注重隐私的委派将披露决策交给本地代理,但移除显式标识符是不够的:准标识符可能通过多轮咨询和反复患者访问积累,从而实现重新识别。我们引入PrivMeSA,一种隐私意识强、自我演进的多代理系统,学习控制披露并保留远程专业知识以供本地重用。本地代理管理每次就诊,并咨询可能请求额外信息的远程专家。强化学习平衡任务准确性与直接披露及基于登记处的重新识别风险,隐私评估基于每次会诊的完整外出转录。本地课程记忆将完成的咨询提炼成通用临床指导,并在传输前检索相关经验,使后续案例无需再次远程交换即可重复使用专业知识。记忆增长无需额外结果标签或参数更新。基于基于MIMIC-IV-ED记录的急诊科基准测试,PrivMeSA将平均任务准确率提升至委派平均15.8个百分点。在同一环境下,PrivMeSA将个人信息披露率从98.0%降至0.2%,并将患者可被缩减至十名或更少登记患者的比例从74%降至0%。
Conditional Generation of Creative Chess Puzzles with Diffusion Models
利用扩散模型的创造性国际象棋谜题的条件生成
- Authors: Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38577
- Pdf link: https://arxiv.org/pdf/2609.38577
- Abstract
While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.
- 中文摘要
尽管现代语言模型展现出令人印象深刻的生成能力,但它们常常在受限且反直觉的创造性任务中遇到困难。为解决这一限制,我们探索国际象棋谜题生成作为计算创造力和推理的严格测试平台,在这一领域中,改变单个棋子可能使整个解法失效。我们提出了一种利用掩蔽扩散模型的条件式国际象棋谜题生成新方法。与以往方法不同,我们的非方向扩散方法允许对特定战术主题和部分棋局局面进行条件化。我们引入了一种新颖的辅助任务——同时最佳走法预测,该任务使解的唯一性提升了11.6%,主题条件准确率提升了2.5%。为进一步优化解的独特性和主题条件,我们建立了基于去噪扩散策略优化(DDPO)的强化学习框架。这种强化学习训练使独特且主题匹配局面的产出提升了89.1%。最后,我们发布了首批用于国际象棋谜题生成的开放权重模型(附录B),为可控且富有创造力的生成提供了新的途径。
Reinforcement Learning with Complex (valued) Memories
复杂(有价值)记忆的强化学习
- Authors: Sathya Kamesh Bhethanabhotla, Efstratios Gavves, André Biedenkapp
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38598
- Pdf link: https://arxiv.org/pdf/2609.38598
- Abstract
Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 $\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: this https URL
- 中文摘要
部分可观测环境在深度强化学习中构成根本挑战,要求智能体从观察中压缩时间信息并保持记忆以做出有效决策。虽然存在多种方法,涵盖门控重现、注意力机制和基于模型的强化学习,但寻找能够捕捉长期依赖关系的有效表征技术仍是一个活跃的研究领域。本研究中,我们重新审视了酉循环网络(uRNNs)[Arjovsky 等,2016,Jing 等,2017],它们展示了优越的梯度流和联想回忆,表达了复杂向量空间中的重现和隐藏状态。它们保持范数的酉动态使信息能够通过长序列传播。为此,我们提出了三种不同版本的uRNN作为循环PPO架构的替换,并证明简单复原和复杂表征相位自由度增加,使得在多个可改善内存任务(包括连续控制)上相较基线取得显著提升。我们还进一步探讨如何通过与量子态测量方式进行类比,为相位感知策略保留复隐藏态的相位信息。我们的方法在Rocksample和Craftax等环境中的回报可达基线的2-3倍,这项工作为强化学习和部分可观测性问题的表征开辟了令人振奋的新方向。代码可在以下网址获取:此 https URL
Vision-Language-Action Autonomous Driving Agent with Language-based Memory
视觉-语言-动作自动驾驶智能体,基于语言的记忆
- Authors: Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.38641
- Pdf link: https://arxiv.org/pdf/2609.38641
- Abstract
Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
- 中文摘要
视觉-语言-行动(VLA)基础模型近年来成为自动驾驶的主流解决方案之一,因为它们可以利用视觉语言预训练中获得的知识,实现准确且可理解的驾驶。然而,由于图像的标记成本高,VLA只能接受有限的视觉输入,这对于确定全程停车到达顺序和长视野驾驶场景理解等依赖内存的任务存在问题。现有解决方案使用通过交叉注意访问的潜在向量记忆,这些记忆既不可解释也不可移植。本文提出AD-Memo,一种通用的VLA驾驶代理,基于语言的记忆。智能体作为其思维链(Chain-of-Thought,CoT)的扩展,输出内存以记录驾驶关键的周围物体;这些记忆成为智能体未来输入的一部分。我们策划基于内存的数据集并用两阶段方案训练VLA:监督微调(SFT)和\textit{Da Capo},这是一种新颖的半闭环强化学习(RL)算法,利用轨迹级优势和驾驶阶级优势,从而更好地分配信用。在全向停车和一般驾驶等场景下,AD-Memo提升驾驶质量,提升驾驶场景中的问答,并为其他模型提供即插即用的内存。
TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion
TERRA:地形感知重建、重定向与肌肉骨骼运动控制
- Authors: Merkourios Simos, Chengkun Li, Bianca Ziliotto, Alexander Mathis
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
- Arxiv link: https://arxiv.org/abs/2609.38653
- Pdf link: https://arxiv.org/pdf/2609.38653
- Abstract
Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and because retargeting terrain interactions to complex musculoskeletal bodies is challenging. We present TERRA, an end-to-end pipeline for terrain-aware retargeting and control of musculoskeletal locomotion. From kinematic trajectories alone, TERRA combines terrain priors, estimated contacts, and negative free-space evidence to recover task-relevant support geometry. TERRA further considers anatomical, tendon-continuity, and contact constraints during retargeting. Using the resulting motion-terrain pairs from five datasets, we successfully train a single muscle-actuated control policy on 9.4 hours of diverse locomotion. Across reconstruction, retargeting, and held-out tracking benchmarks, TERRA improves terrain accuracy, sharply reduces anatomical and interaction violations, and achieves the highest observed completion rate over supported terrain families. Overall, TERRA provides a practical route from scene-less motion data to muscle-actuated locomotion over diverse non-flat terrain. Project website: this https URL
- 中文摘要
肌肉骨骼建模和强化学习的最新进展使肌肉驱动的智能体能够重现日益复杂的人类动作。然而,这些能力主要局限于平地,部分原因是运动数据集很少包含对齐地形几何,且将地形交互重新定位到复杂肌肉骨骼体具有挑战性。我们介绍TERRA,一条端到端的地形感知重定向和肌肉骨骼运动控制流程。仅凭运动轨迹,TERRA结合地形先验、估计接触和负自由空间证据,恢复任务相关的支撑几何。TERRA进一步考虑了重新定向过程中的解剖学、肌腱连续性和接触约束。利用来自五个数据集的运动-地形对,我们成功训练了单一肌肉驱动控制策略,涵盖9.4小时的多样化运动。在重建、重新定位和持续跟踪基准测试中,TERRA提升了地形精度,显著减少了解剖和相互作用违规,并在支持地形家族中实现了最高的观测完成率。总体而言,TERRA为从无场景运动数据到肌肉驱动的多样非平坦地形移动提供了实用路径。项目网站:此链接
Forward-Invariant Policy Classes for Safe Reinforcement Learning in Multicopter Control
多旋翼控制安全强化学习的前向不变策略类
- Authors: Chieh Tsai, Muhammad Junayed Hasan Zahed, Jinzhi Shen, Yi Xie, Ruoshan Lan, Majed Obaid, Salim Hariri, Hossein Rastgoftar
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.38655
- Pdf link: https://arxiv.org/pdf/2609.38655
- Abstract
This paper proposes a reinforcement learning (RL) framework for safe gain scheduling based on forwardinvariance-induced action-space design. Under stated nominalmodel and inversion-domain assumptions, rather than enforcing safety through runtime shielding or penalty-based constraints, safety is embedded directly into the policy class. Specifically, we construct a finite library of feedback controllers sharing a common Lyapunov certificate that establishes forward invariance of a prescribed admissible set under arbitrary switching. Consequently, any policy whose actions are restricted to this library, including policies encountered during RL exploration, inherits the same certificate. Policy optimization can therefore focus on closed-loop performance without runtime safety filtering or action projection. The framework is instantiated for quadcopter hover regulation, where a DQN schedules among certified feedback controllers. Nonlinear MuJoCo simulations demonstrate state-dependent gain scheduling and empirically evaluate robustness to wind, model mismatch, sensor noise, and sensing delay. The results illustrate how safety certification can be separated from policy optimization by learning over a forward-invariant policy class.
- 中文摘要
本文提出了基于前向不变性诱导动作空间设计的安全增益调度(RL)强化学习(RL)框架。在既定的名义模型和反演域假设下,安全性不是通过运行时屏蔽或基于惩罚的约束来强制执行安全,而是直接嵌入策略类中。具体来说,我们构建了一个有限的反馈控制器库,共享一个共享的Lyapunov证书,该证书在任意切换下确立了规定可接受集合的前向不变性。因此,任何行为限制在该库中的策略,包括在强化学习探索中遇到的策略,都继承了相同的证书。因此,策略优化可以专注于闭环性能,而无需运行时的安全过滤或动作投影。该框架用于四旋翼悬停调节,DQN在认证反馈控制器之间调度。非线性MuJoCo仿真展示了状态依赖增益调度,并通过实证评估对风、模型不匹配、传感器噪声和传感延迟的鲁棒性。结果展示了如何通过学习前向不变策略类,将安全认证与策略优化区分开来。
In-Distribution Imagination for Model-Based Offline Reinforcement Learning
基于模型的离线强化学习中的分布式想象力
- Authors: Mintae Kim, Koushil Sreenath
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38673
- Pdf link: https://arxiv.org/pdf/2609.38673
- Abstract
Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Experiments show that trajectory support predicts rollout failure substantially better than transition-level uncertainty, highlighting the importance of trajectory-level rollout control in MBORL.
- 中文摘要
基于模型的离线强化学习(MBORL)通过模型生成轨迹提升样本效率。然而,累积的模型误差可能导致想象轨迹超出离线数据分布,导致合成数据不现实且策略优化不稳定。许多现有方法主要利用过渡级不确定性来控制滚动。我们提出了\emph{分布内想象}(IDI),一种扩展控制框架,用于估计学习表示空间中的轨迹支持,并自适应截断离开离线轨迹流形的滚动。结合轨迹正则化RL(熵正则化RL的扩展),IDI在有限数据环境中持续提升性能。实验表明,轨迹支持比过渡级不确定性更能预测滚动失败,凸显了MBORL中轨迹级滚动控制的重要性。
CEER2: Directional and Tunable End-Effector and Root Compliance for Humanoid Loco-Manipulation
CEER2:人形机动操作的方向性与可调末端执行器和根适应性
- Authors: Xinyuan Luo, Chunyuan Yang, Boyuan Chen, Xianyi Cheng
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.38709
- Pdf link: https://arxiv.org/pdf/2609.38709
- Abstract
Humanoids are increasingly capable of tracking complex whole-body motions, but physical interaction introduces a different challenge. When a robot makes contact with a person or the environment, it needs to respond to external forces while preserving the motion needed for the task. This response can vary across directions in the end-effectors and on the body. For example, an end effector may need to accommodate contact force in one direction while maintaining motion accuracy in another, while the robot body may resist an external force or move with it. We present a compliance framework for humanoid loco-manipulation that combines directional and tunable end-effector (EE) compliance with selectable root compliance for external force rejection or force following. A hierarchical reinforcement learning controller modulates a fixed whole-body tracking policy through high-level EE and root commands, while interaction forces are estimated from proprioceptive history. Our simulation and real-world experiments on a humanoid demonstrate directional stiffness control, online stiffness adjustment, distinct root compliance, compliant manipulation, and collaborative carrying.
- 中文摘要
类人机器人越来越能追踪复杂的全身运动,但物理互动带来了不同的挑战。当机器人与人或环境接触时,需要在保持任务所需运动的同时响应外部力。这种响应在终端执行器和身体各方向上可能有所不同。例如,末端执行器可能需要在一个方向承受接触力,同时在另一个方向保持运动准确性,而机器人身体则可能抵抗外部力或随力移动。我们提出了一个人形机车操控的合规框架,结合了方向性和可调末端执行器(EE)的顺应性,并可选择根顺从以拒绝外部力或跟随力。分层强化学习控制器通过高级EE和根指令调节固定的全身追踪策略,而相互作用力则根据本体感觉历史估计。我们在类人生物上的模拟和现实实验展示了方向刚度控制、在线刚度调整、明显的根部顺应、顺应操作和协作携带。
Code to Control: Synthesizing Parameterized Reactive Controllers
代码到控制:参数化无功控制器的综合
- Authors: Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38733
- Pdf link: https://arxiv.org/pdf/2609.38733
- Abstract
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.
- 中文摘要
近期基于LLM的控制方法要么调用语言模型选择动作,要么合成需要在每个决策时规划的世界模型,从而引入延迟,限制实时使用。我们介绍Code to Control,这是一种合成直接以策略形式执行的Python控制器的方法。Code to Control将程序结构与参数分离。LLM综合控制器结构,而无导数搜索则利用环境反馈拟合其参数以实现持续控制。一旦学习,最终的控制器无需在决策时推理或规划,实现实时游戏体验,并且在我们的时序协议下,动作选择速度比PPO策略更快。在一系列Atari游戏、Flappy Bird和MuJoCo任务中,Code to Control优于基于规划的程序合成方法,在深度强化学习中保持竞争力,同时使用更少的环境交互,跨越环境动态的重大变化,并可扩展到复杂的运动任务。
Learning to Route in Visual Space via Multi-Step Embedding Retrieval
通过多步嵌入检索学习在视觉空间中布线
- Authors: Tianyu Chen, Mingyuan Zhou, Jiaxing Wu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2609.38743
- Pdf link: https://arxiv.org/pdf/2609.38743
- Abstract
LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.
- 中文摘要
LLM代理依赖检索工具访问外部知识,但视觉代理搜索仍被标准单步检索器严重限制。在当前的流程中,代理必须对每个中间步骤发出文本查询,当视觉线索难以描述或检索器未能在顶部结果中呈现必要的中间证据时,表现会很吃力。我们假设将整个嵌入空间的多步导航直接卸载给检索工具,可以解决这一性能瓶颈。为系统研究这一点,我们引入了VHOP,这是一个灵活的数据生成框架和基准测试,包含五个核心难度等级,测试视觉匹配和搜索规划。利用该框架,我们开发了VHOP-Router,一个端到端的训练流水线---结合监督微调、在线模仿学习和强化学习---将标准嵌入模型转变为自回归多步检索器。VHOP-Router直接在视觉潜在空间中运行,只需一次工具调用即可检索链接的图像链,无需代理生成中间文本查询。实验显示,VHOP-Router将检索性能从5%以下提升至76.3%。在代理搜索中,VHOP-Router将任务成功率提升52.7%,平均令牌长度从1886%减少到728%,而升级代理仅提升3.7%。与强基线(代理每步检索前50个结果)相比,VHOP-Router在将上下文图像成本降低23美元、累计API负载减少35%的同时保持了优越性能。模型还能稳健地推广到未见难度和现实测试集。最终,VHOP和VHOP-Router为视觉代理搜索提供了高效且有效的解决方案,同时完全保留了原生LLM的功能。
Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
得分更高,回答更差:通过协议级评分标准缓解基于评分标准的强化学习中的奖励黑客问题
- Authors: Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38847
- Pdf link: https://arxiv.org/pdf/2609.38847
- Abstract
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at this https URL
- 中文摘要
基于评分标准的强化学习(Rubric-RL)训练没有验证者的语言模型。评委检查评分标准的每个标准,并将裁决汇总成奖励,通常通过加权和。我们表明这种加法聚合是弱点。在和项下,标准相互补偿:一个错过关键决策的政策可以用没人问的建议买回分数。在临床咨询中,这种政策得分更高,回答更差。评分标准覆盖率上升,而医生提出的标准适用性低于未训练模型。医学标准并非罪魁祸首。将标准分组以保持一致,保持原有的标准可回收三分之一的损失;简短答案几乎不恢复任何损失。因此,我们提出了协议级评分标准(ProRubric),它保留了标准所要求的内容,并改变了它们的汇总方式。它将检查表分组为几个协议级别的维度。只有当所有标准都成立且失败条款未触发时,一个维度才算数。分组仅在离线时完成一次,优化器保持不变。ProRubric 在不丢失覆盖范围的情况下提高了适宜性 10.8 个百分点,并且在两个量表上均拥有最佳的七个基准平均值。奖励有效性不仅由评分标准验证的内容决定,还取决于其如何汇总。代码可在此 https URL 获取
OpenJev-RLCD: A Working RLCD Implementation
OpenJev-RLCD:一个可运行的RLCD实现
- Authors: Zhimin Gao, Pichao Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.38850
- Pdf link: https://arxiv.org/pdf/2609.38850
- Abstract
Decision models such as Jev answer questions with probabilities, which are only useful if they are calibrated. Open-source reproductions rely on supervised fine-tuning plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model samples a rationale, and we score the answer distribution it commits to afterwards with a strictly proper scoring rule. A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per-rationale objective either switches reasoning off or is drowned out by policy-gradient noise, which leads to a two-stage recipe: calibrate, then reinforce. With Qwen3-1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and GRPO (each temperature-scaled) in accuracy and beats all of them in selective prediction; on GSM8K answer verification a single query decides \gvTwoCovFive\% of the items at $\le$5\% error, versus \gvGrpoCovFive\% for GRPO. When uncertainty comes from annotator disagreement, RLCD provably cannot beat cross-entropy. Code and results: this https URL.
- 中文摘要
像Jev这样的决策模型以概率回答问题,只有在校准后才有用。开源再现依赖监督微调和温度尺度,而可验证奖励强化学习(RLVR)使推理模型过于自信。我们展示了校准决策强化学习(RLCD)推理模型的工作实现:模型采样一个理由,然后用严格正确的评分规则对其承诺的答案分布进行评分。方差恒等式表明,对多个样本混合进行评分会奖励不同意的理由,而RLVR正是这种混合目标,但没有多样性项。在简单优化下,每个理性目标要么关闭推理,要么被策略梯度噪声淹没,导致两阶段方案:校准,然后强化。在Qwen3-1.7B的两个推理任务(3个种子,配对测试)上,RLCD在准确度上匹配或击败SFT、RFT/STaR和GRPO(每个温度尺度后),并在选择性预测中胜过所有这些指标;在GSM8K答案验证中,单一查询决定了\gvTwoCovFive\%的项目以$\le$5\%错误判定,而GRPO则判定\gvGrpoCovFive\%。当不确定性源于注释者分歧时,RLCD可以证明无法战胜交叉熵。代码和结果:此https URL。
GeoNest: Learning to Select Failure-Aware Neighborhoods for the Irregular Knapsack Problem in a Circular Container
GeoNest:学习在圆形容器中选择故障感知邻域以解决不规则背包问题
- Authors: Zhongman Du, Huiming Zhang, Linlin Yang, Sheng Xu, Baochang Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38863
- Pdf link: https://arxiv.org/pdf/2609.38863
- Abstract
The two-dimensional irregular knapsack problem in a fixed circular container is an important combinatorial optimization problem for maximizing material utilization in manufacturing. Conventional geometric packing solvers can produce tightly packed layouts, yet they often partition the residual space into isolated small pockets that cannot fit valuable unplaced polygons. To overcome this late-stage packing bottleneck, we propose a failure-aware large neighborhood search framework named GeoNest, driven by a graph policy trained via reinforcement learning. Specifically, we first construct neighborhoods by pairing failed target polygons with residual pockets. We then use explanatory poses to identify the placed polygons that block candidate insertions. These diagnosed blocking relations define bounded, fixed-item repair subproblems for the underlying geometric solver. Finally, the graph policy selects the most promising subproblem for execution. For evaluation, we introduce CircleNest-Bench, a benchmark comprising 2,391 load-controlled instances from four contour sources, including a held-out industrial CAD source. Experimental results demonstrate that, under the same total time budget, GeoNest improves mean utilization over a state-of-the-art standalone packing solver by about 0.9% on average across the three main test sets and by about 0.6% on the held-out industrial set.
- 中文摘要
固定圆形容器中的二维不规则背包问题是最大化制造材料利用的重要组合优化问题。传统的几何包装求解器可以生成紧凑的布局,但它们常常将剩余空间划分为孤立的小口袋,无法容纳有价值的未放置多边形。为克服后期包装瓶颈,我们提出了一个名为GeoNest的失败感知大型邻域搜索框架,由通过强化学习训练的图策略驱动。具体来说,我们首先通过将失败的目标多边形与残差口袋配对来构建邻域。然后使用解释性姿态识别阻碍候选插入的放置多边形。这些诊断出的阻塞关系为底层几何求解器定义了有界的固定项修复子问题。最后,图策略选择最有前景的子问题进行执行。作为评估,我们介绍了CircleNest-Bench,这是一个基准测试,包含来自四个等高线来源的2,391个负载控制实例,其中包括一个保留的工业CAD源。实验结果表明,在相同的总时间预算下,GeoNest在三大主要测试集的平均利用率相比最先进的独立包装求解器平均提升约0.9%,在保留的工业测试集中提升约0.6%。
DODGER: Safety-Guided Reinforcement Learning for Robot Navigation Among Dynamic Obstacles
DODGER:安全引导强化学习,用于动态障碍物中的机器人导航
- Authors: Sanghyuk Park, Kwanwoo Lee, Taekyung Kim, Seohyeon Lim, Yisoo Lee
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.38873
- Pdf link: https://arxiv.org/pdf/2609.38873
- Abstract
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.
- 中文摘要
在以人为中心的环境中运行的机器人必须安全穿越多个动态障碍物,以避免与人及周围基础设施碰撞。控制障碍函数(CBF)为安全过滤提供了有效的机制,近期基于CBF的强化学习(RL)方法将此类安全信息嵌入到学习的策略中。然而,在训练期间仅执行安全过滤动作可能会限制策略探索,这一限制在动态场景中尤为重要,因为安全依赖于机器人与障碍物的相对运动。我们提出了DODGER,一个安全引导的强化学习框架,直接执行策略生成的动作以推动训练推广,同时利用CBF过滤的引用和约束违规来塑造避免碰撞的策略。我们通过Dubins汽车安全分析评估DODGER,并在全序类人模拟和基于激光雷达的感知下,演示了在多个动态障碍中实现目标导向导航,并利用基于激光雷达的感知实现,无需运行时安全滤波器。
Doing More with Less Tokens: Hierarchical Reinforcement Learning for Efficient Coding Agents
用更少的代币做更多事:高效编码代理的分层强化学习
- Authors: Haobin Li, Liang Jiang, Zhenyu Huang, Mouxing Yang, Xi Peng
- Subjects: Subjects:
Software Engineering (cs.SE)
- Arxiv link: https://arxiv.org/abs/2609.38885
- Pdf link: https://arxiv.org/pdf/2609.38885
- Abstract
Recently, coding agents have emerged as a dominant paradigm for real-world software engineering (SWE) scenarios, which solve complex tasks through multi-turn interactions with development environments. However, frequent interactions with environments would inevitably introduce substantial token overhead, leading to high usage costs and latency. Although recent studies have explored reducing token usage by context manipulation and interaction limits at inference time, these approaches focus on improving token efficiency while overlooking the risk of discarding task-relevant information, thus struggling to balance the trade-off between resolution rate and token efficiency. In this paper, we study a more general paradigm without suffering from the limitation, i.e., training token-efficient coding agents with promising resolution performance, which is a highly-practical yet less-explored problem. To this end, we reveal two core observations in SWE scenarios: i) Efficiency Variation: successful resolution could be achieved with fewer tokens; ii) Entropy Correlation: unproductive behaviors are associated with turn-level entropy. Motivated by observations, we propose a novel reinforcement learning framework, dubbed HERO. Specifically, HERO prioritizes task resolution over token efficiency during policy optimization and encourages efficient reasoning patterns at both trajectory and turn levels. Extensive experiments on SWE-bench Verified and SWE-bench Multilingual demonstrate that HERO achieves a favorable trade-off between resolution rate and token efficiency compared with state-of-the-art coding agents and reinforcement learning methods.
- 中文摘要
近年来,编码代理已成为现实软件工程(SWE)场景中的主导范式,这些场景通过与开发环境的多回合交互解决复杂任务。然而,频繁的环境交互不可避免地会带来显著的令牌开销,导致高使用成本和高延迟。尽管近期研究探讨通过上下文操作和推理时的交互限制来减少令牌使用,但这些方法侧重于提升令牌效率,同时忽视了丢弃任务相关信息的风险,因此难以在解决率与令牌效率之间取得平衡。本文研究了一个更通用的范式,避免了训练具有良好分辨率表现的令牌高效编码代理,这是一个高度实用但较少被探讨的问题。为此,我们揭示了软件工程场景中的两个核心观察:i)效率变异:成功解析可以通过更少的令牌实现;ii)熵相关性:非生产性行为与回合级熵相关。基于观察,我们提出了一种新的强化学习框架,称为HERO。具体来说,HERO在策略优化中优先考虑任务解决而非代币效率,并鼓励在轨迹和转向层面实现高效的推理模式。对SWE-bench Verified和SWE-bench Multilingual的广泛实验表明,HERO在分辨率率和代币效率之间相较于最先进的编码代理和强化学习方法实现了有利权衡。
PRICE the Action Chunks: Physical Relational Credit Assignment for Embodied Reinforcement Learning
为行动块定价:具身强化学习中的物理关系学分作业
- Authors: Yangang Zou, Jiajun Lu, Weitao Zhou, Haibao Yu, Bozhou Zhang, Jiawei Wang, Honglong Tian, Minglei Li, Li Zhang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.38890
- Pdf link: https://arxiv.org/pdf/2609.38890
- Abstract
Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.
- 中文摘要
基于结果的强化学习(RL)在训练后使用终端成功信号进行愿景-语言-行动策略,但为每个行动块赋予相同的轨迹层级优势。失败的事件因此可能惩罚有用的早期行动,仿佛它们导致了失败。现有方法寻求通过学习过的评估者获得更细粒度的反馈,增加任务特定监督或额外的模型训练。我们首次探讨了跨轨迹的物理关系是否能仅凭终端结果在具身强化学习中提供行动块信用,无需辅助评估器。关键见解是,推广达到对应的物理情境可以相互参考:其终端结果为评估局部进展提供了证据。我们引入了从事件中推断学分的物理关系(PRICE),包含两个组成部分:(i)物理关系图,将当前和历史结果在对应区块边界汇聚以估算成功潜力;(ii)基于置信门控学分分配,利用这些潜力的变化细化轨迹级监督。我们的分析将预言机电位变化与终极成功目标联系起来,并为结果无关证据池提供了有限样本方向界限。独立的延续测试显示,PRICE保留的学分与局部进展相符,而在LIBERO、RoboTwin 2.0和真实机器人上的实验显示,任务成功率优于基于结果的基线和更快的学习速度。
From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning
从图像解读到临床推理:上游医生情境感知多模态学习与因果强化学习
- Authors: Jialu Pi, Yanan Ma, Weijie Chen, Owen Crystal, Shubham Trivedi, Stephen Xie, Anna Silverman, Matthew Stib, Chadi Ayoub, Reza Arsanjani, Imon Banerjee
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.38924
- Pdf link: https://arxiv.org/pdf/2609.38924
- Abstract
Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.
- 中文摘要
重大不良心血管事件(MACE)仍是全球死亡的主要原因。利用常规临床数据的机会性筛查,提供了一种可扩展的方法,可以在急性事件发生前识别高风险个体。尽管胸部X光(CXR)捕捉潜在心血管生物标志物和临床病史提供了补充的患者背景,但现有的医学视觉语言模型主要优化为放射学解读,而非预后推理。我们提出了一种因果强化学习框架,用于多模态临床推理,整合CXR和医生撰写的临床病史,用于机会性MACE预测。该框架引入了(1)角色解耦的双LLM架构,将推理与风险预测分离,(2)双行动因果强化学习策略用于证据选择和推理优化,以及(3)因果标记剪枝以学习紧凑的多模态表示。在内部队列、急诊科队列和外部MIMIC数据集中评估时,所提出的框架持续优于单模态基线和最先进的医学视觉语言模型,分别实现了0.720、0.760和0.845的AUROCs。它还显著提升了推理质量,获得了更高的绿色(GREEN)分数和更高的专家偏好,同时在多样化患者群体中保持了稳健的预测表现。
Local-Minimum Escaper: Programmatic Subgoal Generation for Robust Navigation in Unknown Environments
局部最小逃逸器:在未知环境中实现稳健导航的程序化子目标生成
- Authors: Yin Gu, Xinming Zhang, Shanze Wang, Siwei Cheng, Wei Zhang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.38928
- Pdf link: https://arxiv.org/pdf/2609.38928
- Abstract
Mapless navigation in unknown and partially observable environments remains challenging for mobile robots, particularly when local minima prevent the robot from making progress toward its goal. Existing local navigation methods often lack an explicit mechanism for escaping such situations, while deep reinforcement learning (DRL) approaches typically learn recovery behaviors implicitly through reward design and policy optimization. In this work, we propose \textbf{LME} (Local-Minimum Escaper), a programmatic hierarchical framework that explicitly generates and reasons subgoals to guide robots out of local-minimum regions. LME operates solely on local observations and selects candidate subgoals using interpretable heuristic criteria that account for both surrounding obstacle geometry and candidate-location safety. A local planner then generates low-level motion commands toward the selected subgoal. This design enables LME to handle environments both with and without local minima within a unified framework, while remaining independent of the underlying local planner and requiring no additional training. Extensive experiments in simulated and real-world environments demonstrate that LME provides robust navigation performance and generalizes to challenging unseen scenarios. Furthermore, the generated subgoals can be used to guide different local planners, substantially improving their ability to escape local minima. Successful deployments on both differential-drive and quadruped robots further demonstrate the practical applicability and generality of the proposed framework.
- 中文摘要
在未知且部分可观察的环境中,无地图导航对移动机器人来说依然具有挑战性,尤其是在局部极小值阻碍机器人向目标前进时。现有的局部导航方法通常缺乏明确的逃脱机制,而深度强化学习(DRL)方法通常通过奖励设计和策略优化隐式学习恢复行为。本研究提出 \textbf{LME,Local-Minimum Escaper,一种程序化的层级框架,明确生成并推理子目标,引导机器人脱离局部最小区域。LME 仅基于局部观测,并利用可解释的启发式标准选择候选子目标,这些标准兼顾周围障碍几何和候选位置安全。局部规划者随后生成针对所选子目标的低级运动指令。该设计使LME能够在统一框架内处理有无局部极小值的环境,同时保持独立于底层本地规划器,无需额外训练。在模拟和现实环境中的大量实验表明LME提供了稳健的导航性能,并能推广到具有挑战性的未知场景。此外,生成的子目标可用于指导不同的本地规划者,显著提升他们逃离局部最小值的能力。在差动驱动和四足机器人上的成功部署进一步展示了该框架的实用性和通用性。
Robust Risk-Sensitive Reinforcement Learning from Corrupted Human Feedback
从腐败的人类反馈中学习的强化风险敏感强化
- Authors: Xinyi Ni, Lifeng Lai
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38938
- Pdf link: https://arxiv.org/pdf/2609.38938
- Abstract
Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under adversarial preference-label flips. We consider additive linear rewards and a fixed-reference protocol with one comparison per episode and at most $C$ flipped labels over $K$ episodes. We propose weighted streamed-preference CVaR RLHF (WSP-CVaR-RLHF), which combines uncertainty-weighted reward estimation with optimistic augmented-state CVaR planning. For known transitions and normalized rewards, we establish the regret bound $\widetilde{O}\left(\frac{d}{\kappa}\sqrt{\frac{K}{\alpha}}+\frac{dC}{\kappa\alpha}\right)$ up to lower-order terms, where $d$ is the reward-feature dimension, $\alpha$ is the CVaR level, and $\kappa$ characterizes the preference link. The bound separates the clean statistical cost from the penalty caused by corrupted feedback. We further extend the analysis to unknown tabular transitions, where the trajectory distribution entering the CVaR objective must be learned together with the reward. We address the resulting coupled uncertainty using rectangular transition confidence sets, joint optimistic planning, and a history-level CVaR simulation argument. Experiments under four adversarial attacks demonstrate that WSP-CVaR-RLHF consistently reduces cumulative regret relative to its unweighted robust counterpart while preserving confidence-set coverage.
- 中文摘要
人类反馈强化学习(RLHF)通过人类比较学习,这些比较可能被破坏或故意操控。本文研究在线风险敏感RLHF,采用静态条件风险值(CVaR)在对抗性偏好标签翻转下。我们考虑了加法线性奖励和固定引用协议,每集进行一次比较,最多$C$标签在$K美元剧集中翻转。我们提出了加权流偏好CVaR RLHF(WSP-CVaR-RLHF),结合了不确定性加权奖励估计与乐观增强状态CVaR规划。对于已知的转移和归一化奖励,我们建立遗憾界限$\widetilde{O}\left(\frac{d}{\kappa}\sqrt{\frac{K}{\alpha}}+\frac{dC}{\kappa\alpha}\right)$至低阶项,其中$d$为奖励特征维度,$\alpha$为CVaR水平,$\kappa$为偏好链接。该界限将纯净的统计成本与反馈损坏引起的惩罚区分开来。我们进一步将分析扩展到未知的表格转移,其中必须与奖励一起学习进入CVaR目标的轨迹分布。我们通过矩形转移置信集、联合乐观规划和历史级CVaR模拟论证来解决耦合不确定性。在四种对抗性攻击下的实验表明,WSP-CVaR-RLHF相较于未加权的稳健对应者,能够持续降低累计遗憾,同时保持置信度集的覆盖率。
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
通过Q分数匹配的顺序值恢复实现无环逆强化学习
- Authors: Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38955
- Pdf link: https://arxiv.org/pdf/2609.38955
- Abstract
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
- 中文摘要
逆强化学习(IRL)旨在恢复一个能解释专家演示的奖励函数。现有的IRL方法通常依赖于在奖励学习和策略优化之间交替进行的双级优化过程,导致计算负担和训练不稳定性。在本研究中,我们引入了一条不同的路径,通过利用扩散策略完全消除策略优化。我们的关键见解是,扩散策略编码了最优软Q函数的作用梯度结构,使奖励学习能够被表述为一系列价值恢复问题,从而绕过以往IRL方法固有的奖励策略循环。具体来说,我们的方法分为三个阶段进行:(I)通过动作梯度匹配恢复最优软Q函数,并以Gumbel回归启发的方式估计相应的软价值函数(Q值的LogSumExp);(II)通过推断状态依赖的偏移来校准这些软值;(III)通过强制Bellman一致性提取奖励。这导致了无环逆强化学习(LFIRL),这是一种完全离线的算法,以简单、无环且顺序的方式运行。LFIRL实现简单,显著提升训练效率,同时保持强劲的奖励恢复表现。在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基准测试中,LFIRL在最快基线上实现了2-3倍的加速,同时在奖励恢复质量上与最先进方法相当甚至超越。
OccluDex: Hierarchical 3D Visuo-Tactile Representation Learning for Egocentric Dexterous Manipulation under Self-Occlusion
OccluDex:层级三维视觉触觉表征学习,用于自我封闭下以自我为中心的灵巧操作
- Authors: Ziheng Xu, Yueyuan Chen, Xinyuan He, Guoxing Liu, Yuanshuo Tan, Huiming Pan, Bin He, Shuo Jiang, Peter B. Shull
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.39017
- Pdf link: https://arxiv.org/pdf/2609.39017
- Abstract
Reliable dexterous manipulation requires continuous estimation of object geometry and hand-object contact throughout interaction. With egocentric sensing, however, the manipulating hand frequently occludes task-relevant object surfaces and contact regions, reducing the visual evidence available for state estimation and thereby making robust closed-loop control and generalization to unseen object geometries particularly challenging. To address this, we present OccluDex, a hierarchical 3D visuo-tactile representation learning framework that integrates global geometric structure with local contact information for robust manipulation under dynamic self-occlusion during hand-object interaction. OccluDex adopts multi-scale masked autoencoding to progressively encode partial 3D geometry and fuses tactile contact tokens with high-level geometric features through cross-modal attention. The encoder is pretrained from synchronized human visuo-tactile demonstrations and transferred as a frozen perceptual backbone for downstream reinforcement learning. We evaluate OccluDex on a faucet rotation task, requiring one full clockwise handle revolution, and a tabletop object reorientation task, requiring a 180-degree tabletop object reorientation without toppling. In simulation experiments, OccluDex demonstrated 12.6% higher accuracy for unseen objects and 8.3% higher accuracy for previously seen objects than the strongest state-of-the-art baseline models. Physical experiments were further performed with a Shadow Hand to demonstrate successful zero-shot sim-to-real generalization on unseen physical objects. This results could enable humanoid egocentric object manipulation for seen and unseen objects even when the manipulating robotic hand occludes vision.
- 中文摘要
可靠的灵巧操作需要在整个交互过程中持续估计物体几何体和手与物体接触。然而,在自我感知中,操作手经常遮挡任务相关的物体表面和接触区域,减少可用于状态估计的视觉证据,从而使稳健的闭环控制和对看不见物体几何的泛化尤为困难。为此,我们提出了OccluDex,一种层级三维视觉-触觉表征学习框架,整合全局几何结构与局部接触信息,实现手与物体交互动态自遮蔽下的稳健操作。OccluDex采用多尺度掩码自编码,逐步编码部分三维几何,并通过跨模态关注融合触觉接触标记与高级几何特征。该编码器通过同步的人类视觉-触觉演示预训练,并作为冻结的感知骨干传输,用于后续强化学习。我们在水龙头旋转任务中评估OccluDex,该任务需要顺时针旋转一圈手柄,以及桌面物体重新定向任务,需180度桌面物体重新定向且不倾倒。在模拟实验中,OccluDex对未见物体的准确率提升了12.6%,对已见物体的准确率提升了8.3%,优于最强的最先进基线模型。进一步使用Shadow Hand进行物理实验,展示了对未可见物理物体的零时模拟到现实的成功推广。这一结果可能使类人以自我为中心的物体操作能够实现可见和看不见物体的操作,即使机械手遮挡了视线。
Frame Differential On-Policy Self-Distillation for Video Reasoning
视频推理中的帧差分策略自蒸馏
- Authors: Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.39021
- Pdf link: https://arxiv.org/pdf/2609.39021
- Abstract
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.
- 中文摘要
强化学习(RL)通过可验证的奖励和日益细粒度的视觉或时间赋值,显著提升了多模态语言模型的推理能力。然而,在视频推理中,当前的强化学习方法通常以固定稀疏帧预算训练:增加帧数会使自回归展开成本高昂,而帧太少可能会错过时间局部事件和细粒度的视觉细节。我们介绍了 \textbf{帧差分策略自蒸馏(FD-OPSD)},它在强化学习训练中将稠密帧观察的有用证据转移到稀疏帧策略中。FD-OPSD 比较策略中对同一采样响应在稀疏和稠密视图下的令牌级偏好,并在无外部教师或密集自回归展开的情况下提炼出帧差分信号。该方法保留稀疏帧展开,保持推断不变。在六个视频推理基准测试的Qwen2.5-VL-7B和Qwen3-VL-4B中,FD-OPSD在16、32和64帧评估设置下,整体平均表现优于最强的对应GRPO、T-GRPO或视频-KTR基线。这些结果表明,密集视觉证据可以通过令牌级自我蒸馏在训练过程中选择性转移,同时保留稀疏的帧展开和不变的推断。
Structure-aware Reinforcement Learning for Protein Directed Evolution
蛋白质导向进化的结构感知强化学习
- Authors: Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39048
- Pdf link: https://arxiv.org/pdf/2609.39048
- Abstract
Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
- 中文摘要
蛋白质优化仍是生命科学领域的长期目标。现有机器学习辅助定向进化(MLDE)方法主要依赖仅序列特征,忽视了蛋白质结构中编码的关键空间约束和共进化相互作用。然而,由于可靠突变结构稀缺,直接整合结构信息仍具挑战性。为解决这些问题,我们提出了StructEvo,一种针对蛋白质导向进化的新型结构感知强化学习框架。StructEvo采用δ结构融合编码器,通过特征差异近似突变结构特征,实现空间知识的动态整合。庞大的突变空间随后通过结构对齐的层级行动网络分解为可管理的子空间,几何约束进一步稳定δ特征学习。我们的方法在两个具有挑战性的优化基准测试中,比以往最先进方法分别高出9.2%和16.3%,并进一步识别出GFP中经过实验验证的上位模式,凸显了结构指导对于有效蛋白质导向进化的重要性。
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
LexReward:一个基于分类法的法律语言模型奖励框架
- Authors: Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.39071
- Pdf link: https://arxiv.org/pdf/2609.39071
- Abstract
Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.
- 中文摘要
法律语言模型需要奖励信号,不仅能体现答案的正确性,还能体现法律回应的多维质量。然而,现有的奖励方法往往依赖粗粒度的整体判断,提供了有限的领域特异性和可解释性。我们介绍了LexReward,一个基于分类法的法律奖励建模框架。LexReward将法律响应质量从三个互补维度来描述:风格,涵盖词汇和句法质量;元素,评估法律主题、事实、法规和判决;链条,评估法律推理的顺序、完整性、正确性和非冗余性。对于每个维度,我们制定了定义评估标准和质量水平的评分标准。所得的奖励用于构建直接偏好优化(DPO)和奖励模型训练的两两偏好数据。实验表明,基于评分标准的奖励能够可靠区分不同质量的法律回应,且基于偏好数据的DPO训练提升了三个维度的表现。学习奖励模型LexRM也支持有效的下游优化:每个维度特定的奖励模型通过强化学习提升其对应维度的政策表现,而无需在奖励时提供参考答案。按维度进行的分析进一步支持所提分类法和奖励构建的有效性。
Beyond Prediction: Steering VLM Agents with Retrospective World Modeling
超越预测:用回顾性世界建模引导VLM代理
- Authors: Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39101
- Pdf link: https://arxiv.org/pdf/2609.39101
- Abstract
Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution $P(\hat{a}{t}|s_t, s{t+1})$ for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.
- 中文摘要
为VLM代理配备世界建模能力,已显示出复杂推理和长期规划的强大潜力,同时减少了策略学习对昂贵现实世界交互的依赖。现有方法主要依赖前瞻性仿真来预测候选行动的后果。然而,这种仅向前的范式关注下一步会发生什么,并提供了有限的约束来验证行为是否与观察到的状态转变在因果上一致,可能导致看似合理但物理上不连贯的行为。本文挑战了将世界建模视为仅是前瞻性预测的观点,并引入了回顾性世界建模,这是一种新的代理学习范式,使代理能够通过估计最可能导致某一转移的动作的回溯归因分布$P(\hat{a}{t}|s_t, s{t+1})$来进行推理。基于这一能力,我们制定了自一致性奖励(SCR),这是一种内在信号,衡量策略行动与回顾性解释之间的概率一致性。将SCR整合进强化学习,提供了密集的过渡级反馈,引导智能体朝向既有效又具物理基础的行为。跨越多种智能任务的大量实验表明,我们的方法相比仅前瞻性的世界建模基线,显著提升了策略的鲁棒性和泛化性。
T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning
T-路由器:通过参数高效强化学习学习丘脑路由进行推理
- Authors: Liuxian Ma, Jiale Dai, Jiaqi Li, Lu Mi
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
- Arxiv link: https://arxiv.org/abs/2609.39109
- Pdf link: https://arxiv.org/pdf/2609.39109
- Abstract
Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achieves 83.64 +/- 1.16 MathAvg after GSM8K RL, compared with 73.79 +/- 1.83 for full-parameter GRPO across three evaluation rounds. At a comparable parameter budget and with matched retries, it exceeds LoRA's 77.28 +/- 1.95 MathAvg, improving all three task families and raising mean AIME accuracy from 48.33 to 60.56. Capacity-controlled comparisons favor addressable block changes and recurrent context; separate search training extends the interface to tool-mediated reasoning. These results establish controlled computation reuse as an effective route to parameter-efficient reasoning reinforcement learning.
- 中文摘要
参数高效的强化学习旨在通过一个紧凑且可训练的预训练模型接口来提升推理能力。我们介绍丘脑路由器(T-Router),它专注于完成计算的复用。压缩的可寻址银行保留块变更;深度递归控制器对其选择和相对尺度写回进行条件。这种耦合为丘脑上下文相关路由提供了具体的计算形式:学习接收层使用哪些早期贡献,以及受到何种影响。正确性奖励接口训练,同时保持骨干参数和层序。在8.95B参数骨干上,T-Router分配了41.73M参数(占骨干的0.466%),GSM8K RL后实现83.64 +/- 1.16 MathAvg,而全参数GRPO在三轮评估中为73.79 +/- 1.83。在相当的参数预算和匹配重试下,T-Router超过LoRA的77.28 +/- 1.95 MathAvg,提升了三大任务族,并将平均AIME准确率从48.33提升至60.56。容量控制比较有利于可寻址块变更和循环上下文;独立搜索训练扩展了工具介导推理的接口。这些结果确立了受控计算重用作为参数高效推理强化学习的有效路径。
Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation
通过代币级感知基础优势估计强化多模态推理
- Authors: Zhihan Zhang, Lizi Liao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39168
- Pdf link: https://arxiv.org/pdf/2609.39168
- Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at this https URL.
- 中文摘要
带可验证奖励的强化学习(RLVR)提升了多模态大型语言模型(MLLMs)的推理能力,但现有框架依赖粗略的序列级奖励信号,缺乏对多模态推理链中视觉基础步骤的细致监督。我们通过两个代币级指标:视觉依赖性(即令牌预测对输入图像特征的依赖程度)和预测熵来探讨这一差距。我们的实证分析揭示了两个关键发现:(1)正确的推理链在视觉基础增强时表现出明显更明显的熵减少,相较于错误的推理链;(2)关键令牌,即那些因误判触发推理崩溃的标记,在视觉依赖与预测熵的联合分布中具有统计上异常值。基于这些发现,我们提出了代币层面感知基础优势估计(TPAE),通过测量每个代币与正确展开时视觉熵模式的统计一致性来估算代币层面优势。TPAE 利用这一细分调制序列层面优势,产生可集成到各种 RLVR 框架中的细粒度监督信号。对七个基准测试的广泛实验表明,TPAE 持续优于领先的强基线,从而实现了更稳定高效的多模推理优化。代码公开于此 https URL。
Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
将思想与答案对齐:概率奖励以驯服思维漂移
- Authors: Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.39183
- Pdf link: https://arxiv.org/pdf/2609.39183
- Abstract
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.
- 中文摘要
本文研究视觉语言模型中的\textbf{思维-答案一致性}。我们聚焦于视觉意图基础,模型基于人类意图查询推断目标对象并预测边界框。我们揭示,以往基于IoU的强化学习(RL)框架存在“思维漂移”问题,即模型尽管推理过程指向不同目标对象时仍能生成正确的边界框。因此,我们提出 \textbf{Rita}(\textit{Reinforcing Thinking--Answer consistency})作为一种新的强化学习范式来驯服漂移。具体来说,Rita 引入了两种无推理标签的强化学习奖励,基于参照答案的条件概率构建:\textbf{思考奖励}和\textbf{一致性奖励}。它还采用了一种有难度的\textbf{数据过滤}策略,利用推广错误率和奖励方差,选择有信息量的易中度样本进行强化学习。在EgoIntention和新的RefEgo-Int基准测试上的大量实验表明,Rita的表现始终优于监督式微调方法和普通的强化学习微调框架。
Learning Process Rewards via Reasoning State Propagation
通过推理状态传播获得学习过程奖励
- Authors: Kai Gan, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39220
- Pdf link: https://arxiv.org/pdf/2609.39220
- Abstract
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.
- 中文摘要
过程奖励模型(PRM)通过提供细粒度信号来评估中间推理状态,在测试时间的扩展和强化学习中表现出显著效果,但其训练高度依赖于昂贵的过程注释。缓解这种依赖的自然方法是用可扩展的结果监督补充有限的过程监督。然而,现有的PRM通常独立建模推理前缀,没有明确机制有效利用最终结果指导中间推理状态的学习。我们介绍推理状态传播(RSP),它表示每个推理前缀具有二元效度状态,并建模推理轨迹中连续状态之间的转换。具体来说,RSP预测有效状态变无效的断裂概率和无效状态恢复有效的修复概率。通过传播这些转移,RSP将中间状态与最终状态连接起来,使过程注释能够监督中间状态,而结果标签则监督最终状态,并为前一步提供学习信号。在推理搜索、响应选择和强化学习中,RSP始终优于代表性的PRM基线,在束搜索中平均优于Qwen2.5-Math-PRM为5.6%,在强化学习中均有2.1%的提升。
QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs
QATFactory:一个多功能、部署对齐的框架,用于量化感知的训练与LLMs提炼
- Authors: Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39223
- Pdf link: https://arxiv.org/pdf/2609.39223
- Abstract
Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and this http URL's Q4_K format; dense and mixture-of-experts models; and both full-parameter and LoRA-based training. It exports checkpoints directly to vLLM and this http URL without an additional lossy conversion step or added inference overhead. With QATFactory, we conduct extensive experiments on models ranging from 8B to 230B parameters and evaluate exported checkpoints in production inference engines. Across models and formats, QAD consistently improves deployed-model quality over strong PTQ baselines. On Qwen3.5-9B, QAD achieves average benchmark accuracies of 68.9% under NVFP4 and 66.0% under MXFP4, outperforming the best PTQ results of 65.4% and 56.4%, respectively. Through our experiments, we found that although both FP4 formats quantize weights and activations at deployment, the best training strategy is format-dependent: NVFP4 generally performs better when only weights are quantized during training, whereas MXFP4 benefits from quantizing both weights and activations. At a fixed training token budget, training on fewer 32K sequences improves average accuracy by 1.9 points over training on more 4K sequences. We release the complete QATFactory training code and the resulting checkpoints.
- 中文摘要
大型语言模型(LLM)推断正逐渐朝向更低精度的方向发展,以实现硬件加速器的吞吐量,但激进的训练后量化(PTQ)可能会降低模型质量。我们介绍 QATFactory,一个开源的部署对齐量化感知蒸馏(QAD)和强化学习(QARL)框架。QATFactory 在 BF16 中模拟部署时量化,同时执行矩阵乘法,使模型能够适应量化噪声,而无需原生支持目标格式的训练硬件;例如,它支持 H100 GPU 上的 NVFP4 训练,而 H100 GPU 缺乏 FP4 张量核心。该框架支持 NVFP4、MXFP4 及该 http URL 的 Q4_K 格式;密集模型和专家混合模型;以及全参数和基于 LoRA 的训练。它直接导出检查点到vLLM和该http URL,无需额外的有损转换步骤或额外的推理开销。通过QATFactory,我们对从8B到230B参数的模型进行了大量实验,并在生产推理引擎中评估导出的检查点。在不同模型和格式中,QAD在强PTQ基线下持续提升部署模型的质量。在Qwen3.5-9B中,QAD在NVFP4下平均基准准确率为68.9%,在MXFP4下达到66.0%,分别优于PTQ的最佳结果65.4%和56.4%。通过我们的实验,我们发现虽然两种FP4格式在部署时都会量化权重和激活,但最佳训练策略依赖于格式:NVFP4通常在训练时仅量化权重时表现更好,而MXFP4则通过同时量化权重和激活量化而受益。在固定的训练令牌预算下,使用较少的32K序列训练平均准确率比训练更多4K序列提升1.9个百分点。我们发布了完整的QATFactory训练代码及其产生的检查点。
Rethinking Multi-Image Re-Representation in Multi-Image Understanding
重新思考多图像理解中的多重图像再表征
- Authors: Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39363
- Pdf link: https://arxiv.org/pdf/2609.39363
- Abstract
Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at this https URL.
- 中文摘要
多图像理解不仅要求MLLM识别单个图像的内容,还需要组织分布在这些图像上的视觉证据。我们通过多图像再表征来研究该问题,将提示性思维链推理和智能视觉工具作为推理过程中重新组织视觉证据的不同方式。我们介绍了Mosaic,一种通用的多图像视觉工具,使MLLM能够主动构建具有十个可组合图像操作的视觉中间体。我们比较了现有多图像基准测试和MosaicBench(一种基于基础的新型精细多图像理解基准测试)上的五个再表征设置。我们的实验表明,文本和视觉再表征的相对优势高度依赖任务。视觉再表征对于需要精确视觉证据的任务尤其有效,包括假设检验、精度比较和方向敏感推理,而以高层语义内容为主的任务则获得较小或不一致的收益。基于这一发现,我们训练MosaicAgent-8B使用仅使用准确率和格式化奖励的强化学习Mosaic。没有演示轨迹或特定工具使用的奖励,智能体学会在多个步骤中构建视觉操作,并无提示地展现出多样化的问题解决模式。代码和数据将通过该 https URL 发布。
EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning
EHR-RobustGym:针对强健临床推理的基准与培训代理
- Authors: Yitong Qiao, Yancheng Jin, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39371
- Pdf link: https://arxiv.org/pdf/2609.39371
- Abstract
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.
- 中文摘要
在医院工作流程中,电子健康记录(EHR)常常存在噪声,且可能不包含确认临床查询中事件或测量所需的证据。即使数据库检索成功,临床代理也可能忽略这些差异,返回合理但缺乏支持的答案。我们推出了EHR-RobustGym,这是一个可扩展且互动的环境,用于评估和培训基于噪声EHR的稳健临床人员。EHR-RobustGym基于MIMIC-IV医院记录(36.5万患者,31个表格,超过5亿条记录),包含5,486对纯噪声对,涵盖六个临床意向以及患者层级和人群层级查询。这些匹配测试对记录级、价值级和查询级噪声的鲁棒性,同时交互式SQL/Python执行和结果验证支持轨迹收集和训练。评估多个LLM显示存在显著的稳健性差距:专有和大规模开放权重模型的平均任务成功率从Clean问题的62.2%降至噪声问题的37.9%。在k=4时,大多数评估模型的pass^k一致性低于50%,暴露出临床任务完成的不稳定性。EHR-RobustGym中的监督微调和强化学习提升了性能,这些进步可推广到五个外部EHR基准。这些结果共同使EHR-RobustGym成为评估和提升临床代理证据基础鲁棒性的测试平台。
SkillFM: Generating Skills for LLM Agents via Latent Flow Matching
SkillFM:通过潜在流匹配为LLM代理生成技能
- Authors: Zuming Zhang, Jie He, Yizhe Zhang, Jeff Z. Pan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39382
- Pdf link: https://arxiv.org/pdf/2609.39382
- Abstract
Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill Flow Matching), a generative framework that synthesizes task-conditioned textual skills directly without test-time skill retrieval. Our framework combines a codec for encoding and reconstructing textual skills in a continuous latent space with a conditional flow model trained using improved MeanFlow. At inference time, the learned velocity field enables single-step latent sampling, and an LLM-based decoder converts the sampled representation into textual guidance for a frozen downstream agent. We evaluate the framework on embodied tasks, question answering, and web shopping. On ALFWorld and Search-QA, our method achieves the best overall performance among the compared vector-based skill approaches. Our analyses further demonstrate that latent skill generation is an effective alternative to retrieval-based skill augmentation. Our code and training skill libraries are available at this https URL.
- 中文摘要
文本技能为大型语言模型代理提供了可复用的指导,但现有方法通常依赖手动策划的技能库或带有间接和延迟反馈的强化学习。我们引入了SkillFM(技能流匹配),这是一个生成框架,直接综合任务条件文本技能,无需测试时的技能检索。我们的框架结合了用于在连续潜在空间中编码和重构文本技能的编解码器,与使用改进的均流训练的条件流模型。在推理阶段,学习到的速度场支持单步潜在采样,基于LLM的解码器将采样表示转换为冻结下游代理的文本指导。我们评估该框架在具身任务、问答和网购方面。在ALFWorld和Search-QA上,我们的方法在所有基于向量的技能方法中整体表现最佳。我们的分析进一步表明,潜在技能生成是基于检索的技能增强的有效替代方案。我们的代码和训练技能库可在此 https 网址获取。
From Search to Signal: Online Post-Training in Automatic Heuristic Design
从搜索到信号:自动启发式设计在线后期培训
- Authors: Yilun Yuan, Tianyu Zhou, Zhenzhou Tang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39383
- Pdf link: https://arxiv.org/pdf/2609.39383
- Abstract
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.
- 中文摘要
基于大型语言模型(LLM)的自动启发式设计(AHD)迭代提出并优化启发式方法,将设计理由与可执行代码结合起来。任务特定的评估器评估程序;执行结果和性能评分指导搜索。许多AHD系统将生成器冻结;EvoTune和算法与语言模型共演化(CALM)则从被评估候选对象更新生成器。当这些结果驱动可验证奖励强化学习(RLVR)时,它们形成了一个搜索耦合循环:被评估的候选流同时提供搜索状态更新和训练信号,用于生成未来候选模型。然而,有效性和性能并不能唯一决定有用的模型更新;将其转化为学习信号时,必须考虑生成每个候选的提示和不断演变的搜索状态。我们将AHD中小型开权重LLM的在线后训练作为上下文依赖信号构建,并开发了从程序有效性、任务性能和生成上下文等替代映射以更新信号。通过共享评估的部署和匹配更新预算,跨AHD任务和模型家族的受控实验将这些映射与在线训练后基线进行比较,测试其对效度、有效提案表现及在上下文比较下提升的有效提案产出的影响。补充的检查点、冻结搜索和实时系统评估是否体现提案级提升出现在更新后的检查点行为及后续搜索中,而非仅仅源于累积搜索状态。在预设成本核算下进行资源匹配比较,检验在线更新是否在使用冻结生成器时额外搜索之外增加价值。该设计避免仅将端到端搜索收益视为更强启发式设计能力的证据。
No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation
《没有任务向量是孤岛:一项关于任务向量可组合性的综合研究》——通过策略提炼
- Authors: Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39405
- Pdf link: https://arxiv.org/pdf/2609.39405
- Abstract
Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced by on-policy distillation (OPD) remains largely unexplored. OPD trains a student using teacher feedback on student-generated trajectories, yielding parameter updates that differ from those produced by the teacher model, usually by reinforcement learning (RL). We therefore ask whether OPD task vectors can complement their RL teacher updates and compose effectively across tasks. Across five domains and two model architectures, we find evidence for both forms of composability. Within a task, merging OPD and RL task vectors can outperform both constituent models, even when the OPD student is weaker than its RL teacher. Across tasks, OPD task-vector compositions achieve higher average scores than corresponding RL compositions in seven of eight backbone-merging-rule comparisons. Parameter-space analyses reveal substantial non-collinearity between OPD and RL updates. Experiment in CODE domain on SMOLLM3-3B shows that the combined direction outperforms either constituent direction at the tested global update norm, supporting directional complementarity in this configuration. Across tasks, OPD updates also show lower overlap among the top-10% feed-forward channels ranked by update energy. Together, these results show that weaker standalone performance does not imply weaker task-vector composability. OPD task vectors can complement stronger RL teacher updates and combine effectively across tasks, highlighting composability as a distinct property for understanding and evaluating post-training updates.
- 中文摘要
任务向量为通过模型合并组合学习到的能力提供了一种简单的机制。然而,通过策略提炼(OPD)产生的任务向量的组合性仍然大多未被充分探索。OPD通过教师对学生生成轨迹的反馈来训练学生,从而产生与教师模型(通常通过强化学习(RL)产生的参数更新不同的参数更新。因此,我们探讨OPD任务向量是否能够补充其强化学习教师的更新,并在任务间有效组合。在五个领域和两种模型架构中,我们发现了两种可组合形式的证据。在一个任务中,即使OPD学生比强化学习老师弱,OPD任务向量的合并也能优于这两个组成模型。在各任务中,OPD任务向量组合在八次骨干合并规则比较中有七次平均得分高于对应的强化学习组合。参数空间分析显示,OPD与RL更新之间存在显著的非共线性。在SMOLLM3-3B的CODE域实验显示,综合方向在测试的全局更新规范下优于任一组成方向,支持该配置中的方向互补性。在任务中,OPD更新在按更新能量排名前10%的前馈通道间重叠更少。综合来看,独立性能较弱并不意味着任务向量组合性较弱。OPD任务向量可以补充更强的强化学习教师更新,并在任务间有效组合,凸显可组合性作为理解和评估训练后更新的独特特性。
From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL
从模仿到奖励发现:代理强化学习的政策预热
- Authors: Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39436
- Pdf link: https://arxiv.org/pdf/2609.39436
- Abstract
Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and higher final performance than the alternative baselines. Motivated by this observation, we study On-Policy Warmup (OPW), a teacher-guided stage in which the student trains with teacher supervision on its own interaction trajectories before transitioning to RLVR. Unlike imitation on fixed teacher-generated trajectories, OPW targets states induced by the student's own decisions, including imperfect actions and recovery situations. We provide a theoretical explanation by connecting on-policy reverse-KL distillation to trajectory-level distribution matching. Under a competent teacher and sufficiently small population distillation loss, this connection yields a lower bound on initial verifier success and a corresponding bound on reward-discovery complexity. For group-relative RLVR, we further characterize when increased success probability produces more reward-informative groups. Together, our findings support on-policy distillation as an effective warmup for agentic RLVR and identify initial reward discovery as a mechanism that can contribute to the observed acceleration.
- 中文摘要
带有可验证奖励的强化学习(RLVR)为训练语言模型代理提供了可扩展的方法,但结果奖励稀疏可能导致早期训练缺乏策略改进的信号。我们识别出一个策略启动加速现象:在主要比较中,采用策略提纯初始化的RLVR在训练早期达到高绩效,并在后续RLVR中获得更高的平均表现,最终表现优于其他基线。基于这一观察,我们研究了策略热身(OPW),这是一个教师引导阶段,学生在教师监督下训练自身的互动轨迹,然后再过渡到RLVR。与对固定教师生成轨迹的模仿不同,OPW针对由学生自身决策诱发的状态,包括不完美动作和恢复情境。我们通过将策略上的逆KL蒸馏与轨迹级分布匹配连接,提供了理论解释。在称职教师和足够小的群体蒸馏损失下,这种联系给出初始验证者成功的下界和奖励发现复杂度的下界。对于组相对RLVR,我们进一步表征了成功概率增加时产生更多奖励信息组。综合来看,我们的发现支持策略蒸馏作为能动RLVR的有效热身,并识别初始奖励发现作为促进观察加速的机制。
CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
CAST:因果优势结构化训练,空间基础的组合奖励,适用于扩散模型
- Authors: Shu Yu, Chaochao Lu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39441
- Pdf link: https://arxiv.org/pdf/2609.39441
- Abstract
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
- 中文摘要
在线强化学习已被扩展到扩散模型(DM)图像生成的流匹配。然而,该范式面临三个局限:(1)窗口选择。现有方法手动设置随机微分方程(SDE)采样窗口,即注入探索噪声的去噪步骤。我们则根据每个模型的去噪轨迹确定。(2)奖励饱和。当前方法依赖基于人工注释训练的评分模型;我们发现此类评分极高,在最新的SOTA开源DM中几乎无法区分,使优势估计效果较差。(3)样本效率低下。单一标量奖励将不同失败模式合并为几乎相同的评分,留下最小的梯度指导以实现目标改进。为解决这些问题,我们提出了CAST(因果优势结构化训练)方法,这是一种针对预训练DM的强化学习微调方法,该方法(1)识别每个模型固定对象及其空间排列的去噪步骤,并利用该时间设置SDE窗口;(2)通过因果场景图(CSG)将每个提示分解为可验证的原子,即最小语义单元如对象、计数、属性或空间关系,分别检查,并分别奖励每个原子;(3)通过教师强制关注将签名的原子级优势投射到像素空间,并用它们空间权重SDE策略目标。我们用CAST微调了两款最强的开源DM——FLUX.2-dev和Qwen-Image-2512,并在GenEval 2(合成基准测试)和Qwen-Image-Bench上评估它们的整体质量。在几乎相同的训练预算内,CAST在最具挑战性的GenEval 2提示上相较基础模型的提升是Flow-GRPO的3.07倍,整体生成质量也有所提升。
CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
注意点:一个可控分析测试平台用于编程强化学习中的奖励黑客
- Authors: Shouli Wang, Yanfeng Jia, Zhihao Ou, Zitao Su, Ruize He, Haotong Xie, Hao Peng, Juanzi Li, Xiaozhi Wang
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39533
- Pdf link: https://arxiv.org/pdf/2609.39533
- Abstract
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We introduce CATCH, a controllable testbed for studying reward hacking in coding RL. CATCH deliberately exposes environmental loopholes and provides execution-based gold labels by comparing success under a vulnerable evaluator with task correctness under an independent audit. It also can control the model's initial hacking tendency through supervised fine-tuning data mixtures and the difficulty of earning rewards through reward designing, enabling systematic comparisons of hacking dynamics and interventions. Experiments show that CATCH can produce diverse RL training trajectories with clear reward hacking, and analyses demonstrate that both initial models and reward difficulties shape the emergence of reward hacking. We further evaluate the effectiveness of different reward hacking detection and mitigation methods. A key finding is that a chain-of-thought monitor initially suppresses hacking, but this protection erodes as the policy model learn to mislead the monitor with code comments. This highlights the need to evaluate hacking mitigations throughout training with CATCH. The source code and resources are publicly released at this https URL.
- 中文摘要
在可验证奖励强化学习(RLVR)中,大型语言模型(LLM)可以利用环境中的漏洞,在不提升预期能力的情况下获得高奖励,即奖励黑客。尽管存在训练效率和安全风险,但在培训期间监控和缓解奖励黑客仍然具有挑战性,因为缺乏能够重现并可靠识别黑客行为的测试平台。我们介绍了CATCH,一个可控测试平台,用于研究编程强化学习中的奖励黑客。CATCH有意暴露环境漏洞,并通过比较脆弱评估者的成功率与独立审计下的任务正确性,提供基于执行的黄金标签。它还可以通过监督微调数据混合控制模型的初始黑客倾向,并通过奖励设计降低获得奖励的难度,从而实现黑客动态与干预的系统比较。实验表明,CATCH能够通过明确的奖励黑客生成多样化的强化学习训练轨迹,分析也表明初始模型和奖励困难共同塑造了奖励黑客的出现。我们还进一步评估了不同奖励黑客检测和缓解方法的有效性。一个关键发现是,思维链监控器最初能抑制黑客行为,但随着策略模型学会用代码注释误导监视器,这种保护逐渐减弱。这凸显了在CATCH训练过程中评估黑客缓解措施的必要性。源代码和资源已在此HTTPS网址公开发布。
From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning
从既有到收集证据:纵向医学推理中的能动学习
- Authors: Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.39566
- Pdf link: https://arxiv.org/pdf/2609.39566
- Abstract
Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evidence, together with a tool-use harness and an agentic post-training framework for compact vision-language policy models. We further introduce a longitudinal multimodal benchmark built on UK Biobank, comprising 50,401 clinical questions derived from real-world ICD-10-coded diagnoses of 4,739 participants. Each question links to a patient-specific environment containing clinical context and multi-sequence MRI from baseline and follow-up visits, where agents autonomously select which visits, organs, modalities, slices, and specialist tools to inspect and compare. Supervised fine-tuning transfers evidence-seeking workflows from 14,734 frontier-model interaction trajectories, followed by agentic reinforcement learning on the learner's own environment interactions. Privileged on-policy self-distillation and rubric-based LLM feedback refine evidence-to-conclusion reasoning without prescribing tool sequences. Experiments show that CASE moves beyond question-answer imitation toward transferable investigation policies, strengthening evidence-grounded longitudinal reasoning. Under matched evaluation conditions, our Qwen3-VL-8B based agent achieves over 16% and 10% relative improvements in answer accuracy over GPT-5.4 and Claude Opus 4.8. Code will be available at this https URL.
- 中文摘要
基础模型可通过工具使用工具作为临床代理。然而,传统医学基准评估的是对预选证据的推理能力,而非跨临床记录和纵向影像检索的能力。我们提出CASE:一系列针对特定角色的临床证据寻求代理,以及工具使用工具框架和针对紧凑视觉语言政策模型的代理后培训框架。我们还进一步引入基于英国生物样本库的纵向多模态基准,包含50,401个临床问题,源自4,739名参与者的真实ICD-10编码诊断。每个问题关联患者特定环境,包含临床背景和基础及随访多序列MRI,代理自主选择进行检查和比较的访问、器官、模态、切片和专科工具。监督微调将证据寻求工作流从14,734条前沿模型交互轨迹转移,随后对学习者自身环境互动进行智能体强化学习。特权策略自我蒸馏和基于评分标准的LLM反馈优化了证据到结论的推理,无需规定工具序列。实验显示,CASE超越了问答模仿,迈向可转移的调查策略,强化了基于证据的纵向推理。在匹配评估条件下,我们的基于Qwen3-VL-8B的智能体在回答准确度上相较GPT-5.4和Claude Opus 4.8分别提升了16%和10%。代码可在此 https 网址获取。
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
跳出框架思考:语言模型能否选择性依赖外部指导?
- Authors: Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao, Fang kong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.39578
- Pdf link: https://arxiv.org/pdf/2609.39578
- Abstract
Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
- 中文摘要
代理工具通常能通过人工设计的工作流程改进语言模型,但随着模型能力提升,不可靠的指导会越来越限制其执行。我们称之为在利用有用指导的同时,超越不可靠引导思维的能力。我们引入了Box$^2$-Bench,它保持模型和任务固定,同时调整工作流程可靠性,以隔离模型如何调节对指导的依赖。在Box$^2$-Bench上,前沿模型通常受益于可靠的指导,但当引导误导或变得不可靠时仍存在脆弱性。为测试这种能力是否可被学习,我们训练了两个开放权重模型,使用不良工作流程,将好的工作流程保留用于评估。我们探讨了两种互补训练策略:反事实监督微调提升了鲁棒性,而基于结果的强化学习则可以将平衡转向更广泛使用有用的工作流。我们还发现,这种行为不仅限于工作流程,还能接收其他形式的外部信息,提升了同伴纠正和对损坏记忆的鲁棒性。我们的结果共同指出,选择性依赖可出错的外部信息是代理可靠性的一个维度,而不仅仅是任务表现所能体现的。
GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
GroundingPI:面向具视觉基础的物理智能基础模型
- Authors: Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.39601
- Pdf link: https://arxiv.org/pdf/2609.39601
- Abstract
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
- 中文摘要
精确的接地非常重要。它指定目标对象及其位置,即使是在杂乱和微小物体中,且必须足够快以实现闭环控制。然而,视觉-语言-动作(VLA)和世界动作模型(WAMs)依赖通用视觉语言和视频生成骨干的感知,这些在这些环境中仍然失效。我们介绍GroundingPI,一种4B基础模型,在共享词汇中生成点和框作为量化坐标。训练结合多模态和空间预训练、监督微调和强化学习,使用GRPO,使用公共数据集和专用数据引擎的监督。在涵盖11种感知能力的34个接地基准中,GroundingPI建立了新的技术水平,平均达到73.68%,高于更大的GPT-6 Astra(71.54%)。作为下游视觉骨干,GroundingPI提升了机器人操作和自动驾驶的性能。在RoboTwin 2.0平台上,它在所有四个非分布环境中的主流骨干网表现均优于我们评估的所有主流骨干网,相较于最强骨干网高达24.8%。在RoboCasa-GR1上,GroundingPI使用50%的演示训练后,表现优于训练的基线75%。作为视觉骨干的nuScenes上,GroundingPI的平均开环L2误差达到0.296米。我们系统地分析了GroundingPI在规模和数据组成上的预训练。随着预训练的扩展,下游自动驾驶和机器人操作能力不断提升。分析这11种感知能力的数据配方显示,密集接地对两者都带来了显著益处,OCR也具有作为感知学习催化剂的潜力。这些结果支持将基础化作为感知基础,以及专门的感知预训练作为物理智能基础模型的有前景方向。
ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
ChronoGraph:功能性四维场景图,采用视觉语言模型,用于交互理解和扎根规划
- Authors: Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39665
- Pdf link: https://arxiv.org/pdf/2609.39665
- Abstract
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
- 中文摘要
具身的代理必须确定行动地点,预测场景变化,并解读观察结果以指导后续行动。这需要将4D交互理解(解释过去行为如何改变场景)与空间基础规划结合起来,规划决定如何以及在哪里行动,并预判场景变化。我们介绍ChronoGraph,一个功能性4D场景图,将可感知部分的动作与语义和几何状态变化联系起来。通过以同一形式表示观察到的和预期的转变,它为理解和规划提供了共享基础。我们通过自动数据引擎构建ChronoGraphBench,将人机交互视频和模拟机器人轨迹转换为图形注释问题,用于训练和评估视觉语言模型(VLM)。利用这些注释,我们将预训练VLM分两阶段进行训练。图即思维链监督微调教导模型重建观察到的转变,并以图迹形式预测未来转移,然后再回答。随后的联合四维图强化学习直接奖励图属性和答案正确性。跨模型尺度的实验显示,相较于对应的预训练基线和零样本转移到VLM4D的改进。实际演示进一步表明,基于图的规划和可供性基础支持通过现有机器人技能实现移动操作,无需额外微调。
Validity-Preserving Hierarchical RL for Joint Routing and Switch Placement in EDA
EDA中用于联合路由和交换机布置的保效性层级强化学习
- Authors: Dorian Gailhard, Ugo Lecerf, Enzo Tartaglione, Donatello Conte, Jhony H. Giraldo
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39749
- Pdf link: https://arxiv.org/pdf/2609.39749
- Abstract
Routing and switch placement are fundamental combinatorial optimization problems in chip design, requiring the joint optimization of routing topology and physical placement under strict structural, geometric and logical constraints. Existing approaches typically rely on carefully engineered heuristics that incorporate strong problem-specific biases to navigate the enormous space of possible designs. In this work, we introduce a hierarchical reinforcement learning framework for joint routing and switch placement at the level of logical communication routes. Starting from a minimal routing graph, our method progressively constructs increasingly expressive solutions through three coupled operations: switch expansion, switch placement, and route refinement. These operations preserve routing validity by construction, restricting exploration to feasible configurations where every communicating initiator-target pair has one assigned loop-free route. We explore the induced solution space using Gumbel Monte Carlo Tree Search, showing that neural-guided search substantially improves solution quality over non-learning optimization methods. Furthermore, pretraining across floorplans provides a strong initialization for fine-tuning on unseen instances.
- 中文摘要
路由和交换机布置是芯片设计中基本的组合优化问题,要求在严格的结构、几何和逻辑约束下,实现路由拓扑和物理布置的联合优化。现有方法通常依赖精心设计的启发式方法,结合强烈的问题特有偏差,以应对庞大的可能设计空间。本研究介绍了一个层级强化学习框架,用于逻辑通信路由层面的联合路由和交换机布置。从最小路由图出发,我们的方法通过三项耦合操作逐步构建更具表现力的解:交换机展开、交换机布置和路由细化。这些操作通过构造保持路由有效性,限制探索在可行配置中,即每个通信的发起者-目标对都有一条指定的无环路由。我们利用甘贝尔蒙特卡洛树搜索探索诱导解空间,表明神经引导搜索相比非学习优化方法显著提升了解的质量。此外,跨户型的预训练为未见实例的微调提供了强有力的初始化。
Riemannian Flow Models with Reinforcement Learning for Molecular Crystal Structure Prediction
黎曼流模型与强化学习用于分子晶体结构预测
- Authors: Thomas Egg, Harry Winston Sullivan, Maya M. Martirossyan, Philipp Höllmer, Cheng Zeng, Adrian Roitberg, Mingjie Liu, Richard Hennig, Sapna Sarupria, Ellad B. Tadmor, Stefano Martiniani
- Subjects: Subjects:
Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
- Arxiv link: https://arxiv.org/abs/2609.39773
- Pdf link: https://arxiv.org/pdf/2609.39773
- Abstract
Crystal structure governs material properties, making crystal structure prediction (CSP) a fundamental problem in materials science. Generative models are a promising approach for solving this problem, but the prevalence of polymorphism, coupled with large unit cells and complex packing geometry, makes the molecular CSP task challenging for existing models. To address this, we introduce Coarse-Grained Open Materials Generation (CG-OMatG), an equivariant Riemannian flow-based generative model. CG-OMatG predicts molecular crystal structures \textit{via} a coarse-grained, hierarchical representation. CG-OMatG treats molecules as rigid bodies---performing both inter- and intra-molecular message passing to construct a geometric representation for molecular packings---and learns to reconstruct molecule centroid positions, orientations, and lattice parameters, conditioned on chemical species and conformer geometry. We train the model on subsets of the Open Molecular Crystals (OMC25) and Cambridge Structural Database (CSD) datasets. Further, we fine-tune the model \textit{via} policy gradient reinforcement learning to steer the model towards generating low-energy candidate structures. We validate the generated structures on the CSP blind test benchmark, assessing agreement with experimentally determined crystals using COMPACK packing-similarity analysis. CG-OMatG exhibits strong performance for generative molecular crystal structure prediction, paving the way for accelerated polymorph screening and organic solid-state materials discovery.
- 中文摘要
晶体结构决定材料性质,使晶体结构预测(CSP)成为材料科学中的基础性问题。生成模型是解决该问题的有前景方法,但多态性普遍存在,加上大晶胞和复杂的包装几何,使得现有模型在分子CSP任务中具有挑战性。为此,我们引入了粗粒开材料生成(CG-OMatG),这是一种基于黎曼流的等变生成模型。CG- OMatG 预测分子晶体结构,这是一种粗粒度的层级表示。CG-OMatG 将分子视为刚体---执行分子间和分子内的信息传递,构建分子包装的几何表示---并学习根据化学种类和共融几何结构重建分子质心的位置、取向和晶格参数。我们用开放分子晶体(OMC25)和剑桥结构数据库(CSD)数据集的子集训练模型。此外,我们通过策略梯度强化学习微调模型 \textit{via},引导模型生成低能候选结构。我们在CSP盲测基准上验证生成结构,利用COMPACK包装相似性分析评估与实验确定晶体的一致性。CG-OMatG在生成分子晶体结构预测方面表现出优异表现,为加速多晶型筛选和有机固态材料发现铺平了道路。
Darpan: A Digital Twin Framework for the Next-Generation Computing Continuum
Darpan:下一代计算连续体的数字孪生框架
- Authors: Zhiyu Wang, Rajkumar Buyya
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2609.39799
- Pdf link: https://arxiv.org/pdf/2609.39799
- Abstract
Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, network conditions, and application progress may change, so the decision can be invalid by the time it is executed. Existing runtimes enact predetermined decisions, whereas simulation tools compare alternatives in preconfigured environments; neither explores alternative outcomes directly from the current state of a running application. We propose a Digital Twin framework, called Darpan, that supports both physical and digital execution: the physical side runs real applications and continuously observes their runtime state, while the digital side maintains a continuously updated, executable virtual counterpart based on these physical observations. Darpan evaluates candidate decisions independently from the same starting point and returns the selected decision to the physical side, where its feasibility is validated against the latest physical state before actual execution. Across real Directed Acyclic Graph (DAG) workloads, Darpan predicts physical response time with a mean absolute error of 0.485 s and retains 93.1% scale-out efficiency at 40 nodes. In the state-change trials, Darpan takes about 2 ms on average from state capture to rejection of an invalidated decision, stopping the request before data transfer and avoiding unnecessary physical execution overhead. Under the same physical-DAG budget, Darpan-generated experience enables a weaker Proximal Policy Optimization (PPO) scheduler to outperform two state-of-the-art Deep Reinforcement Learning (DRL) schedulers by 29.1% and 20.2%.
- 中文摘要
计算连续体应用将工作分配到设备、边缘系统、雾资源和云端。在做出放置、调度或恢复决策时,资源可用性、网络状况和应用进度可能发生变化,因此执行时决策可能无效。现有运行时执行预定决策,而仿真工具则比较预配置环境中的替代方案;两者都不直接从当前运行应用状态中探索替代结果。我们提出了一个数字孪生框架,称为Darpan,支持物理和数字执行:物理端运行真实应用并持续观察其运行状态,数字端则基于这些物理观察维护一个持续更新的可执行虚拟对应物。Darpan从相同起点独立评估候选决策,并将所选决策返回物理端,在实际执行前根据最新物理状态验证其可行性。在真实的有向无环图(DAG)工作负载中,Darpan 预测物理响应时间,平均绝对误差为 0.485 秒,并在 40 个节点上保持 93.1% 的扩展效率。在状态变更试验中,Darpan 从状态捕获到拒绝无效决策平均约 2 毫秒,能在数据传输前停止请求,避免不必要的物理执行开销。在相同的物理 DAG 预算下,Darpan 生成的体验使较弱的近端策略优化(PPO)调度器表现优于两个最先进的深度强化学习(DRL)调度器 29.1% 和 20.2%。
KUAISHOU Explorer LLM-Rec Challenge 2026: Reasoning Generative Recommendation
快手探索者LLM-Rec挑战2026:推理生成推荐
- Authors: Jiangxia Cao, Hao Peng, Wenlong Xu, Jiaxin Deng, Zhixin Ling, Xingmei Wang, Kun Shang, Can Tang, Zhihuai Cai, Jun Du, Fang Su, Xiaojuan Liu, Yiling Li, Chenglong Yu, Chongling Rao, Haixuan Gao, Haitao Xu, Jian Liang, Ruiming Tang, Chenglong Chu, Guohong Mu, Honghui Bao, Hui Wang, Jialong Chen, Jiao Ou, Muhao Wei, Peng Zhang, Renpu Liu, Ruochen Yang, Shugui Liu, Xinqi Jin, Yan Sun, Yifan Wang, Yingzhi He, Yufei Ye, Yusen Huo, Tingkuo Wang, Jihong Zhang, Lanxi Zhu, Pengyuan Liu, Zhipeng Yi, Luankang Zhang, Hang Lv, Xuyang Zhi, Tianyu Li, Bintao Wu, Chuang Ou, Siyue Su, Ziyuan Wang, Yuliang Sun, Baiyan Che, Feiyang Xu, Shiwen Zhang, Shiteng Cao, Chongcong Jiang, Yuan Fang, Xiangwu Yang, Hao Deng, Zijian Du, Pengxun Wang, Xiaoming Wang, Shun Qin, Yingqi Song, Tianyi Li, Naixiao Peng, Chenyu Zhou, Qiliang Jiang, Quan Zheng, Cheng Jin, Siying Zeng, Hongjia Xu, Junwu Hu, Teng Fu, Zhengkang Mei, Haijun Yu, Kai Li, Shengyang Zhou, Zhijia Wei, Siyi Xiong, Bo Liu, Zichun Guo, Zhubin Han, Jinpeng Fu, Bingqian Liu, Yuyi Wang, Yu Liu, Qinghai Tan, Ruijie Zhou, Zhuohang Li, Zhijia Zhong, Xiangnan He, Jirong Wen, Min Zhang, Wenwu Ou, Peng Jiang, Han Li, Kaiqiao Zhan, Yanan Niu, Lantao Hu, Kun Gai
- Subjects: Subjects:
Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2609.39828
- Pdf link: https://arxiv.org/pdf/2609.39828
- Abstract
Generative recommendation, has been attracted a surge of attentions in industrial and academic research community, towards to build more smart system to build next-generation recommender. Under the significant developing wave of large language model, our team have been developed Semantic ID based OneRec/OneRec-V2. These models have been widely deployed in production and demonstrate the scaling potential of the autoregressive next-item prediction paradigm for industrial recommender systems. Building on the success of OneRec, we further explored a series of models, including OneRec-Think, OpenOneRec, and OneReason, that connect item Semantic IDs with natural language in a unified representation space and seek to unlock the potential of natural-language chain-of-thought (CoT) reasoning for recommendation. However, our preliminary works found that introducing reasoning CoT does not always improve the recommendation performance. To address this issue, OneReason strengthens the semantic alignment between items and language, introduces structured template-based supervision for interest reasoning, and applies advanced reinforcement learning techniques to make reasoning more beneficial to recommendation. As a frontier topic to building recommendation foundation models, we believe this topic has significant research value and hope to encourage more researchers to explore it together. To this end, together with the SIGIR 2026 community, we organized the KUAISHOU Explorer LLM-Rec Challenge 2026: Reasoning Generative Recommendation.
- 中文摘要
生成式推荐在工业和学术研究界引起了大量关注,致力于构建更智能的系统以构建下一代推荐器。在大型语言模型的重大发展浪潮下,我们的团队开发了基于语义ID的OneRec/OneRec-V2。这些模型已广泛应用于生产环境,展示了自回归下一项目预测范式在工业推荐系统中的扩展潜力。基于OneRec的成功,我们进一步探索了一系列模型,包括OneRec-Think、OpenOneRec和OneReason,这些模型将项目语义ID与自然语言连接在统一的表示空间中,旨在释放自然语言思维链(CoT)推理在推荐中的潜力。然而,我们的初步工作发现,引入推理CoT并不总能提升推荐性能。为解决这一问题,OneReason 加强了条目与语言之间的语义对齐,引入了基于模板的结构化兴趣推理监督,并应用高级强化学习技术,使推理更有利于推荐。作为构建推荐基础模型的前沿话题,我们认为该课题具有重要的研究价值,并希望鼓励更多研究者共同探索。为此,我们与 SIGIR 2026 社区共同组织了 KUAISHOU 探索者 LLM-Rec 挑战赛 2026:推理生成推荐。
Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
学习隐写很容易,但学习数字写法推理很难
- Authors: Julian Schulz, Lukas Fülle, Rieke Fruengel
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39838
- Pdf link: https://arxiv.org/pdf/2609.39838
- Abstract
Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded reasoning (reasoning in an illegible but unconcealed format), have already been shown to emerge under training pressures that occur in real pipelines, such as reinforcement learning against monitors. This suggests that steganographic reasoning too might arise as an unintended side effect of training. Here, we compare how easily models learn steganographic reasoning and these two neighbouring capabilities across three elicitation methods: reinforcement learning, in-context learning, and supervised fine-tuning (SFT). For most tasks, models learn steganographic reasoning only under SFT, while they learn steganographic messaging and encoded reasoning under all three elicitation methods. Even under SFT, steganographic reasoning requires at least twice as much training as messaging, and for several model-task combinations it is not learned at all. However, on a cover task that makes hiding information especially convenient, steganographic reasoning can be successfully learned under all three elicitation methods. Steganographic reasoning is thus much harder than steganographic messaging and encoded reasoning, and learning the latter two does not imply learning the former. Yet it lies within reach: an easy version is learned under every elicitation method, when the cover task is convenient for hiding information.
- 中文摘要
作为人工智能监督和控制方法的思维链监控,正受到隐写推理的可能性威胁,即大型语言模型将推理隐藏在看似无害的文本中。两个相邻能力——隐写消息(传递隐蔽消息)和编码推理(以不可辨识但未隐藏格式推理)已被证明在真实管道中训练压力下出现,例如对监视者的强化学习。这表明隐写推理也可能作为训练的意外副作用出现。本文比较模型在三种引发方法——强化学习、上下文学习和监督微调(SFT)下学习隐写推理及这两种相邻能力的难易程度。对于大多数任务,模型仅在SFT下学习隐写推理,而在三种引发方法下学习隐写信息和编码推理。即使在SFT下,隐写推理所需的训练量至少是消息的两倍,且在某些模型-任务组合中根本没有学习。然而,在一个特别方便隐藏信息的掩体任务中,隐写推理可以通过三种引出方法成功学习。因此,刻写推理比隐写消息和编码推理更难,学习后两者并不意味着必须学会前者。但它触手可及:当掩盖任务适合隐藏信息时,每种引出方法都会学会一个简单的版本。
RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments
用于未知环境中基于概率安全的基于感知的RL导引PAC-NMPC
- Authors: Adam Polevoy, Dillon Capalongo, Katherine Tang, Mark Gonzales, Marin Kobilarov, Joseph Moore
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.39854
- Pdf link: https://arxiv.org/pdf/2609.39854
- Abstract
In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models. We then leverage these probabilistic models in a sampling-based SNMPC framework known as Probably Approximately Correct (PAC)-NMPC, which uses hard constraints to enforce finite-time statistical guarantees on the probability of collision and value function improvement. By ensuring that our finite-horizon SNMPC policies decrease the value function in expectation, we can approach the long-horizon performance of the RL approach while satisfying probabilistic safety constraints. Through simulation experiments, we show that our approach can improve the safety of perception-based RL navigation policies and scale to high dimensional systems with large sensor input spaces and complex nonlinear dynamics. We also demonstrate our approach through hardware experiments, showing improved performance for vision-based navigation with an agile fixed-wing aerial vehicle in unknown environments.
- 中文摘要
本文提出了一种结合随机非线性模型预测控制(SNMPC)和强化学习(RL)的方法,以实现在未知环境中基于概率安全的感知导航。我们的方法首先使用强化学习来训练概率性行为者-批判者预测模型和传感器预测模型。随后,我们将这些概率模型应用于一个基于抽样的SNMPC框架,称为Probably Approximately Correct(PAC)-NMPC,该框架利用硬约束强制执行碰撞概率和价值函数改进的有限时间统计保证。通过确保有限视野SNMPC策略降低期望值函数,我们可以在满足概率安全约束的同时接近强化学习方法的长视野性能。通过仿真实验,我们证明了我们的方法能够提升基于感知的强化学习导航策略的安全性,并可扩展到具有大传感器输入空间和复杂非线性动力学的高维系统。我们还通过硬件实验展示了我们的方法,显示出在未知环境中灵活固定翼飞行器在视觉导航方面的性能提升。
Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding
多线总线保持的完成感知跨保真度离线到在线强化学习
- Authors: Yifan Zhang, Qifan Zhang, Liang Zheng
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39868
- Pdf link: https://arxiv.org/pdf/2609.39868
- Abstract
Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.
- 中文摘要
在运营中的公交车队上进行探索性强化学习(RL)不切实际,而仅基于历史数据训练的策略无法获得新经验。混合离线与在线(H2O)强化学习结合了固定目标重放与模拟器交互,但廉价的在线模拟器在转换和事件时长动态上可能与目标存在差异。我们研究了多线路公交车等待的跨保真度问题,解决了较低的广义乘客时间与不完整的乘客旅程共存的失败模式。
GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
GrammarRL:通过强化学习实现的有效语法约束解码
- Authors: Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov`ı
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39869
- Pdf link: https://arxiv.org/pdf/2609.39869
- Abstract
Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with beam width, while sequence-level probability is only an imperfect proxy for semantic quality. We introduce GrammarRL, a label-free reinforcement learning method that adapts language models to grammar constraints without requiring annotated data. GrammarRL optimizes the model using two complementary self-supervised rewards derived from its own likelihoods: a direct reward, measuring how likely the constrained output is given the input, and a reverse reward, measuring how well the input can be reconstructed from the generated output. We optimize these rewards with a Reinforce Leave-One-Out (RLOO) objective over groups of grammar-constrained rollouts, augmented with the top-1 beam-search hypothesis and regularized towards a frozen base model. We evaluate GrammarRL on sign language gloss translation, hierarchical text classification, and named entity recognition using Llama models ranging from 1B to 8B parameters. GrammarRL consistently outperforms constrained greedy decoding, with an average improvement of 9.8 points and gains of up to 22.8 BLEU. It matches or outperforms beam search on two of the three tasks while preserving greedy-decoding inference cost. Ablations further show that the two rewards are complementary: either reward alone can underperform the untrained baseline, whereas their combination consistently improves upon it.
- 中文摘要
语法约束生成保证语法有效性,但当模型偏好输出与强加语法不匹配时,语义质量会大幅下降。当提示词指定不足或模型指令跟随能力有限时,这种权衡尤为严重。束搜索可以通过探索多个有效序列部分缓解这些失败,但其计算成本随束宽增加增加,而序列级概率仅是语义质量的一个不完美代理。我们介绍GrammarRL,一种无标签强化学习方法,能够在不要求注释数据的情况下调整语言模型以适应语法约束。GrammarRL通过两种互补的自监督奖励优化模型:直接奖励,测量受约束输出获得输入的可能性,以及反向奖励,衡量输入从生成输出中重建的概率。我们通过强化留一(RLOO)目标,针对语法约束的展开组优化这些奖励,辅以顶1束搜索假设,并正则化至冻结基模型。我们利用1B至8B参数范围的Llama模型,评估GrammarRL在手语词汇翻译、分层文本分类和命名实体识别方面的表现。GrammarRL持续优于约束贪婪解码,平均提升9.8分,提升最高22.8 BLEU。它在三项任务中的两个上表现优于或超过束搜索,同时保持贪婪解码推断成本。消融进一步表明两种奖励是互补的:单独奖励可能低于未训练基线,而两者组合则持续优于基线。
Grounding with Confidence: Controllable Generative Video Temporal Grounding
自信接地:可控生成视频时间接地
- Authors: Jinhao Chen, Benlei Cui, Ruijian Jia, Ziheng Wang, Tianyu Wo, Pengfei Sun, Longtao Huang, Hui Xue, Yitong Yang, Haiwen Hong
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.39883
- Pdf link: https://arxiv.org/pdf/2609.39883
- Abstract
Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro [email protected] from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.
- 中文摘要
视频时间基础支持视频搜索、内容审查和通过本地化自然语言描述事件的自动编辑等应用。然而,现有生成模型通常输出时间戳,但没有明确的区间级置信分数,以指导候选人选择。我们通过对原始译码内的单个区间进行评分,将候选生成与接受分离。轻量级置信头读取合并的解码器状态,提供为区间选择训练的显式评分。离线验证器评分监督固定候选序列的头值,时间重叠标签则在强化学习期间调整其以适应当前的推广。GT锚定的候选人池监督和集合级优化训练生成器。所得分数支持排名、基于阈值的选择和拒绝,而无需在推断时调用外部验证器。在固定的OMTG-Bench候选人池中,在10%全局回报预算下,查询宏[email protected]的置信度从9.95%提升至14.42%,在25%预算下从26.48%提升至31.12%。连续评分允许下游应用调整返回预算或接受阈值以匹配其精确回忆偏好,而无需重新生成候选区间。
Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?
更好的目标表征是否能改善目标条件强化学习?
- Authors: Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39901
- Pdf link: https://arxiv.org/pdf/2609.39901
- Abstract
Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline GCRL by constructing an exact temporal-distance goal representation in deterministic mazes. We then systematically corrupt its geometric quality while keeping the downstream learner fixed. Across OGBench navigation tasks and two algorithms, large changes in goal-representation quality produce almost no change in performance. However, applying the same interventions to the agent's current state more than doubles success, revealing the state pathway as the true bottleneck. Building on this insight, we show that simple random Fourier positional encodings substantially improve performance on the hardest navigation tasks without map information or objective modifications. Overall, our findings suggest that in state-based offline navigation, improving how the agent's current state is represented matters far more than refining the goal representation. Code will be released soon.
- 中文摘要
目标条件强化学习(GCRL)高度依赖于目标目标如何被策略表示。虽然近期方法通过时间距离、占用性或可控性来编码目标,但下游性能在多大程度上取决于表示质量,仍不明确。我们在离线GCRL中通过确定性迷宫构建精确的时间-距离目标表示来研究这一点。然后系统地破坏其几何特性,同时保持下游学习器固定。在OGBench导航任务和两种算法中,目标表示质量的大幅变化几乎不会导致性能变化。然而,将同样的干预应用于代理当前状态,成功率可翻倍多,揭示状态路径才是真正的瓶颈。基于此见解,我们证明简单的随机傅里叶位置编码在最困难的导航任务中显著提升性能,且无需地图信息或目标修改。总体来看,我们的发现表明,在基于状态的离线导航中,改进代理当前状态的表示方式远比优化目标表示更为重要。代码将很快发布。
Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling
慢性多病患者的以患者为中心治疗计划:偏好建模的层级强化学习框架
- Authors: Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.39911
- Pdf link: https://arxiv.org/pdf/2609.39911
- Abstract
Patient preference, defined as a patient's demonstrated willingness and capacity to adhere to clinical recommendations, is a primary determinant of therapeutic effect yet remains structurally absent from existing computational treatment planning models. We address this gap by presenting patient-centered factored-action hierarchical option-critic (FAHOC), a hierarchical reinforcement learning (HRL) framework that jointly learns high-level options corresponding to therapeutic strategies and factored intra-option policies that decompose the joint action space into disease- and intervention-specific subcomponents, while imposing a cooperation-aware action masking mechanism. This enables structured exploration, improved credit assignment across hierarchy levels, and more interpretable decision pathways, while enforcing patients' preferences. Formal guarantees establish that cooperative patients achieve higher optimal expected health outcomes than non-cooperative patients, and that the factored Q-function approximation error is provably bounded. The framework is evaluated using longitudinal data collected from approximately 50,000 comorbid hypertension and type 2 diabetes mellitus patients from five hospitals in the Southeast U.S. FAHOC achieves a quality-adjusted life year expectancy equivalent improvement of 0.669 (vs -0.133 observed clinician practice), correctly identifies cooperative patients in 95.9% of cases and never violates a patient's preference in held-out test, demonstrating that HRL with explicit preference constraints can support preference-consistent, clinically safe decision-making in multimorbidity management.
- 中文摘要
患者偏好定义为患者表现出遵守临床建议的意愿和能力,是治疗效果的主要决定因素,但在现有计算治疗计划模型中结构性地缺失。我们通过提出以患者为中心的因式分解行动层级选项批判(FAHOC)来弥补这一空白,这是一种层级强化学习(HRL)框架,该框架共同学习对应治疗策略的高层次选项,并将联合行动空间分解为疾病和干预特异的子组成部分,同时施加合作意识的行动掩蔽机制。这支持结构化探索、层级间的信用分配改善和更可解释的决策路径,同时强制执行患者的偏好。形式保证证明合作患者比非合作患者获得更高的最佳预期健康结果,且分解后的Q函数近似误差是可证实的有界的。该框架利用来自美国东南部五家医院约5万名共病高血压和2型糖尿病患者的纵向数据进行评估。FAHOC在质量调整后预期寿命改善率为0.669(相比临床医生实际观察为-0.133),在95.9%的病例中正确识别合作患者,且在未完成测试中从未违反患者的偏好,证明带有明确偏好约束的HRL能够支持偏好一致、临床安全的多病管理决策。
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
不可学,还是无法测量?关于RLVR难度标签的可靠性
- Authors: Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.40115
- Pdf link: https://arxiv.org/pdf/2609.40115
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
- 中文摘要
带有可验证奖励的强化学习(RLVR)已成为提升训练后推理的重要方法。最新研究表明,即使偶尔产生正确解答,一些困难提示仍对学习具有抵抗力。我们重新审视了这一不可学习性现象,发现受影响的提示确实提升,约为可学习率的三分之一,而用于研究它们的难度定义提示集的可重复性远低于预期。这些困难标签是根据有限的抽样反应估计的。跨种子组合它们可以进一步改变选择的提示,而不仅仅是减少测量噪声。我们开发了一个基于抽样的框架,用于量化这种不稳定性,并确定难度分配可靠重现所需的评估量。我们还重新审视了为解释不可学习性的梯度相似性证据,并指出部分观察到的分离源于困难提示提供的正确展开较少,从而估算其梯度。匹配样本数量可以减弱梯度差异,但并未消除。总体来说,慢学习现象经受了我们的重新分析,而用于定义它的提示和解释它的证据都需要更细致的测量。
Tactile Curiosity Drives Robot Interaction
触觉好奇心驱动机器人互动
- Authors: Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.40134
- Pdf link: https://arxiv.org/pdf/2609.40134
- Abstract
Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which manipulation skills emerge. Existing intrinsic motivation methods based on model disagreement or epistemic uncertainty improve on isotropic noise, but they can also reward uncertainty in functionally irrelevant transitions, such as erratic motions in free space. In this work, we argue that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. By anchoring curiosity to the sense of touch, TacEx drives the robot to discover complex contact dynamics, learning to manipulate and grasp objects without task rewards or expert demonstrations during exploration. The interaction-dense dataset collected through this tactile-driven curiosity supports offline learning of downstream pick-and-place policies without additional environment interaction. We further use tactile-driven exploration to post-train vision-language-action (VLA) models. Although the VLAs are initially pre-trained without tactile feedback, post-training with TacEx substantially improves downstream performance while remaining highly sample-efficient.
- 中文摘要
通过强化学习(RL)掌握机器人操作技能仍然在很大程度上样本效率低下。最常见的强化学习算法依赖随机动作采样来发现新策略,导致智能体将大部分训练预算分配到自由空间中的运动,远离操作技能产生的接触点。基于模型不一致或认知不确定性的现有内在动机方法改进各向同性噪声,但它们也可能奖励功能无关的转移,如自由空间中的不规则运动。本研究论证触觉反馈为探索提供了自然信号,并引入了TacEx框架,该框架通过分解不同感官模态的模型不确定性,将好奇心引导至触觉通道,将触觉纳入认知不确定性驱动的探索中。通过将好奇心锚定于触觉,TacEx驱动机器人发现复杂的接触动态,学习在探索过程中无需任务奖励或专家演示即可操作和抓取物体。通过这种触觉驱动的好奇心收集的互动密集数据集支持离线学习下游的选放策略,无需额外环境交互。我们还进一步利用触觉驱动探索对视觉-语言-动作(VLA)模型进行后期训练。尽管VLA最初是未进行触觉反馈的预训练,但使用TacEx后期训练显著提升了下游性能,同时保持高度样本效率。
Role-Adaptive Policy Optimization for Offline Reinforcement Learning
离线强化学习的角色自适应策略优化
- Authors: Seonvin Cho, Soohyun Choi, Songnam Hong
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.40149
- Pdf link: https://arxiv.org/pdf/2609.40149
- Abstract
Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.
- 中文摘要
离线强化学习中的策略正则化平衡了策略改进与依赖不确定值估计的依赖。这种平衡在选择执行动作和提供批评者自助动作之间可能不同,但如TD3+BC等方法通过共享策略将这些角色耦合。我们提出了角色自适应策略优化(RAPO),根据策略更新系数在价值学习和执行中的角色进行调整。RAPO通过通过基于基础算法的演员目标形成的候选策略更新进行差异化来学习这些系数。对于TD3+BC,RAPO将引导行为者和执行行为者分离,并独立调整其系数:引导目标惩罚策略引发的目标值变化,执行目标评估局部策略改进替代。对于IQL,其价值学习已经独立于执行演员,RAPO保留了原始的值更新,并仅在优势加权策略提取中适配逆温度。D4RL移动和AntMaze任务的实验显示,TD3+BC的提升更大,其RAPO实例化平均优于基线。
Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization
SpTRSV优化中的强化学习引导图变换
- Authors: Buse Yılmaz
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.40159
- Pdf link: https://arxiv.org/pdf/2609.40159
- Abstract
Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. Recent graph transformation techniques address these limitations by modifying the dependency graph of the input matrix to improve parallel execution. Existing graph transformation strategies, however, rely on manually designed heuristics, making their development and adaptation to different optimization objectives challenging. This work proposes a reinforcement learning-guided graph transformation framework for SpTRSV, in which graph transformation is formulated as a sequential decision-making problem and an RL agent learns matrix-dependent transformation policies. Experimental results on real-world sparse matrices demonstrate level reductions of up to 94% and reductions of up to 80% in the coefficient of variation of level costs, while modifying only 1.50% of the rows in the highest case. On average, the RL- guided graph transformation achieves a 23% reduction in the number of levels and a 29% reduction in the coefficient of variation of level costs while rewriting only 0.82% of the matrix rows. Although the heuristic strategies generally achieve more aggressive level reduction(between 31% and 46%), the RL-based approach achieves the largest average reduction in the coefficient of variation of level costs, demonstrating its ability to balance competing graph transformation objectives. The results further show that the learned policies can be transferred to previously unseen matrices through curriculum learning and fine-tuning, while zero-shot experiments provide insights into the limitations of generalizing graph transformation policies across different sparsity patterns.
- 中文摘要
稀疏三角求解(SpTRSV)是众多科学和工程应用中的基础核。然而,稀疏三角矩阵固有的数据依赖性显著限制了可用的并行性,使得高效的工作负载分布具有挑战性。近期的图变换技术通过修改输入矩阵的依赖图来改善这些局限性,以提升并行执行。然而,现有的图转换策略依赖手动设计的启发式方法,使其开发和适应不同优化目标具有挑战性。本研究提出了一个基于强化学习引导的SpTRSV图变换框架,其中图变换被表述为顺序决策问题,强化学习代理学习矩阵依赖变换策略。现实稀疏矩阵的实验结果显示,层级降低可达94%,层级成本变异系数减少高达80%,而最高情况下仅修改1.50%的行。平均而言,RL引导图变换在仅重写0.82%矩阵行的情况下,实现了层级数量减少23%,水平成本变异系数减少29%。尽管启发式策略通常实现更激进的层级缩减(介于31%至46%之间),基于强化学习的方法在层级成本变异系数上实现了最大平均降低,展示了其平衡竞争图变换目标的能力。结果进一步表明,通过课程学习和微调,所学策略可以转移到此前未见的矩阵中,而零样本实验则揭示了在不同稀疏度模式中推广图变换策略的局限性。
Social-WM: Safety-Aware Latent World Models for Robot Social Navigation
Social-WM:机器人社会导航的安全意识潜世界模型
- Authors: Zhihao Zheng, Mooi Choo Chuah
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.40177
- Pdf link: https://arxiv.org/pdf/2609.40177
- Abstract
Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.
- 中文摘要
安全的社交导航不仅需要机器人预见其行为的未来后果,还需要判断名义上的行动是否能在周围的物理和社会约束下实际执行。我们提出了Social-WM,一种由自我中心的RGB视频序列训练的高效潜在世界模型规划框架。我们的关键观察是,社交导航体验中名义动作与可实现动作之间存在系统性差异:名义上的前进动作可以在自由空间中完全执行,但在朝向行人或障碍物时需要加以约束。Social-WM通过动作条件未来预测直接学习这些安全相关后果,其中目标是每次指令后实际观察到的未来。我们还进一步引入了一个可实现的逆动力学目标,将观察到的潜在转变与实际实现的动作联系起来,而非名义上的动作。部署时,候选动作通过潜在世界模型想象,逆动力学模型则估计其实现性;名义上可实现的差异在执行前提供安全信号。学习的动态与可实现性模型保持目标无关性,支持位置和图像目标导航。在Social-HM3D上,Social-WM实现63.77%成功率,将人与碰撞降低至21.67%,在零射击转移至Social-MP3D时保持强劲性能,无需显式行人追踪、特权人类状态或在线强化学习。
MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
MemLife:策划与推理长期自我中心的视频记忆
- Authors: Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Aidong Zhang, Xin Luna Dong
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.40195
- Pdf link: https://arxiv.org/pdf/2609.40195
- Abstract
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.
- 中文摘要
长期自我中心的视频使个性化AI助手能够推理日常生活。然而,随着视频历史长达数百小时、跨月甚至数年,每次查询重新处理原始片段的计算量变得极为困难。记忆系统通过将视频压缩成文本表示提供了可扩展的替代方案,但常在实际基准测试中失败:要么内存无法保存关键证据,要么由于搜索空间不断扩大的检索竞争,检索器未能找到相关条目。为应对这些挑战,我们引入了MemLife,一种多模态记忆系统,构建基于实体的第一人称文本片段,并通过时间索引代理读取器检索。在没有训练或查询时视频访问的情况下,MemLife在四个长期视野基准中相较最强无训练基线提升4.6%至12.0%。为了进一步提升记忆质量,我们提出了MemOpt,一种强化学习框架,优化记忆编写者,使其产生忠实、信息丰富且可检索的记忆。MemOpt在不同视频和问题分布中,持续提升MemLife2.7%至5.0%,且这些提升在写作者和读者骨干及记忆系统中具有普遍性。
PhantomEnvironments: Training LLM Agents in Fictional Worlds
PhantomEnvironments:虚构世界中的大型语言模型代理培训
- Authors: Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q. Weinberger
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.40221
- Pdf link: https://arxiv.org/pdf/2609.40221
- Abstract
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
- 中文摘要
用强化学习(RL)训练LLM代理时,环境必须提供可验证的奖励、支持长期互动并低成本扩展,成为瓶颈。现有方法依赖于昂贵的人工策划数据或存在幻觉和基准污染风险的LLM生成环境。我们展示了,LLM可以通过完全由规则生成的合成环境训练成具备能力的搜索代理,这些生成无需LLM,且边际成本为零。我们构建了PhantomEnvironments,这是基于虚构世界的多回合强化学习环境,代理必须搜索模板化文章语料库以回答多跳问题。尽管这些极其简单的环境能生成能转移到现实多跳搜索基准的代理,且常常优于新基准测试的现实训练数据。训练有素的代理会推广到看不见的虚构宇宙,Qwen模型则学会与问题难度大致线性地扩展搜索预算,这表明仅靠环境互动就能实现涌现的搜索扩展。消除环境复杂性表明,跳数驱动的传递远不止约束或比较:即使是最简单的规则生成环境,也是训练可推广LLM代理的高效且免费的资源。
EviRover: Reinforcing Agentic Perception Beyond a Glance
EviRover:强化超越一瞥的智能感知
- Authors: Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.40230
- Pdf link: https://arxiv.org/pdf/2609.40230
- Abstract
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
- 中文摘要
视觉感知通常被表述为从单次一瞥图像中做出的一次性预测,假设图像内容和模型的参数化知识足以解决查询。但这一假设在依赖细微视觉细节或需要知识密集型且最新信息的现实场景中常常失效。我们将此类情况称为 \textit{证据不足下的感知},并将感知表述为一种能够获得超越单次一瞥信息的代理过程。为解决该场景数据的缺失,我们设计了两条专用的数据生成流程,分别用于训练,分别是 EviRover-SFT-5K 和 EviRover-RL-12K。我们还构建了 EviLens,这是一个由人类验证的基准测试,包含五个感知类别的 688 个实例。基于这些数据,我们介绍了据我们所知,EviRover是首个明确训练通过交互来解析感知查询的感知代理,采用监督微调和代理强化学习。实验显示,4B EviRover在EviLens上平均性能优于其骨干30分,性能可与先进专有模型媲美。这些提升超越了EviLens,延伸至WebEyes、传统感知基准测试和通用多模态基准测试,包括BrowseComp-VL提升15分。所有代码、模型和数据均已发布。
Semifactual Credit-Augmented Policy Optimization
半事实性信用增强政策优化
- Authors: Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.40360
- Pdf link: https://arxiv.org/pdf/2609.40360
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at this https URL.
- 中文摘要
带有可验证奖励的强化学习(RLVR)提升了大型语言模型(LLMs)的推理能力,但其预测仍对任务无关提示特征敏感。我们通过半事实提示干预来研究这种敏感性,保留了底层问题及其答案。我们的分析揭示了代币级敏感性存在显著差异,并显示在解码过程中抑制高漂移的代币候选人可以在不更新模型权重的情况下提高推理准确性。这些发现凸显了群相对策略优化(GRPO)的局限性,该方法赋予每个响应代币相同的结果衍生优势,可能在有用推理的同时强化潜在的虚假依赖。基于这一观察,我们引入了半事实性增强策略优化(SCAPO),这是一种因果启发的GRPO变体,将半事实稳定性融入代币级信用分配。SCAPO在半事实干预下测量固定反应的代币概率漂移,并使用归一化稳定性评分来降低早期训练中相对不稳定代币的优势,同时不额外给予稳定性加分。在Qwen3-4B-Base和Qwen3-1.7B-Base上,SCAPO分别将AIME 2024-2026的准确性提升了5.63个百分点和4.17个百分点。在这两个模型尺度下,SCAPO在大多数评估的数学基准测试以及所有评估的非分布基准测试中均取得了最佳成绩。这些结果表明,半事实稳定性为通过RLVR中更细粒度的信用赋值,为提升推理和泛化提供了有效的训练信号。该代码可在此 https 网址获取。
Keyword: diffusion policy
Conditional Generation of Creative Chess Puzzles with Diffusion Models
利用扩散模型的创造性国际象棋谜题的条件生成
- Authors: Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.38577
- Pdf link: https://arxiv.org/pdf/2609.38577
- Abstract
While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.
- 中文摘要
尽管现代语言模型展现出令人印象深刻的生成能力,但它们常常在受限且反直觉的创造性任务中遇到困难。为解决这一限制,我们探索国际象棋谜题生成作为计算创造力和推理的严格测试平台,在这一领域中,改变单个棋子可能使整个解法失效。我们提出了一种利用掩蔽扩散模型的条件式国际象棋谜题生成新方法。与以往方法不同,我们的非方向扩散方法允许对特定战术主题和部分棋局局面进行条件化。我们引入了一种新颖的辅助任务——同时最佳走法预测,该任务使解的唯一性提升了11.6%,主题条件准确率提升了2.5%。为进一步优化解的独特性和主题条件,我们建立了基于去噪扩散策略优化(DDPO)的强化学习框架。这种强化学习训练使独特且主题匹配局面的产出提升了89.1%。最后,我们发布了首批用于国际象棋谜题生成的开放权重模型(附录B),为可控且富有创造力的生成提供了新的途径。
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
通过Q分数匹配的顺序值恢复实现无环逆强化学习
- Authors: Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.38955
- Pdf link: https://arxiv.org/pdf/2609.38955
- Abstract
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
- 中文摘要
逆强化学习(IRL)旨在恢复一个能解释专家演示的奖励函数。现有的IRL方法通常依赖于在奖励学习和策略优化之间交替进行的双级优化过程,导致计算负担和训练不稳定性。在本研究中,我们引入了一条不同的路径,通过利用扩散策略完全消除策略优化。我们的关键见解是,扩散策略编码了最优软Q函数的作用梯度结构,使奖励学习能够被表述为一系列价值恢复问题,从而绕过以往IRL方法固有的奖励策略循环。具体来说,我们的方法分为三个阶段进行:(I)通过动作梯度匹配恢复最优软Q函数,并以Gumbel回归启发的方式估计相应的软价值函数(Q值的LogSumExp);(II)通过推断状态依赖的偏移来校准这些软值;(III)通过强制Bellman一致性提取奖励。这导致了无环逆强化学习(LFIRL),这是一种完全离线的算法,以简单、无环且顺序的方式运行。LFIRL实现简单,显著提升训练效率,同时保持强劲的奖励恢复表现。在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基准测试中,LFIRL在最快基线上实现了2-3倍的加速,同时在奖励恢复质量上与最先进方法相当甚至超越。
Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization
文本到三维策略:用于看不见规范泛化的细粒度语言-行为对齐
- Authors: Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.39599
- Pdf link: https://arxiv.org/pdf/2609.39599
- Abstract
3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
- 中文摘要
三维视觉运动策略为空间精确操作提供了坚实基础,但当前的文本到三维策略难以遵循演示覆盖之外的未见细粒度行为规范。我们将这一挑战作为看不见的规范推广来研究,语言指定了策略训练中缺失的行为显著变异,如目标位置、位移或发明状态。我们发现预训练语言表示和传统全局行为-语言对齐捕捉了粗略的任务语义,但常常模糊了需要不同行为的附近规范。我们介绍了T3DP,一种用于细粒度语言-行为对齐的文本到三维策略框架。T3DP不将每个指令和演示压缩成单一全局嵌入,而是保留其局部结构,并在语言元素与行为片段之间建立双向令牌级对应关系。这直接奠定了其影响行为组件的细微语言差异基础,防止密切相关的规范在表示空间中崩溃。由此产生的规范敏感语言表示条件是基于点云的三维扩散策略,使得对看不见的行为规范进行更精确的控制,而无需修改底层策略架构。在Meta-World、ManiSkill和RoboTwin中,T3DP在全球语言-行为对齐方面平均提升了预约规范成功率+11.0-14.2个百分点,且在所有15个任务族均有提升;在真实机器人任务中,它进一步将平均成功率从47.5%提升至65.0%(+17.5分)。表示和动作探针分析显示,细粒度对齐更能更好地保留规范几何和动作相关变异,将局部行为的基础与下游控制联系起来。