生成时间: 2026-10-06 22:54:38 (UTC+8); Arxiv 发布时间: 2026-10-06 20:00 EDT (2026-10-07 08:00 UTC+8)
今天共有 97 篇相关文章
Keyword: reinforcement learning
What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study
可验证的奖励教会视频语言模型关于时间的什么?一项受控多模型研究
- Authors: Avyay Sadhu, Patrick Cooper
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.03792
- Pdf link: https://arxiv.org/pdf/2610.03792
- Abstract
Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group relative policy optimization under three data recipes: verified (synthetic CLEVRER questions with exact answer and event-order rewards), unverified (self-supervised pretext tasks over 43,751 real web videos), and a 1:1 mixture, plus a verified+real arm that adds 4,000 verifiable questions on real video. Each cell is evaluated in-domain and on out-of-domain real video (a NExT-QA temporal stress set and an MVBench subset), with frames in order, shuffled, and absent. (1) Verified training yields large in-domain gains that shrink as base competence grows (+14 to +19 points on weaker models; +6 on the strongest). (2) Much of the gain is non-visual: accuracy with no frames rises nearly as much as with frames. (3) Verified-only training can severely degrade out-of-domain accuracy with no sign during training: Qwen3-VL-8B loses 26.7 and 25.2 points on the two real-video sets, while the mixture never significantly degrades a model trained on it. Adding real verified questions removes that loss (-2.3 points, within noise of base) and keeps a +9.3 in-domain gain, so the cause is narrow synthetic-only data, not verification. (4) No recipe induces temporal-order grounding: across 41 evaluations the ordered-versus-shuffled gap is indistinguishable from zero in 39 and marginal in two, despite an event-order reward. Verifiable rewards improve benchmark accuracy without temporal understanding. Report no-frame controls, and mix in real video to guard against out-of-domain degradation.
- 中文摘要
可验证奖励强化学习(RLVR)在语言模型中带来了显著的推理提升,且可验证的视频基准也使其适用于因时视频问答。我们研究RLVR如何教授视频语言模型关于时间的内容。我们通过小组相对策略优化微调了四个开放模型(Qwen3-VL-8B/4B,Qwen2.5-VL-7B,Gemma-3-12B),采用三种数据配方:已验证模型(合成CLEVRER题,具有精确答案和事件顺序奖励)、未验证模型(对43,751个真实网络视频进行自我监督的前述任务)、1:1混合模型,加上一个已验证+真实模型,增加了4,000个可验证的真实视频问题。每个单元在域内和域外实视频(NExT-QA时间应力集和MVBench子集)上评估,帧顺序、洗牌和缺失。(1)经过验证训练的领域内收益大幅下降,但随着基础能力提升而缩小(弱模型为+14至+19点;最强模型为+6)。(2)大部分提升为非视觉效果:无帧准确率提升幅度几乎与有帧时相同。(3)仅验证训练在训练期间无迹象时会严重降低域外准确性:Qwen3-VL-8B在两个实视频集中分别损失26.7点和25.2分,而混合后从未显著降低基于该模型的模型。添加真实验证问题可消除该损失(-2.3分,在基础噪声范围内),保持+9.3的领域内收益,因此原因是狭义的合成数据,而非验证。(4) 没有任何方法能引发时间顺序基础:在41次评估中,有序与洗牌的差距在39个中为零,2个为边缘,尽管有事件顺序奖励。可验证的奖励在不理解时间的情况下提升基准测试准确性。报告无帧控制,并混合真实视频以防止域外退化。
Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking
晶体发生器离散组成通道上的强化学习:验证的收益与奖励黑客
- Authors: Pawan Prakash, Philipp Höllmer, Addis Fuhr, Peter Hirschfeld, P. Ganesh, Stefano Martiniani, Richard Hennig
- Subjects: Subjects:
Machine Learning (cs.LG); Materials Science (cond-mat.mtrl-sci)
- Arxiv link: https://arxiv.org/abs/2610.03880
- Pdf link: https://arxiv.org/pdf/2610.03880
- Abstract
Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimization (GRPO) to align a generative model based on stochastic interpolants and discrete flow matching with general black-box reward functions through reinforcement learning. Atom types are generated by a discrete flow and the policy gradient of our generalization of GRPO directly acts on the likelihoods of the atom-type transitions, which differentiates our work from previous reinforcement-learning approaches for diffusion and flow-based generative models of crystalline materials. We introduce a reward function that raises the yield of metastable, unique and novel structures (mSUN) from 13.4% for the pretrained model to 45.5% for the reinforced model, as evaluated by a community benchmark. Our reward also improves the performance of a reinforcement learning framework for crystalline materials based on latent denoising diffusion models. At the same time, we find that directly reinforcing atom-type transition likelihoods enables reward exploitation that has to be prevented with explicit guards. The same analysis also exposes a gap in the community metric. Single-element structures in distinct packings are counted as metastable, unique and novel materials and inflate mSUN without yielding any new compounds. A stability claim is only as good as its reference hull. We report every result split by the number of reference phases behind it and argue that benchmarks should do the same.
- 中文摘要
逆材料设计是计算材料发现的长期目标。晶体材料生成模型通常训练以匹配结构数据库的分布,而其训练目标中没有指向特定设计目标,如目标性质。我们使用群相对策略优化(GRPO),通过强化学习将基于随机插值和离散流匹配的生成模型与一般黑箱奖励函数对齐。原子类型由离散流生成,我们对GRPO推广的策略梯度直接作用于原子类型跃迁的似然度,这使我们的工作区别于以往针对扩散和基于流动的晶体材料生成模型的强化学习方法。我们引入了一种奖励函数,使亚稳、独特和新颖结构(mSUN)的产率从预训练模型的13.4%提升到强化模型的45.5%,这是基于社区基准评估的结果。我们的奖励还提升了基于潜在去噪扩散模型的晶体材料强化学习框架的性能。同时,我们发现直接强化原子型跃迁似然能够实现必须通过显式保护防止的奖励利用。同一分析还揭示了社区指标中的空白。不同包装中的单元素结构被计为亚稳、独特且新颖材料,并膨胀mSUN而不产生任何新化合物。稳定性声明的价值取决于其参考包壳。我们报告每个结果按其后参考相数划分,并主张基准测试应如此。
Barrier-Shaped Recurrent Reinforcement Learning for Autonomous Landing on a Heaving Ship Deck
障碍状反复强化学习,用于在颠簸船甲板上自主着陆
- Authors: Ritwik Shankar, Chiranjeev Prachand, Abhishek, Soumya Ranjan Sahoo
- Subjects: Subjects:
Systems and Control (eess.SY); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.03924
- Pdf link: https://arxiv.org/pdf/2610.03924
- Abstract
This paper addresses autonomous landing of unmanned aerial vehicle (UAV) rotorcrafts on a heaving ship deck using a recurrent policy trained via asymmetric actor-critic reinforcement learning: the critic sees 6 s of future deck height during training, while the actor sees only what the onboard sensors provide in flight, a 17-dimensional state made of the vehicle's position, velocity and attitude relative to the deck and outputs world-frame velocity and yaw-rate commands. Training is performed across 4096 parallel simulated environments, followed by fine-tuning in 16 environments with an onboard vision pipeline in the loop. A control-barrier-function (CBF) stopping margin on the deck-relative vertical state is incorporated at two stages: as a reward term during training, where it halves the median simulated contact speed relative to a policy trained without it, and as a runtime safety filter at deployment, evaluated at every control step to abort and retry the descent when the margin is violated. Because the autopilot's disarm logic cannot detect the vehicle resting on a moving deck, proximity-based thrust cutoff at touchdown is commanded directly in the landing pipeline. The proposed approach is validated using a parallel-manipulator-platform-based deck emulator that reproduces the heaving motion of the ship deck (scaled to 0.70 m peak-to-peak, 7.5 s mean period) and a quadcopter UAV with an onboard camera. Across 31 motion-capture and 20 vision-based trials, the UAV landed every time, with median times to contact of 6.5 and 7.3 s; 77% and 50% landed on the first attempt, with mean deck-relative speeds of 0.29 and 0.27 m/s, respectively, at the instant of thrust cutoff, which is the last speed under the policy's control. Supplementary video: this https URL
- 中文摘要
本文探讨了无人机旋翼机在颠簸船甲板上的自主着陆,采用通过非对称演员-批判者强化学习训练的循环策略:批判者在训练中看到未来甲板高度的6秒,而执行者仅看到飞行中机载传感器提供的17维状态,即飞行器相对于甲板的位置、速度和姿态,并输出世界帧速度和偏航率指令。训练跨越4096个并行模拟环境进行,随后在16个环境中通过机载视觉流水线进行微调。在甲板相对垂直状态上引入控制障碍函数(CBF)停止余裕,分为两个阶段:作为训练期间的奖励项,使模拟接触速度中位数相对于未训练策略的目标减半;作为部署时的安全过滤器,在每个控制步骤评估以中止并重试下降,当余距被违反时。由于自动驾驶仪的解除武装逻辑无法检测飞行器停靠在移动甲板上,着陆时基于接近的推力切断直接在着陆管道中指令。所提方法通过并行机械臂平台甲板模拟器验证,模拟舰甲板的起伏运动(峰峰高度缩放至0.70米,平均周期7.5秒),以及配备机载摄像头的四旋翼无人机。在31次动作捕捉和20次视觉测试中,无人机每次着陆,中位数接触时间分别为6.5秒和7.3秒;77%和50%成功首次着陆,推力切断瞬间的平均甲板相对速度分别为0.29和0.27米/秒,这是政策控制下的最后速度。补充视频:此链接
Network Adaptation in IRS-Aided Hybrid RF/VLC Systems Using Cooperative Multi-Agent DRL
利用合作多智能体日程学习(DRL)的IRS辅助混合射频/极低噪声系统中的网络适配
- Authors: Ahrar N. Hamad, Ahmad Adnan Qidan, Taisir E.H. El-Gorashi, Jaafar M. H. Elmirghani
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.03964
- Pdf link: https://arxiv.org/pdf/2610.03964
- Abstract
Hybrid radio frequency (RF) and visible light communication (VLC) networks have emerged as a promising solution for high-capacity indoor wireless connectivity in sixth-generation (6G) systems. However, the limited optical coverage and vulnerability of VLC links to line-of-sight (LoS) blockage under user mobility remain fundamental challenges. In this work, a mirror-based intelligent reflecting surface (IRS) is deployed to assist a dynamic indoor hybrid RF/VLC network, where each mobile user is exclusively assigned to either the IRS-enhanced VLC subnetwork or the RF subnetwork through a binary selection decision. A joint optimization problem is then formulated to maximize proportional fairness by jointly optimizing the RF/VLC technology selection, power allocation, and IRS mirror roll and yaw orientation angles. To enable real-time adaptability, the problem is reformulated as a Markov decision process (MDP) and solved using a cooperative multi-agent deep reinforcement learning (DRL) algorithm based on centralized training with decentralized execution. Simulation results demonstrate the superior performance of the optimized hybrid network compared with optimized standalone VLC and RF networks. The results further validate the practicality and effectiveness of the proposed DRL framework compared to widely adopted DRL algorithms as well as conventional model-based optimization approaches.
- 中文摘要
混合射频(RF)和可见光通信(VLC)网络已成为第六代(6G)系统中高容量室内无线连接的有前景解决方案。然而,VLC链路在用户移动性下有限的光学覆盖和易受视距(LoS)阻挡的脆弱性仍是根本性挑战。在这项工作中,部署了基于镜像的智能反射面(IRS)以辅助动态室内混合射频/VLC网络,每个移动用户通过二元选择决策独占分配到IRS增强型VLC子网或射频子网。随后提出联合优化问题,通过联合优化RF/VLC技术的选择、功率分配以及IRS镜面的滚动和偏航方向角,最大化比例公平性。为实现实时适应性,问题被重新表述为马尔可夫决策过程(MDP),并基于基于集中训练和去中心化执行的协作多智能体深度强化学习(DRL)算法求解。模拟结果表明,优化后的混合网络优于优化的独立VLC和射频网络。结果进一步验证了所提DRL框架相较于广泛采用的DRL算法及传统基于模型优化方法的实用性和有效性。
Exploration-Preserving Policy Optimization
勘探保全策略优化
- Authors: Hangzhan jin, Mohammad Hamdaqa, Doina Precup
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04011
- Pdf link: https://arxiv.org/pdf/2610.04011
- Abstract
Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using prompt-relative, length-normalized response surprisal and prompt pass rate. ExPPO combines bounded shaping with shared normalization to preserve verifier polarity and approximately maintain each prompt group's total absolute sequence-advantage mass. Our analysis characterizes response-level credit allocation alongside sampled mode updates, deriving local conditions for gains in entropy and correct-mode discovery. Experiments show improved in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and strong coverage at large sampling budgets. A controlled multi-answer evaluation further demonstrates increased correct-mode yield and gains in diversity among verified-correct responses. Code is available at this https URL
- 中文摘要
带有可验证奖励的强化学习提升了推理能力,而学习的分配则使得哪些解在反复抽样下仍可访问。群体相对目标将同等优势分配给同样奖励的响应,使得总功劳与抽样模式频率成正比。我们引入了探索-保全策略优化(ExPPO),这是一条轻量级优势塑造规则,利用提示相对、长度归一化的反应惊喜和即时通过率重新分配信用。ExPPO结合了有界整形与共享归一化,保持验证者极性,并大致保持每个提示组的总绝对序列优势质量。我们的分析将响应级信用分配与抽样模式更新并列,推导出熵和正确模式发现的局部条件。实验显示,域内外推理覆盖率提升,汇总响应准确率更高,且在大抽样预算下覆盖率强。受控多重答案评估进一步展示了正确模式产率的提高和验证正确答案多样性的提升。代码可在此 https URL 获取
Reinforcement Learning with Comparative Evidence for Social Intelligence
社会智能的强化学习与比较证据
- Authors: Keane Ong, Yuriel Ryan, Sabri Boughorbel, Vladimir Necula, Jack Wei Lun Shi, Rui Mao, Roy Ka-Wei Lee, Adriel Kuek, Nancy F. Chen, Erik Cambria, Gianmarco Mengaldo, Paul Pu Liang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04072
- Pdf link: https://arxiv.org/pdf/2610.04072
- Abstract
Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreover, core social targets such as affect, intent, preference, and pragmatic meaning are often ambiguous. The same behavior can support multiple plausible interpretations, making it difficult to verify which is best supported. To address this challenge, we introduce Reinforcement Learning with Comparative Evidence (RLCE), a reinforcement learning method that learns social understanding from unlabeled training data without constructing rewards from ground-truth annotations. Given distinct answers in a rollout group, RLCE constructs evidence tests that identify observable evidence favoring an answer over another, validates these tests against the input sample, and aggregates test outcomes to determine the best-supported interpretation. Tests are regenerated as the policy produces new answers, enabling them to evolve with the policy. Across four benchmarks spanning affect, pragmatics, communicative intent, and preference, RLCE attains the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Gains over the strongest baseline reach up to +18.93 points. Analyses further show that RLCE exhibits a larger share of reward variation between correct and incorrect predictions than compared rubric methods, can overturn erroneous policy-derived preferences, and benefits from pairwise test construction, compositional test aggregation, and on-policy test evolution.
- 中文摘要
开发具有社会智能的人工智能仍然高度依赖人工注释数据,限制了社会理解模型所能获得的规模和广度。从未标记数据中推导训练信号的方法提供了超越这种依赖的路径,但社会预测缺乏数学和编码中可用的验证预言机。此外,情感、意图、偏好和语用意义等核心社会目标往往存在模糊性。同一行为可能支持多种合理解释,使得验证哪种解释更为有效变得困难。为应对这一挑战,我们引入了带比较证据的强化学习(RLCE),这是一种强化学习方法,通过未标记的训练数据学习社会理解,而无需基于真实注释构建奖励。在推出组中,给定不同答案时,RLCE构建证据测试,识别有利于某一答案的可观察证据,验证这些测试与输入样本的对照,并汇总测试结果以确定最有支持性的解释。随着政策产生新答案,测试会重新生成,使其能够随政策演进。在涵盖情感、语用学、交际意图和偏好的四个基准中,RLCE在七种不使用实地训练标签的奖励方法中表现最优,包括共识、政策LLM-评判验证、多模态共进化和基于评分标准的奖励。超过最强基线的提升可达+18.93分。分析进一步表明,RLCE在正确与错误预测之间的奖励差异比例高于对比的评分标准方法,能够推翻错误的策略衍生偏好,并受益于两对测试构建、组合测试聚合和策略内测试演进。
IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation
IdeaScientist:策划扎实科学构思的代理
- Authors: Jiarui Liu, Renjie Tao, Yiwei Liao, Chuanyang Jin, Kai Sun, Xiao Yang, Xinyuan Zhang, Xilun Chen, Zhuangqun Huang, Lechen Zhang, Yongjin Yang, Yinghui He, Weihao Xuan, Rakesh Wanga, Anuj Kumar, Mona T. Diab, Wen-tau Yih, Xin Luna Dong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.04074
- Pdf link: https://arxiv.org/pdf/2610.04074
- Abstract
Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Accordingly, we introduce IdeaScientist, which decomposes ideation into gap finding, innovation, and report writing, and trains each role with reinforcement learning. These roles identify limitations in related work, draw solution intuitions from analogous problem settings, and develop those intuitions into complete research proposals. To facilitate discovery of insights across domains, we construct the Svalbard Idea Vault, a corpus of 2.77M decomposed research ideas for retrieval, training, and temporally controlled evaluation. Our evaluation restricts access to literature available before a cutoff date and assesses how closely proposed directions align with those later explored in 15K papers authored by human researchers. On Qwen3.6-27B, IdeaScientist outperforms the strongest open-source autoresearch baseline by 14.0%, driven mainly by gains in novelty. On this 27B open backbone, IdeaScientist even outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4, by up to 5.9%.
- 中文摘要
尽管科学研究自动化迅速进展,但生成有前景且扎实的研究解决方案仍是核心挑战。我们将研究构思作为独立任务,基于直觉构建:一个领域的挑战通常可以通过解决另一个类似挑战的机制来实现。因此,我们引入了IdeaScientist,将创意分解为空白发现、创新和报告撰写,并通过强化学习培训每个角色。这些角色识别相关工作的局限性,从类似问题环境中提取解决方案直觉,并将这些直觉发展成完整的研究提案。为促进跨领域洞见的发现,我们构建了斯瓦尔巴群岛创意库,这是一个包含277万个分解研究想法的语料库,用于检索、培训和时间控制评估。我们的评估限制了截止日期前可获得文献的访问,并评估提出的方向与后来由人类研究人员撰写的1.5万篇论文中探讨的方向的契合度。在Qwen3.6-27B阶段,IdeaScientist的表现优于最强的开源自研基线14.0%,主要得益于新颖性提升。在这27B开放骨干网中,IdeaScientist甚至比Claude Code SDK的Claude-4.8-Opus和Codex SDK的GPT-5.4高出5.9%。
PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO
PB-GRPO:通过偏好批处理GRPO从人格驱动模拟学习社会适应LLM代理
- Authors: Jingquan Wang, Jun Yin, Xu Han, Yongsheng Mei, Jie Hao, Bin Guo
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04132
- Pdf link: https://arxiv.org/pdf/2610.04132
- Abstract
Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.
- 中文摘要
构建能够社交表现良好的大型语言模型,而不仅仅是正确,需要构建表现良好的大型语言模型,不仅仅提供本地有用的回应。具备社交能力的智能体必须推断用户未明说的目标,尊重他们的偏好,并随着对话展开进行调整。这些行为本质上是多回合且社交性的,难以优化:真实交互数据稀缺,用户偏好通常是潜在的,而非直接可观察。为应对这些挑战,我们基于人格驱动的社交模拟环境(包括角色库、基于LLM的用户模拟器和从[0,1]范围的用户满意度评分系统组成),引入偏好批处理GRPO(PB-GRPO),这是一种从对话级反馈中学习社会适应策略的后训练算法。与普通GRPO相比,PB-GRPO通过对一桶具有相似偏好的用户进行归一化估计,计算优势,稳定了多样化社会人群的训练。实证证据表明,PB-GRPO在模拟环境中优于强强化学习基线,提升了模型的社会行为。
How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws
强化学习如何重塑大型语言模型推理:可转移性、覆盖率与扩展性规律
- Authors: Ziheng Cheng, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.04158
- Pdf link: https://arxiv.org/pdf/2610.04158
- Abstract
Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand an LLM's reasoning boundary, or merely reweight its existing reasoning space? We revisit these phenomena across Qwen and Gemma model families, showing both cross-domain gains and forgetting, while coverage at large sampling budgets increases on some tasks and decreases on others. Detailed analysis of solution traces before and after RL indicates a shift in the reasoning strategies the model employs, motivating a two-stage autoregressive policy model that separates \emph{strategy selection} from problem-specific execution. Within this framework, we prove how RL's implicit bias reshapes strategy preferences, allowing gains on some tasks while suppressing strategies required by others. This mechanism can also broaden or narrow coverage at a given sampling budget even without expanding strategy support. We further provide theoretical justifications for log-sigmoid and log-linear scaling laws in RL compute, and evaluate their predictive power. Together, these results connect changes in strategy selection to cross-domain transfer, reasoning coverage, and compute scaling.
- 中文摘要
近期关于强化学习(RL)的研究报告了关于大型语言模型(LLM)推理的证据似乎相互矛盾。数学训练可以提升其他领域的表现,但Pass@1的提升可能伴随着低于基础模型Pass@$N美元。这引出了一个根本性问题:强化学习是扩展了大型语言模型的推理边界,还是仅仅重新加权其现有推理空间?我们回顾了Qwen和Gemma模型家族中的这些现象,展示了跨域的增长和遗忘现象,同时在大范围采样预算下的覆盖率在某些任务上增加,而在其他任务中减少。对强化学习前后解痕迹的详细分析表明模型采用的推理策略发生了转变,促使采用了两阶段自回归策略模型,将策略选择与问题特定执行分离。在此框架下,我们证明了强化学习隐性偏见如何重塑策略偏好,使某些任务获得收益,同时抑制其他任务所需的策略。即使不扩展策略支持,这一机制也能在给定采样预算内扩大或缩小覆盖范围。我们还进一步为强化学习计算中的对数S形和对数线性尺度律提供了理论依据,并评估其预测能力。这些结果共同将策略选择的变化与跨域转移、推理覆盖和计算尺度联系起来。
Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks
异步几乎可以自由地用于长视野能动任务的演化策略
- Authors: William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04196
- Pdf link: https://arxiv.org/pdf/2610.04196
- Abstract
LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-training addresses this inefficiency by consuming trajectories as they arrive, but introduces policy lag and off-policy optimization. Evolution strategies (ES) offer a backpropagation-free alternative for LLM post-training, yet it relies on a larger number of rollouts and existing practices have remained largely synchronous. In this short-form paper, we introduce bounded-staleness asynchronous ES and demonstrate it on Endless Terminals benchmark using Qwen2.5-7B-Instruct. Across three evaluation seeds, natural Async-1 matches synchronous ES, achieving 25.9\% versus 25.4\% held-out success. Controlled schedules that delay 10\% of each update cohort by four or eight policy updates reduce success by only 1.6 and 3.1 percentage points, respectively, without explicit off-policy correction. GRPO performs better overall, reaching 29.0\% held-out success, but importantly our results show that ES tolerates moderate policy staleness with limited degradation, opening possibilities for future improvement of ES-based post-training with asynchronous algorithms. To the best of our knowledge, we are the first to demonstrate the effectiveness of sync and async ES on a multi-turn terminal style agentic coding task.
- 中文摘要
基于LLM的长视野代理后训练常被推出生成瓶颈:轨迹跨越多个交互回合,完成时间差异显著,同步更新障碍使更快的工人等待落后者。被LLM后训练采用的异步强化学习通过消耗轨迹来解决这种低效率,但引入策略滞后和非策略优化。进化策略(ES)为LLM后训练提供了无反向传播的替代方案,但依赖更多次展开,现有实践基本保持同步。在这篇短文中,我们介绍了有界陈旧异步ES,并在Endless Terminals基准测试中演示,使用Qwen2.5-7B-Instruct。在三个评估种子中,自然异步-1与同步ES匹配,达到25.9%对25.4%的延迟成功率。控制计划将每个更新队列延迟10%的策略更新四次或八次,分别仅使成功率降低1.6个百分点和3.1个百分点,且无明确的非策略修正。GRPO整体表现更好,保持成功率达到29.0%,但重要的是我们的结果显示ES能容忍中度策略陈旧且性能下降有限,为未来基于ES的异步算法后训练改进打开了可能。据我们所知,我们是首个展示同步与异步ES在多回合终端风格代理编码任务中有效性的机构。
ROOT: Discovering Rewards for User-Specified Embodied Behaviors
ROOT:发现用户指定具身行为的奖励
- Authors: Eren Sadikoglu, Aditya Taparia, Xinyuan Liu, Ransalu Senanayake
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.04250
- Pdf link: https://arxiv.org/pdf/2610.04250
- Abstract
Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over a persistent experiment tree that stores reward programs, trained policies, and rollout observations, together with behavioral insights distilled by a video-language model, to diagnose behavioral failures and guide subsequent reward refinements. We evaluate ROOT on seven tasks across four embodiments: simulated Hopper, HalfCheetah, Ant, Unitree Go2, and as well as the real-world Unitree Go2. ROOT produces behaviors that better align with user intent than those generated by existing LLM-based reward-generation methods, achieving up to 86.8% locomotion-completeness accuracy and improving Vid-LLM behavioral alignment from 3.56/5 to 4.14/5, a 16.5% improvement over baselines. Human evaluations further support these results, with ROOT preferred in 51-63% of pairwise comparisons.
- 中文摘要
针对具身控制的强化学习仍受限于奖励指定的困难。尽管最新的大型语言模型(LLM)方法可以从自然语言描述中综合奖励函数,但它们往往无法捕捉人类关心的微妙行为属性,如自然步态、姿势和运动风格。这一限制源于许多期望行为更容易通过视觉识别而非用奖励函数编码。我们介绍了可观察树奖励优化(ROOT),这是一个用于发现能使所学策略与用户指定的具身行为对齐的奖励函数的框架。ROOT将奖励设计视为对持续实验树的观察引导搜索,树中存储奖励程序、训练策略和推广观察,结合视频语言模型提炼的行为洞见,用以诊断行为失败并指导后续奖励细化。我们在四个实例中评估了七个任务:模拟的Hopper、HalfCheetah、Ant、Unitree Go2以及现实世界的Unitree Go2。ROOT产生的行为比现有基于LLM的奖励生成方法更符合用户意图,达到高达86.8%的移动完整性准确率,并将Vid-LLM行为对齐率从3.56/5提升至4.14/5,较基线提升16.5%。人工评估进一步支持这些结果,在51%-63%的两对比较中更倾向于ROOT。
Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry
RLVR在融合格罗莫夫-瓦瑟斯坦几何上的层级学分作业
- Authors: Qi Yu, Ruizhong Qiu, Zhichen Zeng, Xuying Ning, Yanjun Zhao, Dongqi Fu, Yinglong Xia, Hong Li, Hanghang Tong
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.04344
- Pdf link: https://arxiv.org/pdf/2610.04344
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works refine credit assignment of GRPO based on local signals such as token locations or entropy, they often fail to capture the global semantic novelty of a reasoning behavior relative to the current policy. In this work, we propose a hierarchical credit assignment approach for group-based RLVR methods, called HarA, which identifies and encourages semantically novel reasoning behaviors during RLVR. HarA represents each sampled rollout as a distribution over the hidden states and locations of tokens, and computes the Fused Gromov-Wasserstein (FGW) barycenters of all rollouts with the same outcome, capturing the internal reasoning patterns in the latent space under the current policy. The semantic novelty of a reasoning element can then be measured by its contribution to the FGW distance between the current rollout and the barycenter. While solving the FGW formulation is expensive, we introduce an anchor-guided linearization that turns it into a Wasserstein formulation solvable via the Sinkhorn algorithm efficiently. By reweighing token-level advantage of group-based RLVR methods based on the novelty signals, HarA highlights novel reasoning behaviors at flexible granularities to encourage fine-grained LLM exploration. Extensive experiments across three group-based RLVR methods show that our plug-and-play method effectively enhances the exploration of LLMs, outperforming existing methods across diverse reasoning benchmarks.
- 中文摘要
带有可验证奖励的强化学习(RLVR)已被证明能提升大型语言模型(LLMs)在不同推理任务中的推理能力。然而,基于群体的RLVR方法,如GRPO,在相同结果的推广中,会为所有代币分配统一优势。虽然现有研究基于标记位置或熵等局部信号对GRPO的信用分配进行了细化,但它们往往未能捕捉推理行为相对于当前策略的全局语义新颖性。本研究提出了一种基于群体的RLVR方法的分层式信用分配方法,称为HarA,旨在识别并鼓励RLVR期间语义上的新颖推理行为。HarA将每个抽样的展开表示为标记隐藏状态和位置上的分布,并计算所有结果相同的整合Gromov-Wasserstein(FGW)重心,捕捉当前策略下潜在空间的内部推理模式。推理元素的语义新颖性可以通过其对当前滚出与质心之间FGW距离的贡献来衡量。虽然解决FGW公式成本高昂,我们引入锚点引导线性化,将其转化为可通过Sinkhorn算法高效求解的Wasserstein表述。通过基于新颖信号重新权衡基于群组RLVR方法的代币级优势,HarA突出了灵活粒度的新推理行为,鼓励细粒度LLM的探索。在三种基于小组的RLVR方法中进行的大量实验表明,我们的即插即用方法有效提升了对大型语言模型的探索,在多种推理基准测试中优于现有方法。
Large Language Models and Augmented Democracy
大型语言模型与增强民主
- Authors: Jairo Gudiño-Rosero
- Subjects: Subjects:
Computers and Society (cs.CY); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2610.04412
- Pdf link: https://arxiv.org/pdf/2610.04412
- Abstract
Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as intermediaries in augmented democracy, focusing on individual preference representation, collective representation of political organizations, and the vulnerability of those representations to attackers. First, using data from an online experiment in Brazil, I examine whether personalized DTs can predict citizens' preferences for unseen policy proposals. Second, I extend the DT framework from individuals to political organizations. Using Swiss parliamentary data, I build topic-specific knowledge graphs from lawmakers' legislative records and connect them to LLM-based lawmaker agents, which are organized into party-level DTs representing collective positions. Agentic deliberation among these agents tests whether aggregated party representations capture a broader range of intra-party perspectives than official party communications. Finally, I study the vulnerability and robustness of LLM-mediated deliberation against prompt-injection attacks that amplify viewpoints, suppress opinions, or redirect consensus. Using data from a 2023 deliberative experiment in the United Kingdom, I analyze how attack effectiveness varies with the distribution of opinions and rhetorical strategies, and evaluate a pipeline combining injection detection, structured opinion representations, and reinforcement learning to improve resistance. These findings characterize the opportunities and challenges of LLM-based digital twins in augmented democracy, stressing accurate preference representation, faithful aggregation, and robustness to strategic interaction.
- 中文摘要
人工智能使计算智能体能够代表政治偏好并参与集体决策。在本论文中,我探讨基于大型语言模型(LLMs)的数字孪生(DT)作为增强民主中介的机遇与挑战,重点关注个人偏好代表、政治组织的集体代表性以及这些代表对攻击者的脆弱性。首先,利用巴西一项在线实验的数据,我考察了个性化DT是否能预测公民对未公开政策提案的偏好。其次,我将DT框架从个人扩展到政治组织。利用瑞士议会数据,我从立法记录构建主题特定知识图谱,并将其连接到基于LLM的立法者代理,这些代理被组织成代表集体立场的党级DT。这些代理之间的代理性审议测试了聚合政党代表是否比官方党内沟通更广泛地涵盖了党内观点。最后,我研究了LLM介导的审议在放大观点、压制意见或重定向共识的快速注入攻击中的脆弱性和稳健性。利用2023年英国一项审议实验的数据,我分析了攻击效果如何随意见和修辞策略的分布变化,并评估了结合注入检测、结构化意见表达和强化学习以提升抵抗力的流程。这些发现描绘了基于LLM的数字孪生在增强民主中的机遇与挑战,强调偏好代表的准确性、忠实的聚合以及战略互动的稳健性。
CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies
CORE-RL:基于信心的黑盒强化学习策略可靠性评估
- Authors: Santhosh GS, Ananya Ravi, Devika Jay, Abhishek Sarkar, Perepu Satheesh Kumar, Saurav Prakash, Kaushik Dey, Balaraman Ravindran
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04418
- Pdf link: https://arxiv.org/pdf/2610.04418
- Abstract
The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policies infeasible. To address this gap, CORE-RL: Confidence-Oriented Reliability Evaluation of black box RL policy is proposed in this paper. The CORE-RL pipeline introduces a Unified Reliability Metric that formally integrates early task termination and safety constraint violations, preventing unsafe policies from masking failures through premature episode halts. By subjecting the policy to a noise certification envelope of perceptual noise, actuation noise and change in environment dynamics, the pipeline computes the finite-sample Clopper-Pearson bounds on unified reliability metric and Hoeffdings' lower bound on reward and safety cost. The pipeline then defines safe operational design domain to report high-confidence certificates for safety and expected performance. Experiments on continuous control tasks demonstrate the CORE-RL pipeline's ability to automatically reject non-compliant policies and map the safe Operational Design Domain (ODD) of safety-aware policies. Thus CORE-RL provides an evaluation framework towards a quantitative, transparent and reproducible, statistical rationale necessary to safely evaluate, compare, and deploy black box RL solutions.
- 中文摘要
在关键领域部署强化学习(RL)代理之前,必须先配备流水线,以评估强化学习代理与复杂多目标规范的对齐性以及在现实环境漂移下的稳健性。然而,为了保护知识产权,强化学习代理可能以不透明可执行或远程API的形式交付评估,这使得基于策略内部的传统评估技术变得不可行。为弥补这一空白,本文提出了CORE-RL:黑箱RL策略的信心导向可靠性评估。CORE-RL流水线引入了统一可靠性指标,正式整合了早期任务终止和安全约束违规,防止不安全策略通过过早事件停止掩盖失败。通过对策略施加感知噪声、驱动噪声和环境动态变化的噪声认证包络,流水线计算统一可靠性指标上的有限样本Clopper-Pearson界限和Hoeffdings在奖励和安全成本上的下界。随后,流水线定义安全操作设计领域,以报告安全性和预期性能的高信度证书。连续控制任务的实验展示了CORE-RL流水线自动拒绝不合规策略并将安全操作设计领域(ODD)映射安全意识策略的能力。因此,CORE-RL为安全评估、比较和部署黑箱RL解决方案提供了定量、透明且可重复的统计依据。
AgroGround: Multi-Granularity Grounded Recognition in Agriculture
AgroGround:农业中的多粒度基础认可
- Authors: Abdulla Alshehhi, Zongyan Han, Rao Anwer
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04425
- Pdf link: https://arxiv.org/pdf/2610.04425
- Abstract
Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels but rarely link them to image regions, and adding such annotations by hand is costly at scale. We introduce AgroGround, a large-scale dataset for grounded agricultural recognition: identifying plant diseases and other agricultural targets and localizing their image regions. An automated pipeline converts the labels of eight agricultural VQA datasets into annotations for disease lesions and whole objects, producing 794,850 instruction examples. Healthy images provide negative supervision for disease queries, teaching the model to return empty predictions. We fine-tune a shared vision-language model on known-target grounding instructions combined with instructions requiring both recognition and localization. We evaluate predicted identities, regions, joint correctness, and healthy-image abstention on 1,480 human-verified images disjoint from all training data. Grounding-only fine-tuning reduces recognition accuracy from 51.8\% to 29.1\%, while adding recognition-and-localization instructions raises it to 72.6\%. With images and annotations held fixed, combining the two formats raises joint accuracy from 19.2\% to 43.3\% at comparable grounding. Healthy negatives raise abstention on healthy images to 95.0\%, and reinforcement learning improves lesion-level grounding. The resulting 2B model exceeds its annotation teacher in grounding F1 on our benchmark and on the external PlantSeg test set. AgroGround establishes a benchmark for grounded agricultural recognition, measuring joint correctness of identity and localization along with abstention on healthy images. The code is available at this https URL.
- 中文摘要
农业视觉模型通常通过识别或定位进行评估,但可靠诊断需要识别存在的因素并对证据进行定位。农业视觉问答(VQA)数据集带有丰富的语义标签,但很少将其与图像区域关联,手工添加此类注释在大规模上成本较高。我们介绍了AgroGround,一个大规模的基于农业识别数据集:识别植物病害及其他农业目标,并定位其图像区域。一条自动化流程将八个农业VQA数据集的标签转换为病害病变和整体对象的注释,产生794,850个指令示例。健康的图像为疾病查询提供负面监督,教模型返回空预测。我们基于已知目标接地指令与需要识别和定位的指令,微调共享视觉语言模型。我们评估了1480张与所有训练数据不相交的人类验证图像,预测的身份、区域、关节正确性和健康图像的省略率。仅基于基础的微调将识别准确率从51.8%降至29.1%,而添加识别和定位指令则提升至72.6%。在图像和注释保持固定的情况下,结合两种格式在可比基础下将联合准确率从19.2%提升至43.3%。健康阴性使健康图像的缺失率提升至95.0%,强化学习则改善病灶级的接地。最终的2B模型在基准测试和外部PlantSeg测试集中的F1基础化方面超越了其注释教师。AgroGround为扎实的农业识别树立了基准,衡量身份和本地化的联合正确性,以及对健康图片的禁欲。该代码可在此 https 网址获取。
LocusRL: Diagnosing LLM Reward and Policy Interventions in Competitive Games
LocusRL:诊断竞技游戏中的大型语言模型奖励与政策干预
- Authors: Chengyu Luan, Bo Xin, Songyan Guo, Yuxiang Zuo, Ahmed Yazdan, Jiahang Li, Yicheng Liu
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04441
- Pdf link: https://arxiv.org/pdf/2610.04441
- Abstract
Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavior. We introduce LocusRL, a diagnostic framework that connects controlled reward-policy comparisons with audits of reward judgments, signal delivery, optimization objectives, and executed actions. The framework traces performance differences to testable explanations and checks targeted corrections through executable rules and counterfactual replay. Across two evaluation batches covering ten Connect Four training seeds, we uncover seed-dependent reversals in intervention effects and show how tracing actual updates changes their interpretation: historical Qwen training operates through reward-weighted teacher-action likelihood. A separate matched three-seed reward-direction experiment distinguishes sensitivity to a learning signal from its usefulness. With terminal rewards held fixed, a sign-reversed dense oracle yields a 2.8% aggregate win rate, compared with 57.2% for terminal-only training and 46.7% for the positive dense oracle. Thus, a reward can strongly influence learning without improving performance. At the decision level, counterfactual replay verifies a winning correction to a diagnosed action error. Complementary experiments in Leduc and reward-validation studies in Goofspiel extend the analysis to imperfect-information settings, revealing how reference-label definitions and validation-data exposure affect intervention assessment. Together, these findings show why evaluating LLM interventions requires tracing how their outputs become learning signals and actions. LocusRL turns aggregate outcomes into actionable diagnoses and verifiable corrections.
- 中文摘要
大型语言模型可以通过奖励设计和动作选择介入强化学习,但总体表现对这些干预的实际作用并不完整。类似的回报可能隐藏不同的学习机制,而合理的奖励也可能诱发不良行为。我们介绍了LocusRL,这是一个诊断框架,将受控奖励-策略比较与奖励判断、信号传递、优化目标和执行动作的审计连接起来。该框架追踪性能差异至可测试的解释,并通过可执行规则和反事实重放检查有针对性的纠正。在涵盖十个Connect Four训练种子的两批评估中,我们揭示了干预效果中依赖种子的逆转,并展示了追踪实际更新如何改变其解释:历史Qwen训练通过奖励加权教师行动似然实现。另一个匹配的三种子奖励方向实验区分了对学习信号的敏感度与其实用性。在终端奖励固定的情况下,符号反转的密集预言机产生的总胜率为2.8%,而仅终端训练的57.2%和正密集预言机的46.7%。因此,奖励可以强烈影响学习,而不会提升表现。在决策层面,反事实回放验证了对已诊断动作错误的成功纠正。Leduc中的补充实验和Goofspiel中的奖励验证研究将分析扩展到不完美信息环境,揭示了引用标签定义和验证数据暴露如何影响干预评估。这些发现共同说明了评估LLM干预需要追踪其输出如何成为学习信号和行动。LocusRL将汇总结果转化为可操作的诊断和可验证的纠正。
Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?
LLM代理能否自动化文本转语音的强化学习?
- Authors: Xuanjun Chen, Zixiong Su, Hao Shi, Chang Zeng, Kai Li, Jyh-Shing Roger Jang, Hung-yi Lee
- Subjects: Subjects:
Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
- Arxiv link: https://arxiv.org/abs/2610.04488
- Pdf link: https://arxiv.org/pdf/2610.04488
- Abstract
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a collaborative workflow that structures human guidance and agentic execution around a shared workspace, applied to CosyVoice2-0.5B. To measure what the agent automates, we audit its trajectory stage by stage against the published recipe. To measure what it exploits, we score its policies with held-out observers hidden from the agent. Our results show that the agent recovers an underspecified recipe, improves it, and, when gains stall, surveys the literature unprompted and pivots from the LM carrier to the flow carrier, halving Bad cases. However, its autonomy exposes three traps across the data, proxy, and algorithm axes: the held-out set leaks through a channel the contract never reads, a self-shaped reward inflates the proxy where it is scored, and separately tuned policies do not compose additively. These findings show that the binding constraint is measurement rather than reasoning, and can inform the design of harnesses whose contracts read every channel the agent does.
- 中文摘要
尽管强化学习(RL)后训练修复了零样本文本转语音(TTS)的局部片段错误,但得出可行配方仍依赖繁琐的手动调优,且LLM代理能否接管这一研究流程尚不明确。我们用AgenticTTS-Forge这一协作工作流研究这个问题,该流程围绕共享工作空间构建人类指导和代理执行,应用于CosyVoice2-0.5B。为了衡量代理自动化的内容,我们逐阶段审计其轨迹与已发布的配方。为衡量其利用内容,我们用隐藏在代理之外的观察者对策略进行评分。结果显示,代理恢复了一个未明确说明的配方,进行改进,当进展停滞时,会主动调查文献,并从LM载体转向流转载体,将不良案例减半。然而,其自主性暴露了数据、代理和算法轴上的三个陷阱:保留的集合通过合约从未读取的通道泄漏,自形成的奖励使其评分的代理膨胀,以及单独调整的策略不构成加法。这些发现表明,约束力在于测量而非推理,并可指导设计那些合同读取代理所访问的每条通道的约束。
DreamTest: World-Model Surrogates for Search-Based Testing of Deep Reinforcement Learning Agents
DreamTest:基于搜索测试的深度强化学习代理的世界模型替代品
- Authors: Qinghua Xu, Guancheng Wang, Boxi Yu, Liting Lin, Lionel Briand
- Subjects: Subjects:
Software Engineering (cs.SE); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04494
- Pdf link: https://arxiv.org/pdf/2610.04494
- Abstract
Testing deep reinforcement learning (DRL) agents in cyber-physical systems aims to uncover diverse failures before deployment, but each execution can be expensive. Surrogate-assisted testing reduces this cost by learning to predict which test configurations are likely to fail. Prior surrogates treat the system as a black box and predict pass or fail outcomes directly; we instead model how a test unfolds and estimate failure from an imagined episode. We introduce DreamTest, a world-model surrogate for testing DRL agents. DreamTest adapts a recurrent state-space model to learn agent behaviour and environment dynamics from the agent's training log. Given a candidate configuration, imagined rollouts produce a failure score that guides search without executing every candidate in a simulator or real system. We evaluate DreamTest for failure prediction, test generation, and failure diversity on Parking, Humanoid, and DonkeyCar. Mean area under the precision-recall curve (AUPRC) exceeds the strongest baseline by 97%, 12%, and 39%, respectively, and gains on five out-of-distribution test sets reach 145%, 29%, and 44%. Under the same simulator-validation budget, the best "DreamTest + search" combinations find 29%, 22%, and 79% more novel failures on average. Across clusterings with k = 2-40, failures generated with DreamTest cover the most behavioural clusters for almost all k, indicating that DreamTest consistently discovers behaviourally diverse failures.
- 中文摘要
在网络物理系统中测试深度强化学习(DRL)智能体,旨在在部署前发现各种故障,但每次执行都可能成本高昂。代理辅助测试通过学习预测哪些测试配置可能失败,从而降低成本。先验代理将系统视为黑箱,直接预测通过或失败结果;我们则对测试展开进行建模,并从想象中的事件中估算失败。我们介绍了DreamTest,一个用于测试DRL代理的世界模型替代品。DreamTest采用循环状态空间模型,从智能体的训练日志中学习代理行为和环境动态。给定候选配置,想象式部署会产生失败评分,指导搜索,而无需执行模拟器或真实系统中的每个候选配置。我们评估DreamTest在Parking、Humanoid和DonkeyCar上的失败预测、测试生成和失败多样性。精度回忆曲线(AUPRC)下的平均面积分别超过最强基线97%、12%和39%,五个非分布测试集的增益分别达到145%、29%和44%。在相同的模拟器验证预算下,最佳“DreamTest + 搜索”组合平均发现了29%、22%和79%的新失败。在k=2-40的集群中,DreamTest生成的失败覆盖了几乎所有k个的行为集群最多,表明DreamTest持续发现行为多样化的失败。
DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
DiffGate:针对政策提炼的困难门槛教师指导
- Authors: Karn Tiwari, Varnith Chordia, Prathosh A P
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04596
- Pdf link: https://arxiv.org/pdf/2610.04596
- Abstract
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines \emph{which trajectories} receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by $+1.7$ and $+1.8$ points and pass@8 by $+1.6$ and $+5.7$ points, respectively. On mathematics, avg@8 remains within $0.5$ points of GRPO while pass@8 improves by $+1.1$ and $+3.9$ points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.
- 中文摘要
策略提纯(OPD)已成为大型语言模型后训练中广泛使用的范式,通过监督学生自身生成的轨迹,减少传统蒸馏的训练-测试不匹配。然而,现有的OPD目标仍大多为标记-局部且结果无关性,尽管推理质量在轨迹层面决定,但每个前缀处仍优化师生共识。带可验证奖励的强化学习(RLVR),尤其是群体相对策略优化(GRPO),提供了互补的结果级监督,但存在奖励稀疏和粗学分分配的问题。我们表明,OPD和RLVR存在互补盲点:教师信号提供密集的局部指导,但与推广正确性对应较弱,而组相对奖励捕捉任务成功,但提供粗糙的代币级学分,且在所有失败组中消失。我们引入了DiffGate,一种结果门槛目标,结合了GRPO与选择性、有界教师指导。教师监督仅应用于失败轨迹,按组难度调整,并平滑界定,以防止极端师生差异主导优化。验证者因此决定哪些轨迹接受教师指导,教师则在这些轨迹内提供密集的标记级更新方向。在Qwen3-0.6B和Qwen3-1.7B学生中,DiffGate使代码avg@8分别提升了匹配GRPO的+1.7$和$+1.8$点,pass@8提升了$+1.6和$+5.7$。在数学方面,avg@8 距离 GRPO 不到 0.5 美元点,而 pass@8 则分别提升了 $+1.1 和 $+3.9$ 点。总体而言,DiffGate 在四个模型域设置中都提升了pass@8,展示了在我们的评估协议下更优的解覆盖率。
Anticipating the Consequences of Curriculum Decisions with Large Language Models
用大型语言模型预判课程决策的后果
- Authors: Octavio Pappalardo, Nathan Herr, Tim Rocktäschel
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04604
- Pdf link: https://arxiv.org/pdf/2610.04604
- Abstract
Automatic curriculum learning can improve the effectiveness of reinforcement learning by selecting the training experiences presented to the agent over time. Predicting the consequences of such decisions can, however, be difficult. We analyze automatic curriculum learning as a sequential decision-making problem, highlighting a gap between the quantities that determine the value of curriculum decisions and the information captured by local learning signals commonly used to guide them. We then investigate whether Large Language Models (LLMs) can exploit richer information about the learning problem to better anticipate the consequences of curriculum decisions. We introduce a method that combines online learning-progress estimates with LLM-informed estimates of (i) the potential downstream benefits of learning on each task and (ii) whether direct training on a task is currently likely to produce progress. We evaluate the approach on a custom benchmark of 256 textual goals in Craftax under different curriculum objectives. We observe the strongest gains when optimizing for individual target tasks. When optimizing across the full task set, the benefits vary across learners with different mechanisms for cross-task transfer, ranging from modest improvements in learning speed to larger gains that persist through the end of training.
- 中文摘要
自动课程学习可以通过选择随时间呈现给主体的训练体验,从而提高强化学习的有效性。然而,预测此类决策的后果可能较为困难。我们将自动课程学习分析为一个顺序决策问题,突出决定课程决策价值的量与常用的本地学习信号捕获的信息之间的差距。随后,我们探讨大型语言模型(LLMs)是否能利用关于学习问题的更丰富信息,更好地预测课程决策的后果。我们引入了一种方法,将在线学习进展估计与基于LLM的估计相结合,评估(i)每个任务学习的潜在后续收益,以及(ii)直接训练当前是否可能产生进展。我们在Craftax中基于不同课程目标下的256个文本目标的定制基准评估了该方法。我们观察到针对单个目标任务进行优化时,收益最为显著。在全任务集中进行优化时,不同学习者的跨任务转移机制带来的益处各异,从学习速度的适度提升到持续到训练结束时持续的更大收益不等。
Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning
用于从非规范密度抽样的评分校准流,应用于生成在线强化学习
- Authors: Zeyang Li, Yunan Wang, Risheek Garrepalli, Mohammad Ghavamzadeh, Navid Azizan
- Subjects: Subjects:
Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.04696
- Pdf link: https://arxiv.org/pdf/2610.04696
- Abstract
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
- 中文摘要
扩散和流模型为在线强化学习(RL)提供了表达式策略类,实现多模态行为并提升性能。然而,训练这些策略仍然具有挑战性:批评者指定了目标策略为未归一化的玻尔兹曼密度,但未直接提供样本。许多现有方法依赖重要性抽样来构建训练信号,这可能导致高方差、增加计算成本并使训练不稳定。我们提出了分数校准流(SCF),这是一种简单高效的算法,用于训练生成模型从未归一化密度中采样,无需重要性采样或通过采样轨迹进行反向传播。我们通过强制自一致性来学习所需流程,绕过目标后验均值估计。通过结合利用规定的目标评分和流量匹配结构,我们将这些自一致性要求作为分数校准的最优条件,先针对终端密度,随后针对可训练速度场。我们证明了它们的唯一解分别是目标密度和条件流匹配(CFM)在目标样本可用时恢复的理想流量模型。我们将速度条件表述为不动点方程,并利用其条件-期望结构构建一个停止梯度目标来强制执行。最终的训练过程保留了CFM中可扩展的样本-插值-回归结构,尽管没有目标样本,使用由当前流生成的端点。对于在线强化学习,批判梯度在生成动作时提供目标分数,提供了直接的演员训练方法。强化学习基准测试的实验表明,SCF能够匹配甚至提升最先进的生成策略基线,同时显著缩短训练时间。
Learning to Clarify Underspecified Intents Under Limited Interaction
在有限互动下学习澄清未明确的意图
- Authors: Pranav M R, Manuel Cherep, Pattie Maes, Nikhil Singh
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.04719
- Pdf link: https://arxiv.org/pdf/2610.04719
- Abstract
AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information problem: the assistant should acquire information whose absence causes the greatest avoidable loss in user utility. This is rarely known ex ante; rather, assistants must predict it in order to optimally allocate limited user interactions. We instantiate this problem in image generation and derive a reinforcement learning framework using multi-turn simulated users to maximize utility recovery under uncertainty. In a preregistered study with 456 interactive sessions across 76 human participants, this helped users significantly better match reference images with significantly fewer questions, less total interaction time, and lower cost. This points toward a simple and scalable framework for training language model assistants to better disambiguate user intent by asking more informative questions.
- 中文摘要
AI助手收到的请求中,会遗漏实现良好结果所需的信息,例如用户的偏好或目标。然后,他们必须进行推测或请求更多信息,才能继续前行。我们将此重新构想为信息价值问题:助手应获得那些缺失导致用户效用损失最大且可避免的信息。这一点很少事先被揭示;相反,助手必须预测这些信息,以优化有限的用户互动。我们将这一问题实例化为图像生成,并基于多回合模拟用户推导出强化学习框架,以最大化在不确定性下效用回收。在一项预注册的研究中,涉及456个互动会话,76名真人参与者,这帮助用户显著更好地匹配参考图像,问题数量显著减少,交互时间更短,成本更低。这表明这是一个简单且可扩展的框架,用于训练语言模型助手,通过提出更具信息性的问题更好地消除用户意图。
PatternDex: Learning Interaction Patterns to Guide Reinforcement Learning of Bimanual Dexterous Manipulation of Articulated Objects
PatternDex:学习互动模式以指导双手灵活操作关节物体的强化学习
- Authors: David Minkwan Kim, Runfa Blark Li, Beckham Po-Ju Lee, Nikolay Atanasov, Truong Nguyen
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.04765
- Pdf link: https://arxiv.org/pdf/2610.04765
- Abstract
In this paper, we develop a method that enables bimanual dexterous hands to manipulate articulated objects with a high success rate without suffering from an embodiment gap. We observe that the correlation between hand motions and object motions is dictated by the object rather than the hands and can be learned from human-object demonstrations. Based on this observation, we propose PatternDex, a method that learns this correlation and represents it as a token sequence, which we call an interaction pattern. From this pattern, PatternDex estimates the wrist motions and contact points that fit the target robot, and then trains a reinforcement learning policy that exploits these estimates as guidance. Since the guidance fits the target embodiment, the policy explores only the actions that the target robot can execute and thus achieves high success rates. PatternDex also requires only simple fine-tuning to train a new robot, since it can reuse the learned interaction pattern. We evaluate PatternDex with bimanual dexterous hands on human demonstrations from the ARCTIC dataset. PatternDex achieves, on average, a 92.8% success rate with Allegro hands, while the state-of-the-art baseline achieves 52.2%. Also, it achieves success rates above 70% with three other robot hands after fine-tuning alone. Furthermore, we verify that the learned policy transfers well to a real-world task of opening a microwave. Videos and additional results are available at this https URL
- 中文摘要
本文开发了一种方法,使双手灵巧手能够高成功率操作关节物体,且不会出现具身缺口。我们观察到手部动作与物体动作之间的相关性由对象决定,而非手部,且可通过人与物的演示学习。基于这一观察,我们提出了PatternDex方法,该方法学习该相关性并将其表示为标记序列,称为交互模式。基于该模式,PatternDex估计适合目标机器人的手腕动作和接触点,然后训练强化学习策略,利用这些估计值作为指导。由于引导符合目标具象,策略仅探索目标机器人能执行的动作,从而实现高成功率。PatternDex还只需简单微调即可训练新机器人,因为它可以复用所学的交互模式。我们通过来自北极数据集的双手灵巧手操作人体演示来评估PatternDex。PatternDex在Allegro手部下平均成功率为92.8%,而最先进基线则为52.2%。此外,仅通过微调,PatternDex在另外三只机器人手上成功率超过70%。此外,我们验证了所学策略在打开微波炉这一现实任务中表现良好。视频和更多结果可在此 https URL 获取
Multi-Agent Spectrum Sharing
多智能体频谱共享
- Authors: Job Elliott, Graduate Student Member, IEEE, Justin G. Metcalf, Golnaz Habibi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04802
- Pdf link: https://arxiv.org/pdf/2610.04802
- Abstract
This project explores how multiple cognitive radars can learn to share limited wireless spectrum with other radio users without interfering with one another. Using machine learning (ML), each device independently decides where and how widely to transmit within a fixed 100 MHz band. The system analyzes real or simulated signal activity to detect which parts of the spectrum are currently in use and which are open. Based on this information, the devices adapt their transmission choices to avoid crowded frequencies while making efficient use of available space. The goal is to develop a flexible, scalable approach to spectrum sharing that could support future wireless communication systems. Experimental results using both over-the-air software-defined radio (SDR) recordings and simulated environments demonstrate that the proposed meta-learning approach consistently balances competing objectives better than conventional reinforcement learning (RL) methods in multi-agent spectrum-sharing scenarios. Across five multi-agent benchmark environments, our proposed method achieved the highest average reward among the primary baseline algorithms while simultaneously maintaining low collision rates and stable transmission behavior.
- 中文摘要
本项目探讨了多个认知雷达如何学习与其他无线电用户共享有限的无线频谱,而不互相干扰。利用机器学习,每个设备独立决定在固定的100 MHz频段内传输的位置和多广。系统分析真实或模拟的信号活动,以检测频谱中哪些部分正在使用,哪些是开放的。基于这些信息,设备调整传输选择,以避免拥挤频率,同时高效利用可用空间。目标是开发一种灵活、可扩展的频谱共享方法,以支持未来的无线通信系统。利用空中软件定义无线电(SDR)录音和模拟环境的实验结果表明,所提元学习方法在多智能体频谱共享场景中,比传统强化学习(RL)方法更有效地平衡了多个竞争目标。在五个多智能体基准环境中,我们提出的方法在主要基线算法中获得了最高的平均奖励,同时保持了低碰撞率和稳定的传输行为。
CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery
CURIO:以好奇心驱动的测试时间学习,实现开放式发现
- Authors: Tao Feng, Fangxu Yu, Zijie Lei, Jiaru Zou, Changjiang Jiang, Yi Yan, Jiaxuan You, Pan Lu
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.04851
- Pdf link: https://arxiv.org/pdf/2610.04851
- Abstract
Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement learning (RL) enables such adaptation; however, strongly favoring high-reward trajectories may suppress low-reward yet potentially promising directions too early. We introduce CURIO, a curiosity-driven test-time learning framework that complements task feedback with an Intrinsic Curiosity World Model (ICWM). The ICWM learns transitions in the policy's hidden-state representation and supplies prediction-error bonuses at sampled tokens outside the policy's top-k choices. Epoch normalization and an annealed weight regulate their contribution to the policy update. On six mathematical discovery tasks and single-cell denoising with Qwen3 backbones from 8B to 235B, three-run means improve over a matched task-only RL control on five mathematical objectives, match the best reported performance on Circle Packing, and improve denoising Score and mean squared error (MSE) on both held-out corpora at every tested scale. Relative gains reach 18.3% on Hadamard and 10.8% on denoising Score. Code-diversity measurements show greater structural variation among generated programs, supporting curiosity as a complementary exploration signal for learning in open-ended discovery.
- 中文摘要
开放式发现需要通过反复尝试学习,同时继续探索价值尚未显现的方向。使用冻结大型语言模型(LLM)进行搜索可以在上下文中重用之前的解法,但无法从测试问题的成功和失败中更新模型。强化学习(RL)支持这种适应;然而,强烈偏向高奖励轨迹可能会过早抑制低回报但潜在有前景的方向。我们介绍CURIO,一种基于好奇心驱动的测试时学习框架,它与内在好奇心世界模型(ICWM)补充任务反馈。ICWM学习策略隐藏状态表示中的转移,并在策略前k选择之外的抽样令牌提供预测错误加成。历元归一化和退火权重调节它们对策略更新的贡献。在6个数学发现任务和用Qwen3骨干从8B到235B的单胞去噪中,三次运行均值在五个数学目标上优于匹配的仅任务强化学习对照,匹配圆圈打包的最佳报告表现,并在所有测试量表下提升了两个保留语料库的去噪分数和均方误差(MSE)。Hadamard的相对提升达到18.3%,去噪Score的10.8%。代码多样性测量显示生成程序间结构性差异更大,支持好奇心作为开放式发现学习的补充探索信号。
Rewrite What Matters: Adaptive Multilingual Query Rewriting for Reasoning via Agentic Reinforcement Learning
重写重要内容:通过智能强化学习实现自适应多语言查询重写推理
- Authors: Rui Qi, Yufeng Chen, Yunlong Liang, Chuan Meng, Sijin Lu, Ge Shi, Jinan Xu, Fandong Meng, Kaiyu Huang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.04899
- Pdf link: https://arxiv.org/pdf/2610.04899
- Abstract
In multilingual scenarios, queries with equivalent semantics but in different languages could guide the model into different reasoning trajectories, leading to performance disparities. To mitigate this gap, previous studies typically apply a one-size-fits-all query rewriting strategy, such as translation, which overlooks the fact that different scenarios require diverse types of semantic transformations. In this paper, we propose mRewriter-R1, an agentic multilingual query rewriting framework with reinforcement learning. Unlike single-turn rewriting, mRewriter-R1 formulates multilingual query rewriting as a multi-turn sequential decision-making process, where the model dynamically performs multi-aspect optimization through adaptive operator selection. Experimental results demonstrate that mRewriter-R1 outperforms all strong multilingual rewriting baselines on different large reasoning backbones. Further analyses show that the learned policy can adaptively decide on rewriting operators according to query characteristics, exhibiting strong generalization ability across diverse reasoning tasks, and plug-and-play compatibility with heterogeneous reasoning language models.
- 中文摘要
在多语言场景中,语义相同但语言不同,可能导致模型走向不同的推理轨迹,导致性能差异。为弥补这一差距,先前研究通常采用一刀切的查询重写策略,如翻译,忽视了不同场景需要多种语义转换的事实。本文提出mRewriter-R1,一种具强化学习的代理多语言查询重写框架。与单回合重写不同,mRewriter-R1将多语言查询重写表述为多回合顺序决策过程,模型通过自适应算子选择动态执行多方面优化。实验结果表明,mRewriter-R1在不同大型推理骨干上表现优于所有强多语言重写基线。进一步分析表明,所学策略能够根据查询特性自适应地决定是否重写运算符,展现出在多种推理任务中强强的泛化能力,并且与异构推理语言模型实现即插即用兼容性。
Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning
残余视觉学分优化:多模态强化学习的保守证据路由
- Authors: Lin Qiu, Yao Liu, Diyi Hu, Hanqing Zeng, Onur Gungor, Chujie Chen, Jiayi Liu, Jianyu Wang, XueLin Zheng
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.04918
- Pdf link: https://arxiv.org/pdf/2610.04918
- Abstract
Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing problem. A controlled visual intervention yields a per-token evidence response; robust within-trajectory coordinates remove incidental scale; and a budgeted entropic router distributes a fixed amount of sequence utility according to perceptual dependence. A residual support path guarantees positive credit at every valid position, and an analytic correction restores the prescribed credit mass exactly. The resulting field is selective, bounded, full-support, and invariant to response-local score shifts, and recovers hard token selection as a limiting case. Across four model families and seven reasoning benchmarks, RVCO improves accuracy over strong RLVR baselines while maintaining late-stage optimization stability, corruption robustness, and competitive training cost. Rewards, rollouts, and the group-relative advantage estimator are unchanged; only the geometry of token-level credit differs.
- 中文摘要
带有可验证奖励的强化学习可扩展多模态推理,但结果奖励说明轨迹的价值,而非该价值在产生该轨迹的决策中如何分布。我们引入残差视觉信用优化(RVCO),将代币信用视为守恒的路由问题。受控视觉干预产生每个代币的证据响应;稳健的轨迹内坐标消除偶发尺度;预算熵路由器根据感知依赖分配固定数量的序列效用。残差支持路径保证每个有效位置的正信用,分析修正准确恢复规定的信用质量。所得字段是选择性、有界、全支持且对响应局部得分变化不变的,并恢复硬代币选择作为极限情况。在四个模型家族和七个推理基准中,RVCO在保持后期优化稳定性、腐败鲁棒性和竞争训练成本的同时,提升了强有力的RLVR基线精度。奖励、推广和群体相对优势估计器保持不变;仅代币级信用的几何形状有所不同。
PWM: Personalized World Models with Online Reinforcement Learning
PWM:带有在线强化学习的个性化世界模型
- Authors: Zhexin Lou, Guancheng Lu, Zeyu Zhang, Yi Zhang, Yang Zhao, Hao Tang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.04920
- Pdf link: https://arxiv.org/pdf/2610.04920
- Abstract
Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We introduce Personalized World Models (PWM), a framework for customizing interactive world models from short scene videos through online reinforcement learning. In PWM, the support trajectory and its associated controls provide reward feedback on continuations sampled from the current policy. In the GRPO instantiation, group-relative optimization updates a compact LoRA adapter using a unified reward for scene appearance, visual continuity, and motion, while base-policy anchoring regularizes changes to the pretrained generation prior of a frozen Yume-5B backbone. The same adaptation procedure is applied across real and rendered environments. We also instantiate PWM with DiffusionNFT as an alternative reward-guided optimization method for learning the scene-specific adapter. We also introduce PWM-Bench, comprising 150 customization tasks across Indoor, Outdoor, and Gaming, with paired evaluation on held-out continuations. The GRPO and DiffusionNFT instantiations of PWM improve customization over native Yume in 71.3% and 65.3% of the evaluated scenes, respectively, with positive mean gains across all three domains. For the GRPO instantiation, matched SFT comparisons further demonstrate higher mean customization gains and better mean image-quality scores in every domain, while retaining frame-level visual quality close to the pretrained model.
- 中文摘要
预训练世界模型可以生成多样化的环境,但用户通常希望探索由自己视频指定的特定场景。这需要在保持动作条件生成质量的前提下,学习场景的视觉身份。我们引入了个性化世界模型(PWM),这是一个通过在线强化学习定制短场景视频互动世界模型的框架。在PWM中,支持轨迹及其相关控件对当前策略中采样的续写提供奖励反馈。在GRPO实例化中,群相对优化通过统一的场景外观、视觉连续性和运动奖励更新紧凑的LoRA适配器,而基础策略锚定则规范了冻结Yume-5B骨干链预训练生成的变更。相同的适应过程应用于真实和渲染环境。我们还用DiffusionNFT实例化PWM,作为学习场景特定适配器的替代奖励引导优化方法。我们还引入了PWM-Bench,涵盖室内、室外和游戏领域的150个定制任务,并对未完成的续写进行配对评估。GRPO和DiffusionNFT的PWM实例分别在评估场景中71.3%和65.3%的自定义提升,均值均为正。对于GRPO实例化,匹配SFT的比较进一步显示了更高的平均定制提升和每个领域的平均图像质量评分,同时帧级视觉质量保持接近预训练模型。
A Unified Dynamics Framework for Reinforcement Learning and Classical Control of a Six-DOF Pipeline-Tracking ROV in NVIDIA Isaac Sim
用于增强学习和经典控制的统一动力学框架,用于NVIDIA Isaac Sim中六自由度流水线跟踪ROV
- Authors: Cheng Siong Chin, M. Venkateshkumar, Jianhua Zhang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.04949
- Pdf link: https://arxiv.org/pdf/2610.04949
- Abstract
Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training. This paper presents a pipeline-tracking architecture for a six-degree-of-freedom remotely operated vehicle (ROV) in which one Universal Scene Description (USD) scene supplies the real BlueROV2-Heavy mass, added-mass, damping, buoyancy, and thruster parameters to both halves of the system: a vectorized NumPy implementation of Fossen's marine-craft equations, and an interactive NVIDIA Isaac Sim deployment applying the identical equations as PhysX forces at every step. The Coriolis-centripetal term is the primary dynamics model in both branches; a controlled ablation on PPO and TRPO shows that including it does not destabilize either algorithm and modestly improves tracking, about 19 percent lower standoff RMS error for TRPO. Five reinforcement learning algorithms, PPO, soft actor-critic, TD3, DDPG, and TRPO, are trained against one environment, reward, and randomized evaluation harness through a checkpoint-compatibility layer scoring any policy with the same code. The pipeline is extended with six classical baselines, PID, sliding-mode, fuzzy logic, feedback linearization, model predictive control, and an adaptive neuro-fuzzy inference system, driven by the same guidance geometry and thruster allocation as the learned policies. Under Coriolis-enabled dynamics, PPO, TRPO, and feedback linearization reach the strongest combination of 100 percent success and competitive accuracy; PID, fuzzy control, and the neuro-fuzzy baseline also reach 100 percent success with looser tracking; DDPG and TD3 each show a specific, explainable failure mode rather than a general weakness of off-policy learning; and classical control remains a strong baseline against the best learned policies.
- 中文摘要
水下飞行器的强化学习控制器通常基于一种物理表现训练,并部署于另一种物理表现,因此报告的性能不一定能描述训练外的行为。本文提出了一种六自由度遥控潜水器(ROV)的流水线跟踪架构,其中一个通用场景描述(USD)为系统两半提供真实的BlueROV2-Heavy质量、加质量、阻尼、浮力和推进器参数:Fossen海洋-船舶方程的矢量化NumPy实现,以及在每一步应用相同方程的交互式NVIDIA Isaac Sim部署。科氏向心项是两分支的主要动力学模型;对PPO和TRPO进行受控消融显示,加入该算法不会使任何算法不稳定,且适度提升跟踪性能,TRPO的距离均方根误差降低约19%。五种强化学习算法PPO、软演员批判、TD3、DDPG和TRPO通过检查点兼容性层对同一代码的策略进行训练,基于同一环境、奖励和随机评估工具进行训练。流水线扩展为六条经典基线、PID、滑动模式、模糊逻辑、反馈线性化、模型预测控制和自适应神经模糊推断系统,这些系统由与所学策略相同的指导几何和推进器分配驱动。在Coriolis驱动的动力学下,PPO、TRPO和反馈线性化达到了100%成功率和竞争准确性的最佳组合;PID、模糊控制和神经模糊基线在较松散的跟踪下也能达到100%成功率;DDPG和TD3各自显示出特定且可解释的失败模式,而非策略外学习的普遍弱点;经典控制仍然是对最佳策略的强有力基线。
How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation
教师应如何准备?关于学生诱导的政策提炼状态的现实学习
- Authors: Xiaoyu Ma, Haoyue Liu, Zhichao Wang, Jionghao Zhu, Xiaoying Tang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.04950
- Pdf link: https://arxiv.org/pdf/2610.04950
- Abstract
On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effectively? Prior work shows that when student prefixes follow reasoning paths that differ from the teacher's own or contain errors, teachers can be less accurate when continuing from these prefixes than when solving problems independently. To this end, we propose Prep-OPD, which uses reinforcement learning (RL) before distillation to train the teacher to adapt to the student's existing reasoning state and correct course when errors arise. Training optimizes teacher continuations from fixed student prefixes using final-answer correctness as the reward. The prepared teacher then trains the student through trajectory guidance and token-level supervision. We evaluate Prep-OPD on eight mathematical reasoning benchmarks, using Qwen3-4B-Instruct-2507 as the teacher and Qwen3-0.6B and Qwen3-1.7B as students. With the 4B teacher and 1.7B student, Prep-OPD improves average accuracy over standard OPD and the strongest baseline, Relay-OPD, by 8.28 and 2.30 percentage points, respectively. Controlled experiments further show that teacher RL conditioned on student-generated prefixes yields higher student accuracy than problem-start teacher RL with and without handoff on Qwen3-1.7B. Reusing the same prepared teacher also improves Qwen3-0.6B.
- 中文摘要
策略提炼(OPD)通过令式级教师对学生生成轨迹的监督,提升小型语言模型的推理能力。然而,擅长独立解决问题的教师能否有效引导学生推理呢?先前研究表明,当学生前缀遵循与教师自身推理路径不同或包含错误时,教师从这些前缀继续前缀时的准确性可能低于独立解决问题时。为此,我们提出了Prep-OPD,即在提炼前使用强化学习(RL)训练教师适应学生现有推理状态,并在出现错误时纠正方向。培训通过奖励以最终答案正确性为奖励,优化教师从固定学生前缀的延续。准备教师随后通过轨迹指导和标记级监督培训学生。我们以Qwen3-4B-Instruct-2507为教师,Qwen3-0.6B和Qwen3-1.7B作为学生,基于八个数学推理基准评估Prep-OPD。对于4B教师和1.7B学生,Prep-OPD相比标准OPD和最强基线Relay-OPD分别提高了平均准确率8.28个百分点和2.30个百分点。受控实验进一步表明,以学生生成前缀为条件的教师RL比问题起始型教师RL在Qwen3-1.7B上交接或未交接时,学生准确率更高。重复使用同一准备老师也提升了Qwen3-0.6B。
AlphaPADI: Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion
AlphaPADI:通过池感知层级离散扩散实现公式化Alpha发现
- Authors: Yanzheng Jin, Pengyang Shao, Yunshan Ma, Haowen Pan, Naixin Zhai, Chen-Hui Song, Fei Shen, Kenji Kawaguchi
- Subjects: Subjects:
Computational Engineering, Finance, and Science (cs.CE); Computational Finance (q-fin.CP)
- Arxiv link: https://arxiv.org/abs/2610.04959
- Pdf link: https://arxiv.org/pdf/2610.04959
- Abstract
Formulaic alpha discovery seeks symbolic expressions that predict cross-sectional asset returns. In deployment, multiple formulas are combined into an alpha pool, where each formula is valued through the complementary information it contributes to joint predictive performance. While Reinforcement Learning and Generative Flow Networks have emerged as promising paradigms for generating formulaic alphas, existing frameworks face three related challenges. First, generating formulas individually leaves pool context and inter-formula complementarity outside the generative state. Second, formula-wise generation lacks a unified mechanism for preserving and revising structures at different levels. Third, pool-level rewards jointly reflect predictive performance and redundancy but cannot be differentiated directly through symbolic evaluation to train the generator. To overcome these challenges, we introduce AlphaPADI (Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion), a novel framework built around three components: (1) grammar-constrained buffer initialization that constructs syntactically valid pool candidates, (2) pool-aware hierarchical diffusion that reconstructs complete pools at multiple structural scales under the current pool context, and (3) reward-guided pool refinement that evaluates joint predictive performance and inner diversity, updates the elite buffer, and trains the reverse model through reconstruction and preference learning. Empirical results on the Chinese and U.S. stock markets demonstrate that AlphaPADI outperforms the evaluated baselines in both predictive and portfolio performance, thereby validating pool-aware generation as an effective framework for automated alpha discovery.
- 中文摘要
公式化α发现寻求符号表达,以预测横断面资产回报。在部署中,多个公式被组合成一个α池,每个公式通过其对联合预测性能贡献的互补信息进行估值。虽然强化学习和生成流网络已成为生成公式化α的有前景范式,但现有框架面临三个相关挑战。首先,单独生成公式使池上下文和公式间互补性超出生成状态。其次,按公式生成缺乏统一机制来保存和修订不同层级的结构。第三,池级奖励共同反映预测性能和冗余性,但无法通过符号评估直接区分以训练生成器。为克服这些挑战,我们引入了AlphaPADI(通过池感知层级离散扩散实现公式Alpha发现),这是一个围绕三个组成部分构建的新框架:(1)语法约束缓冲区初始化,构建语法有效的池候选,(2)池感知层级扩散,在当前池语境下多个结构尺度重建完整池,(3)奖励引导池细化,评估联合预测性能和内部多样性,更新精英缓冲区,并通过重构和偏好学习训练逆向模型。中国和美国股市的实证结果表明,AlphaPADI在预测和投资组合表现上均优于评估基线,验证了池感知生成作为自动Alpha发现的有效框架。
Outcome-Guided On-Policy Self-Distillation
以结果为导向的政策自我提炼
- Authors: ZheXu Wang, Mao-Lin Luo, Yankun Hong, Zi-Hao Zhou, Bo Ye, Jian Zhao, Xialiang Tong, Min-Ling Zhang, Tong Wei
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05070
- Pdf link: https://arxiv.org/pdf/2610.05070
- Abstract
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
- 中文摘要
策略自提纯(OPSD)提供更密集的代币级监督和比可验证奖励强化学习(RLVR)更好的计算效率。然而,这种更密集的监督可能引入大量噪声和训练不稳定性。现有改进通常依赖高方差每个令符统计量,并引入额外的超参数和权衡。基于RLVR中的优势表述,我们从相同视角分析OPSD目标,纳入结果正确性信号。我们发现,原版OPSD对错误轨迹施加的惩罚不足且奖励过高,因为它无论结果正确性如何都应用固定的发散目标。此外,教师监督的可靠性与轨迹结果和部署过程中累计平均教师熵相关。基于这些观察,我们提出了结果引导策略自蒸馏(OG-OPSD),该方法根据二元结果奖励和累计平均教师熵动态调整发散目标和蒸馏位置。大量实验表明,OG-OPSD在1.7B、4B和8B尺度以及Qwen3-VL-2B尺度上,持续提升普通OPSD和多强基线在数学推理、多模态推理及分布外任务中的表现。
Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning
在线目标条件强化学习的方向条件政策
- Authors: S K Swaminathan, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.05087
- Pdf link: https://arxiv.org/pdf/2610.05087
- Abstract
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.
- 中文摘要
对比强化学习(CRL)学习估计目标可达性的表征,但其策略仍基于原始目标,因此不直接利用批评者编码的几何结构。我们介绍方向条件策略(DCP),这是一种基于CRL小幅修改的方法:DCP在在线训练中选择先前访问过的状态作为路径点,并根据其在表示空间中的方向和距离来设定策略条件。部署时,DCP将相同界面直接应用于最终目标,无需选择或规划路径点。在九个导航和操作任务中,DCP在七个任务中取得更高的最终成功率,七个任务中在目标附近停留的时间更长。受控迷宫实验进一步表明,DCP更准确地捕捉最短路径几何,且提供的方向因果地影响行为者行为。我们指出路径点覆盖率和排名是探索的限制,并证明学习到的候选生成能提升两个受控迷宫中的目标达成。
Small Agents with Semantic Search: Efficient Multilingual Code Localization
具语义搜索的小型代理:高效的多语言代码本地化
- Authors: Maxence Lasbordes, Aarush Sinha, Raphael Sourty, Amélie Chatelain, Djamé Seddah
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.05099
- Pdf link: https://arxiv.org/pdf/2610.05099
- Abstract
Locating relevant files from natural-language requests is a core subtask for agents operating over code repositories. We investigate whether this task can be delegated to compact, specialized models to enable on-device search while reducing the token usage, latency, and inference cost of larger agents. We show that semantic search improves file localization, with gains in accuracy, cross-language transfer, and inference efficiency. To study this setting, we introduce a training framework for file-localization agents built around ColGREP, a local semantic search tool based on late-interaction retrieval models. Our recipe combines weighted supervised fine-tuning on teacher trajectories, assigning turn-level credit based on retrieval outcomes, with reinforcement learning on localization quality. We train three model families with fewer than two billion parameters to formulate search queries, inspect retrieved content, and identify relevant files. On localization tasks derived from SWE-bench Lite and Multi-SWE-bench Flash, ColGREP-equipped agents substantially improve over their base models and outperform corresponding GREP-based agents. In addition to improving localization accuracy, ColGREP reduces mean end-to-end trajectory latency by 44.1\% on CPU while using 29.1\% fewer tokens, and enables better generalization to programming languages unseen during fine-tuning. These results suggest that compact, tool-specialized localization agents can provide an efficient interface between natural-language requests and large codebases.
- 中文摘要
从自然语言请求中定位相关文件是代理在代码仓库中操作的核心子任务。我们研究该任务是否可以委托给紧凑、专业化的模型,以实现设备端搜索,同时降低令牌使用、延迟和推理成本,尤其是大型代理。我们证明语义搜索能提升文件本地化,提升准确性、跨语言传输和推理效率。为研究这一背景,我们引入了基于晚期交互检索模型的本地语义搜索工具ColGREP构建的文件本地化代理训练框架。我们的方案结合了对教师轨迹的加权监督微调、基于检索结果分配回合级功劳,以及本地化质量的强化学习。我们训练了三个参数少于20亿的模型家族,用于构建搜索查询、检查检索内容并识别相关文件。在源自SWE-bench Lite和Multi-SWE-bench Flash的本地化任务中,ColGREP代理相较基础模型有显著提升,并优于相应的基于GREP的代理。除了提升本地化准确性外,ColGREP还将CPU端到端平均轨迹延迟降低了44.1%,同时使用了29.1%的令牌数,并实现了对微调中未发现的编程语言的更好推广。这些结果表明,紧凑且专门的工具的本地化代理能够在自然语言请求与大型代码库之间提供高效的接口。
Arithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning
算术演员头与非分布强化学习的训练稳定
- Authors: Yifan Zhang, Liang Zheng
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05143
- Pdf link: https://arxiv.org/pdf/2610.05143
- Abstract
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.
- 中文摘要
强化学习(RL)策略在分布外(OOD)幅度变化下可能恶化。从软演员批评(SAC)及其贝叶斯遗忘分段稳健(BAPR)前身出发,我们研究因果符号BAPR(CS-BAPR)家族。实用方法结合了六种训练稳定设置和替代角色头:神经加法单元(NAU)与神经乘法单元(NMU)启发的二次修正、Kolmogorov-Arnold Network(KAN)或多层感知器(MLP)带有整流线性单元(ReLU)或双曲切激活。
The Law of DeepSeek
深搜法则
- Authors: Jyh-An Lee, Xuan Sun
- Subjects: Subjects:
Computers and Society (cs.CY)
- Arxiv link: https://arxiv.org/abs/2610.05238
- Pdf link: https://arxiv.org/pdf/2610.05238
- Abstract
Amid the intensifying competition in artificial intelligence between the United States and China, the emergence of the DeepSeek-R1 model has sent significant ripples through the technology sector, capital markets, and policy circles. This Article offers a comprehensive analysis of legal and policy landscape surrounding DeepSeek, drawing upon its key technical features-including reinforcement learning, mixture-experts architecture, multi-head latent attention mechanism, knowledge distillation, and open-source approach. Firstly, DeepSeek's success raises crtitical questions about the efficacy of the United States' increasingly robust export control measures on chips and semiconductors, components essential for training large language models. Secondly, akin to Chinese technology giants such as TikTok and Huawei, DeepSeek is confronted with information security scrutiny in the United States and other jurisdictions. Thirdly, despite its domain-specific capabilities and competitive API pricing, DeepSeek has faced criticism for producing outputs laden with political and ideological biases, igniting debates over free speech and censorship. Further more, this Article delves into intellectual property and contractual concerns stemming from the knowledge distillation technique employed by DeepSeek. Ultimately, it concludes that geopolitical consideration will continue to exert a profound influence on the legal challenges and prospective solutions related to the DeepSeek models.
- 中文摘要
在美中人工智能竞争日益激烈的背景下,DeepSeek-R1模型的出现在科技领域、资本市场和政策圈掀起了巨大反响。本文全面分析了围绕DeepSeek的法律和政策环境,结合其关键技术特征——包括强化学习、专家混合架构、多头潜在注意力机制、知识提炼和开源方法。首先,DeepSeek的成功引发了关于美国对芯片和半导体日益严格的出口管制措施有效性的质疑,这些措施是训练大型语言模型的关键组件。其次,类似于TikTok和华为等中国科技巨头,DeepSeek在美国及其他司法管辖区面临信息安全审查。第三,尽管DeepSeek具备领域特定能力和具有竞争力的API定价,但因其产出充斥政治和意识形态偏见的产出而受到批评,引发了关于言论自由和审查的争论。此外,本文深入探讨了DeepSeek所采用的知识提炼技术引发的知识产权和合同问题。最终,文章得出结论,地缘政治考量将继续对DeepSeek模型相关的法律挑战和潜在解决方案产生深远影响。
MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation
MGPO:任务感知数据集蒸馏中的流形引导扩散比对
- Authors: Yunyi Chen, Chenru Wang, Xinyi Ye, Zexin Zheng, Chi Zhang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05252
- Pdf link: https://arxiv.org/pdf/2610.05252
- Abstract
Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.
- 中文摘要
基于扩散的数据集蒸馏(DD)存在一个根本的客观不匹配问题:似然驱动扩散模型优先考虑密度近似,而非下游任务所需的判别决策边界。除了语义不匹配外,仅依赖密度还会导致几何覆盖率丧失,生成的样本崩溃为少数高密度模式,无法覆盖流形的结构多样性。我们提出了流形引导策略优化(MGPO),将DD重新表述为一个多目标强化学习问题,并通过像素空间的判别奖励和由类别极小生成树(MST)引导的潜在空间几何奖励实现双空间对齐。判别性奖励强化类别可分离性,而基于MST的几何奖励鼓励生成的潜在变量覆盖每个类别的稀疏几何骨架,同时解决两种失败模式。我们还提供了理想化分析,支持基于MST的奖励,包括无关数据集大小的豪斯多夫近似界限和子抽样界限。奖励模块化设计通过替代冻结任务奖励模型,扩展到物体检测和分割等结构化任务。大量实验表明MGPO持续优于现有方法,包括在低预算条件下分割获得+8.0%的mIoU收益。
Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes
有证据的答案:一致性意识的接地视觉问答,适用于路边交通场景
- Authors: Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningwei Ouyang, Shaofeng Liang, Heyi Lin, Jinjing Zhu, Yang Shi, Dongming Wu, Daizong Liu, Henghui Ding, Hui Xiong
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
- Arxiv link: https://arxiv.org/abs/2610.05274
- Pdf link: https://arxiv.org/pdf/2610.05274
- Abstract
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and boxes separately leave this failure unpenalized. We trace the mismatch to the conventional answer-then-ground factorization, which commits to a numerical claim before any object is enumerated. To measure it, we build RoadSceneVQA-G, a benchmark of 34.7K question-answer pairs in which every free-form answer is linked to the set of boxes that witnesses it, and we propose the Answer-Grounding Consistency (AGC) evaluation suite. To address it, we introduce Enumerate-then-Answer (EtA), which reverses the generation order so that answer-evidence agreement becomes a property of the output structure, and Enumeration-Consistent Policy Optimization (ECPO), a reinforcement learning stage that uses the union of multiple rollouts as a recall teacher without ground-truth boxes. EtA raises say-point consistency from 26.6\% to 93.7\% and grounding F1 from 52.2\% to 73.0\%, and ECPO further increases F1 to 75.6\% without per-box supervision. On gRefCOCO, the same framework outperforms the strongest compared method, indicating that it transfers beyond traffic scenes. The project is available at \url{this https URL}.
- 中文摘要
路边交通推理要求每个自由形式文本主张都必须有视觉证据支持。现有的基础多模态大型语言模型(MLLM)经常表现出说点不匹配,即文本答案与模型定位的边界框相矛盾。分别对答案和框进行评分的评估指标则未对此失败进行惩罚。我们将不匹配追溯到传统的答案-然后基础分解,该方法在枚举对象之前就承诺一个数值主张。为此,我们构建了RoadSceneVQA-G,这是一个包含34.7K问答对的基准测试,其中每个自由形式答案都关联到见证它的框集合,并提出了答案-基础一致性(AGC)评估套件。为此,我们引入了枚举然后回答(Enitemrate-then-Answer,EtA),它将生成顺序反转,使答案与证据的一致成为输出结构的属性,以及枚举一致策略优化(ECPO),这是一个强化学习阶段,利用多次推广的并集作为回忆教师,无需地面真实框。EtA将说点一致性从26.6%提升到93.7%,将F1的基准从52.2%提升到73.0%,ECPO进一步将F1提升到75.6%,无需每框监督。在gRefCOCO上,同一框架优于最强的对比方法,表明其跨越了交通场景。该项目可在 \url{this https URL} 访问。
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Red-TTT:自动化越狱大型语言模型的测试时训练
- Authors: Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu, Shuyao Xu, Ning Zhang, Ziyang Li, Yinzhi Cao, Bryan Hooi, Chaowei Xiao
- Subjects: Subjects:
Computation and Language (cs.CL); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2610.05282
- Pdf link: https://arxiv.org/pdf/2610.05282
- Abstract
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9\% to 72.4\% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at this https URL
- 中文摘要
大型语言模型仍然容易受到越狱威胁,而自动化红团队是大规模发现越狱的标准方法。目前的方法要么通过搜索、重写和树扩展在测试时抽取更多样本,要么通过强化学习离线训练更强的攻击者。两者都有一个局限:一旦对特定目标行为发动攻击,攻击者的权重会被冻结。它收集到的任何关于该行为的信号都会停留在其上下文窗口中,之后被丢弃。攻击者从不在攻击中调整其提案分布,因此成功几乎完全取决于采样预算,在可大规模负担的预算下,许多行为保持不被破坏。我们提出了Red-TTT,它在攻击过程中更新攻击者对每个行为的参数。每轮攻击者会抽取一组候选人,将其与受害者的回复进行评分,并采取策略梯度步骤后再抽取下一组,因此对当前受害者的了解被整合为权重,而非累积为上下文。我们还将训练目标调整为红队,成功率以单一最佳样本而非平均值衡量。Red-TTT只需对受害者进行抽样访问,并整合进现有攻击管道,不做其他更改。与Best-of-N基线相比,Red-TTT在120个样本预算下将攻击成功率从平均55.9%提升至72.4%,在每种配置中均优于基线,破解了许多之前方法无法破解的行为。代码可在此 https URL 获取
Erased, Rerouted, or Rescaled? Post-Training and the Causal Quotient of a Language Model's Belief State
是被抹除、重新定向还是重新调整比例?训练后与语言模型信念状态的因果商
- Authors: Weihan Li, Tianshi Zheng, Junhao Wu, Xinlei Chen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05292
- Pdf link: https://arxiv.org/pdf/2610.05292
- Abstract
What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still represented, or rescaled to occupy less variance while still represented and used. We make these fates identifiable in models whose pretraining recovers Bayesian belief states. A reward that reads only a coarse function of the hidden state defines an exact reward-null kernel. The kernel lets us separately measure whether the information remains recoverable, whether decisions causally depend on it, and how much activation variance it occupies. Theory says what is protected: KL-anchored reinforcement learning preserves the reference policy's log-odds among equally rewarded outputs, supervised and unanchored objectives carry no such constraint, and spectral compression implies neither erasure nor loss of use. In controlled worlds, post-training mostly reroutes or rescales reward-null information and leaves it decodable. Without an anchor decisions can stop using it although the representation survives, and with one they keep using it. Erasure appears only under prolonged weight decay, for distinctions that neither reward nor next-token prediction can see. Open language models show the same dissociation: in-context belief geometry stays decodable under late-layer spectral compression, and within-class behavior depends on the anchor. Post-training thus selects a causal quotient of the pretrained belief state: the reward defines decision-equivalence, the anchor and the state update protect part of what it ignores, and optimization decides whether the rest is erased, rerouted, or rescaled.
- 中文摘要
当预训练模型不再奖励使用该信息时,一个已编码的信息会发生什么?表示压缩的常用语言混淆了三种命运:信息可能被删除、在仍被表示时被重新定向、在仍被表示时占据更小的方差。我们使这些命运在预训练恢复贝叶斯信念状态的模型中可识别。仅读取隐藏状态粗函数的奖励定义了一个精确的奖励零核。核使我们能够分别衡量信息是否仍可恢复,决策是否因果依赖于它,以及其激活方差的程度。理论说明了受保护的部分:基于KL锚定的强化学习保留了参考策略在同等奖励输出中的对数赔率,监督和非锚定目标没有此类约束,谱压缩既不意味着删除也不意味着使用损失。在受控环境中,后训练大多是重新路由或重新缩放奖励空信息,使其可解码。没有锚点,决策可以停止使用,尽管该表示存续,且有锚点后仍会继续使用。擦除仅在长期权重衰减时出现,这些差异既是奖励也无法预测的,也无法预测下一标记。开放语言模型也表现出同样的解离:上下文中的信念几何在后期层谱压缩下仍可解码,类内行为依赖于锚点。因此,后训练选择预训练信念状态的因果商:奖励定义决策等价性,锚点和状态更新保护其忽略部分内容,优化决定其余部分是否被擦除、重定向或重新缩放。
RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation
RubricArmor:对抗进化提升基于LLM的评分标准生成
- Authors: Haocheng Yang, Yuchao Zhang, Licheng Pan, Jiajun Fan, Maolin Wang, Kangning Zhang, Shuai Shao, Shijian Wang, Yuan Lu, Chunyuan Zheng, Hao Wang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.05308
- Pdf link: https://arxiv.org/pdf/2610.05308
- Abstract
Rubric-based reinforcement learning (RL) provides interpretable rewards for aligning large language models (LLMs) by evaluating responses against query-specific evaluation criteria. To construct rubrics at scale, a straightforward approach to LLM-based rubric generation is to prompt an LLM to generate a rubric directly from the query. However, rubrics directly generated by LLMs are vulnerable to reward hacking, since omitted or underspecified criteria allow the policy to obtain high rubric rewards with low-quality responses. Existing LLM-based rubric generation methods improve the granularity and coverage of the generated criteria but do not proactively guard against reward hacking. To address this limitation, we propose RubricArmor, an adversarial framework that exposes and mitigates potential reward hacking at the rubric generation stage before it occurs in subsequent RL. Specifically, RubricArmor performs adversarial evolution, in which an attack step and a repair step alternate over multiple rounds. The attack step simulates the reward hacking of the policy by constructing adversarial responses that satisfy the current rubric but fail to properly complete the task. The repair step then revises the rubric to detect the response defects exposed by the attack step while preserving other valid criteria. Extensive experiments demonstrate that RubricArmor outperforms competitive rubric generation baselines and translates into more effective downstream rubric-based RL.
- 中文摘要
基于评分标准的强化学习(RL)通过根据查询特定评估标准评估响应,为大型语言模型(LLM)的对齐提供可解释的奖励。为了大规模构建评分标准,一种简单的方法就是让LLM直接从查询生成评分标准。然而,由LLM直接生成的评分标准容易受到奖励黑客攻击的影响,因为遗漏或不明确的标准使策略能够获得高评分奖励但回答质量较低。现有基于LLM的评分标准生成方法提升了生成标准的粒度和覆盖范围,但并未主动防范奖励黑客行为。为解决这一限制,我们提出了RubricArmor,这是一种对抗性框架,可以在评分标准生成阶段暴露并缓解潜在的奖励黑客行为,防止后续强化学习发生。具体来说,RubricArmor 执行对抗演化,即攻击步骤和修复步骤交替进行多轮。攻击步骤通过构建满足当前评分标准但未能正确完成任务的对抗性反应,模拟策略的奖励黑客行为。修复步骤随后修订评分标准,以检测攻击步骤暴露的响应缺陷,同时保持其他有效标准。大量实验表明,RubricArmor 优于竞争性的评分标准生成基线,并转化为更有效的基于评分标准的后续强化学习。
On Semi-Markov Suboptimality in Hierarchical Reinforcement Learning
关于分层强化学习中的半马尔可夫次优性
- Authors: Bingyun Liu, Yuheng Jing
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05338
- Pdf link: https://arxiv.org/pdf/2610.05338
- Abstract
Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execution trees distinguish reward objectives from policy choices and decision interruption. A Unified Value Function for HRL and a four-stage Generalized Hierarchical Bellman Equation then support a common analysis of both losses. Under bounded rewards and uniform termination, we establish hierarchical policy and execution improvement results. With the remaining node policies fixed, task-subtree compatibility and node-policy optimality under the original execution mode establish when Markov execution is optimal. The resulting decomposition leads to independent execution choices for behavior, targets, and deployment. We instantiate this principle through execution improvement and one-stage or two-stage policy improvement at arbitrary hierarchy depth. Option-based and goal-conditioned experiments demonstrate complementary gains from changing execution and changing the learning target. Controlled stochastic environments show how these gains depend on stochastic transition strength and spatial structure. This framework makes execution design an explicit component of hierarchical policy optimization.
- 中文摘要
层级强化学习利用时间扩展子任务进行探索,但承诺执行它们可能会限制部署和策略学习。我们识别并区分最终的执行和策略次优性。任务和执行树区分奖励目标、策略选择和决策中断。HRL的统一价值函数和四阶段广义层级贝尔曼方程支持对这两种损失的共同分析。在有界奖励和均匀终止条件下,我们建立了层级策略和执行改进结果。在剩余节点策略确定后,任务子树兼容性和原始执行模式下节点策略的最优性确定了何时马尔可夫执行最优。由此分解产生行为、目标和部署的独立执行选择。我们通过执行改进以及任意层级深度的一阶段或两阶段策略改进来实现这一原则。基于选项和目标条件的实验展示了改变执行和改变学习目标带来互补的收益。受控随机环境表明这些收益依赖于随机过渡的强度和空间结构。该框架使执行设计成为层级策略优化的显式组成部分。
Optimal Control with Learned Critics under Unmodeled State Dependencies
在未建模状态依赖下,学习批评者的最优控制
- Authors: Philipp Schoch, Markus Ryll
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05359
- Pdf link: https://arxiv.org/pdf/2610.05359
- Abstract
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
- 中文摘要
模型预测控制(MPC)提供了结构化且约束感知的决策机制,但其对优化友好型分析动力学模型的依赖限制了其在接触及其他难以建模状态依赖的任务中的应用。无模型强化学习避免显式建模假设,但通常需要大量交互数据。我们提出了一个基于学习的MPC框架,结合了局部模型规划的数据效率和结构,并结合了学习到的组件,以补偿不完全的动态和有限视野的近视。该方法用残差动力学网络补充名义分析模型,该网络学习数据中缺失的状态依赖效应,并将生成的规划器与学习到的动作值批评器结合,将长视野MDP结构注入局部iLQR优化中。为了在强化学习尺度下实现这一目标,我们开发了一款GPU加速的批量iLQR求解器,能够在最优控制环路内评估已学到的动力学和批判网络,并并行解决数千个轨迹优化问题。完整系统集成到机器人模拟器中,实现在不完全动力学下可扩展的基于模型的强化学习。对偏置和不完全建模控制任务的实验表明,该方法在保持高效受限轨迹优化所需的基于模型的结构的同时,提升了闭环控制性能。
AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness
AIProver:通过证书驱动的进化工具实现数学研究的能动自形式化
- Authors: Prithwish Jana, Viet Bach Hoang, Logan Luna, Viresh Pati, Akash Singirikonda, Cy Xie, Lisa Carbone, Wuyang Chen, Walter Moreira, Joe Stubbs, Sriram Vishwanath, Vijay Ganesh
- Subjects: Subjects:
Logic in Computer Science (cs.LO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05367
- Pdf link: https://arxiv.org/pdf/2610.05367
- Abstract
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
- 中文摘要
证明自形式化将自然语言(NL)定理和证明转换为如精益(Lean)的形式语言(FL),从而实现机械性验证。尽管进展迅速,研究级证明往往依赖于领先证明辅助库(如Lean的Mathlib)缺失的概念,成功编译并不保证翻译能保持定理的意义或证明的推理。此外,NL-FL对齐训练数据稀缺,领先代理常依赖昂贵的前沿模型和手动设计的束缚。为解决上述问题,我们介绍AIProver,一个用于自主证明自动形式化与证明综合(AFPS)的代理框架,它联合后期训练119B开权重语言模型,并与HarnessEvolve共同演进其代理工具调用框架。验证者评估类型正确性、证明完整性和语义正确性,返回奖励和诊断证书,推动模型微调和通过符号反馈交替强化学习,以及HarnessEvolve——一种基于证书驱动的整个约束控制流程进化搜索,重新调整框架以适应更新模型。在研究级培训和评估方面,我们引入了LoCoBench、来自Mathlib、CSLib、Mizar数学库的58.9k实例,以及一本有界算术教材,其中771实例验证,其定理证明对无公开的精益形式化。面对涵盖AFPS代理、前沿大型语言模型和编码代理的39个框架,AIProver将其Leanstral-1.5基数的语义正确性pass@4从15.7%提升至36.7%,并优于所有其他开放权重系统和亚里士多德。作为Claude代码和Codex技能,其语义正确率分别从41.9%和34.1%提升至79.8%和62.4%。此外,它比Numina精益代理便宜24%,推动了研究级AFPS的准确性成本前沿。
Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks
Sibyl:一个高效的小-大模型协作框架,用于长期任务
- Authors: Zhewei Fang, Yuxin Zhang, Zhenwei Shao, Mengze Li, Zheng Lin, Long Chen, Zhou Yu, Zhe Chen, Zhiwen Chen, Zhaode Wang, chengfei lv
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05383
- Pdf link: https://arxiv.org/pdf/2610.05383
- Abstract
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.
- 中文摘要
小型语言模型(SLMs)通过低延迟、资源高效的推断为设备内代理提供了有前景的基础,但有限的推理和规划能力限制了其在需要多步与环境交互的长期任务中的表现。SLM与更大型云托管模型之间的步骤级协作可以弥合这一差距,但识别需要云协助的状态仍然具有挑战性:每个云调用的贡献与后续动作纠缠在一起,只能从最终任务结果中评估。加剧这一挑战的是,SLM必须平衡两个相互竞争的目标:最大化任务成功率和最小化云调用次数。为此,我们提出了Sibyl算法,该算法训练SLM代理在步骤层面有选择性地咨询云模型,并将其指导内化为后续决策,实现强大的任务性能,同时对云的依赖最小化。Sibyl 遵循三阶段训练流程,(1) 通过无协商的自我演化强化学习(RL)构建稳健的基础策略;(2) 通过决定性分歧状态挖掘实现冷启动协商行为;(3) 通过协商感知强化学习共同优化协商决策和指导内化。在 ALFWorld 和 WebShop 上的实验表明,Sibyl 仅使用 0.6B 参数模型,其成功率分别高于 95.2% 和 80.4%,且每轨迹平均仅有 0.8 和 3.9 次云调用。
Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking
未提及的检查清单发现改变了强化学习改善胸部X光报告检查的方式
- Authors: Ali Vosoughi, Akhil Kasturi, Chenliang Xu, Axel Wismueller
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05425
- Pdf link: https://arxiv.org/pdf/2610.05425
- Abstract
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.
- 中文摘要
放射科报告的自动检查可能依赖AI生成的检查表,而这些检查结果未被提及。我们使用强化学习训练视觉语言模型,在未看到被测句子的情况下,从胸部X光片中填写12项发现清单;另一个检查模型从检查表中判断句子。对于未被保留的患者,基于规则的检查和一个独立医疗检查器(均未用于培训)测量了辨别力提升(Youden指数)分别为12.6%和11.8%;只有基于规则的检查符合预设的误报标准。切换到训练格式(固定发现顺序并未提及的发现为缺失)后,提高了训练检查器的测量增益,降低了独立检查者的,预设比较得出6.2%(95%间隔2.0%到10.5%),事后对未检测患者获得7.7%。在8个检查模型中,对未提及发现的标签一致否定陈述接受率从1.0%到97.0%不等。标签是报告生成的,而非放射科医生裁定的。
Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning
层级时间感知自助,用于非策略子目标价值学习
- Authors: Bingyun Liu, Yuheng Jing
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05446
- Pdf link: https://arxiv.org/pdf/2610.05446
- Abstract
Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.
- 中文摘要
非策略层级强化学习必须在低层策略变化时估计高层决策的价值。HIRO通过子目标重新标记调整重放数据,但标签变更后,值更新目标目标是重新标记的子目标,而非高层策略原本需要更新的子目标。我们提出层级时间感知引导(HTB),在当前低级策略下评估指定子目标,同时保留累积的任务奖励。剩余执行时间区分了子目标延续与新的高层决策。结合原始动作条件,它支持基于平稳环境转换律的非策略贝尔曼更新。HTB将这些一步更新与多步后缀返回和截断重新标记结合,减少对中间值估计的依赖。共享值组件支持跨动作学习,而非负残差则限制相对于该组件的向上修正。在固定混合权重0.95下,HTB在1000万环境步长下,5对种子匹配本地HIRO的AntFall成功率为9.6%。消融法识别递归延续和混合监督的贡献;固定策略检验显示,对于那些返回被排除在拟合之外的动作,预测更为准确。
Groupwise Distortion Guarantees for Preference-Based Alignment
基于偏好的群组失真保证
- Authors: Jacob Brodkey, Roberto Tamez, Aaron Roth
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05450
- Pdf link: https://arxiv.org/pdf/2610.05450
- Abstract
Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal utility), which comparisons alone need not identify. Gölz, Haghtalab, and Yang (GHY) measure the gap by distortion: the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. Under individual Bradley--Terry comparisons, GLHF asymptotically matches GHY's optimal population distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. A sharper guarantee for groups with similar preferences approaches distortion of one when members share a feasible favorite response. In experiments using human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every evaluated group and substantially reduces worst-group distortion relative to NLHF and other group-agnostic baselines.
- 中文摘要
基于偏好的对齐方法,如人类反馈强化学习(RLHF)和纳什人类反馈学习(NLHF),汇聚成对偏好以学习LLM策略,但其自然目标是最大化社会福利(平均基数效用),仅靠比较不必识别。Gölz、Haghtalab和Yang(GHY)通过扭曲来衡量差距:最佳固定彩票(分布与响应比例)的福利比值与学习式彩票的最坏情况比率。他们表明,当每个用户都收到相同的彩票时,NLHF是最优的。然而,基于账户的LLM拥有用户信息,可以为不同人群提供不同的彩票。我们给出了一个高效的算法GLHF,通过每个用户的一次比较学习单一群体条件策略。在单独的Bradley-Terry比较下,GLHF渐近地匹配GHY在预指定且可能重叠的集合中每个组同时绑定的最优群体畸变,样本复杂度在组数上呈对数增长,组质量最小时呈反比。对于偏好相似的组,当成员共享可行的偏好反应时,更明确的保证是采用一个失真。在使用人类咖啡评分和合成LLM生成评分的实验中,GLHF降低了每个评估组的失真,并显著降低了相较于NLHF及其他无组基线的最差组畸变。
AI Safety via Debate is Compromised by Cognitive Biases
通过辩论进行的人工智能安全受到认知偏见的影响
- Authors: Gefei Liu, Sonya Rashkovan, Sophia Lloyd George, Isaac Sheidlower, Serena Booth
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05461
- Pdf link: https://arxiv.org/pdf/2610.05461
- Abstract
Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.
- 中文摘要
来自人类反馈的强化学习(RLHF)在使大型语言模型对人类指令做出响应中发挥了核心作用。然而,人类评估者往往更倾向于奉承或说服性的回答而非真实的回答,这为模型吸引评估者而牺牲准确性产生了激励。通过辩论实现人工智能安全被提出以改善语言模型监督的一种方式:在这种范式中,两个代理争论对立立场并相互质疑对方的主张,可能向裁判者揭露虚假信息。通过辩论实现人工智能安全的核心前提是,在对抗性审视下,真实的论点比虚假的论点更容易辩护。本研究中,我们探讨了当辩手使用利用人类判断偏见的修辞策略时,这种优势是否依然存在。受竞争性辩论启发,我们构建了68个由大型语言模型生成的关于已知罪犯侦探悬疑的对话,涵盖四种干预方式:锚定、谬误监督、专业术语和冗长。我们将每种干预应用于支持真正罪犯的一方或支持无辜嫌疑人的一方,从而区分对裁决的影响与正确性。在一项包含369名参与者的研究中,我们发现,这些干预在不同偏见类型间显著偏向控方的判断。这些发现暴露了基于辩论的监督中的一个漏洞:人类裁决对操控性修辞策略非常敏感。
Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents
语言模型代理的层级强化学习与稳定时间抽象
- Authors: Shayan Mohajer Hamidi, Yize Cheng, Yuanda Xu, Zhengze Zhou, Alborz Geramifard
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05473
- Pdf link: https://arxiv.org/pdf/2610.05473
- Abstract
Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by $8.1$ and $7.9$ points on ALFWorld and WebShop with Qwen3-0.6B, and by $23.5$ and $15.8$ points with Llama-3.2-1B-Instruct.
- 中文摘要
分层强化学习通过围绕持久子目标组织原始动作并在多个时间尺度上分配功劳,从而提升了长视野控制。最新的分层语言智能体通过显式区分子目标规划与动作执行,将这些优势带入交互任务。然而,我们观察到,显式层级本身并不能决定最终时间抽象的稳定性:学习的边界策略几乎每回合都会替换子目标,使其实际上是暂时的,或者在子目标不再适用后仍保留该子目标。我们称之为时间抽象不稳定性。我们提出了通过受限优化实现稳定时间抽象(STAC),这是一种受限边界策略优化方法,将过早的重新规划和陈旧持久化视为约束成本。STAC仅将所得的拉格朗日成本应用于抽样边界决策,底层算法的奖励、批评目标、子目标优势和原始动作优势保持不变。在两个骨干和两个基准测试中,STAC在ALFWorld和WebShop上,Qwen3-0.6B的成功率提升了8.1美元和7.9美元,Llama-3.2-1B-Instruct提升了23.5美元和15.8美元。
An LLM-in-the-loop RL Framework for Bioinformatics Feature Selection
一个用于生物信息学特征选择的LLM在环中强化学习框架
- Authors: Xinyuan Wang, Deepti Agrawal, Yanjie Fu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05600
- Pdf link: https://arxiv.org/pdf/2610.05600
- Abstract
High-dimensional bioinformatics data, characterized by a large number of features relative to the number of samples, pose major challenges such as the ``curse of dimensionality,'' leading to overfitting, high computational cost, and poor generalization. Traditional feature selection methods often suffer from limited scalability and adaptability in such domains. We propose an LLM-in-the-loop reinforcement learning (RL) framework for bioinformatics feature selection, where the RL agent formulates feature selection as a sequential decision-making task, while the large language model (LLM) enhances the process in two ways: (1) guiding exploration through domain-informed advice, and (2) providing hybrid rewards that integrate data-driven performance with knowledge-driven evaluation. The LLM also produces explanations to improve interpretability for human experts without altering the RL policy update. Experiments on diverse bioinformatics datasets show that the LLM-in-the-loop framework outperforms baselines, achieves stable performance across downstream models, and converges faster than pure RL.
- 中文摘要
高维生物信息学数据,特征数量相对于样本数量而言数量庞大,带来了诸如“维度诅咒”等重大挑战,导致过拟合、高计算成本和泛化能力差。传统特征选择方法在此类领域常常存在扩展性和适应性有限的问题。我们提出了一种生物信息学特征选择的LLM环路强化学习(RL)框架,其中强化学习代理将特征选择作为顺序决策任务,而大型语言模型(LLM)则通过两方面增强这一过程:(1)通过领域导向的建议引导探索,(2)提供结合数据驱动性能与知识驱动评估的混合奖励。LLM还生成解释,以提升人类专家的可理解性,同时不改变RL策略更新。在多种生物信息学数据集上的实验表明,LLM环路框架优于基线,在下游模型间实现稳定性能,且收敛速度快于纯强化学习。
SCOUT: Supply-Aware Cold-Start Proactive Query Suggestion for Travel Search
SCOUT:供应意识冷启动的主动查询建议,适合旅行搜索
- Authors: Hao Li, Shashank Reddy, Kedar Bellare, Ashish Jain, Stephanie Moyerman
- Subjects: Subjects:
Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2610.05619
- Pdf link: https://arxiv.org/pdf/2610.05619
- Abstract
Generative query suggestion, powered by Large Language Models (LLMs), has become increasingly popular in search and conversational systems to reduce user friction and guide intent formulation. Existing approaches align suggestions with user preferences (e.g., clicks or conversions). This works for open-ended applications like chatbots and personal assistants, where the result space is unconstrained or historical user free-text queries are abundant. However, applying these methods to travel search presents two limitations. First, travel search is fundamentally constrained by physical inventory; a query (e.g., "romantic beachfront villa") may yield abundant results in Bali but few in Tokyo, so aligning with user preferences is not by itself grounded in what can be offered. Second, travel platforms traditionally rely on faceted search interfaces with no free-text queries. This creates a cold-start problem: without historical query logs there is no demand-side data for alignment, and without a seed query at request time, suggestions must be generated proactively from structured context alone. To address these challenges, we propose SCOUT, a bootstrapping framework for supply-aware proactive query suggestion. SCOUT overcomes the data gap by substituting missing demand-side user feedback with supply-side system feedback. It treats the search engine as a reinforcement learning environment, deriving a dense reward from the production reranker's query-listing match scores, and optimizes the policy with Group Relative Policy Optimization (GRPO). SCOUT improves inventory match rate (IMR@18) by 12.3% while preserving diversity, matching a compute-intensive best-of-8 policy at zero marginal inference cost and making supply-aware suggestion deployable on a real-time travel search path.
- 中文摘要
由大型语言模型(LLM)驱动的生成式查询建议在搜索和对话系统中日益流行,以减少用户摩擦并指导意图形成。现有方法将建议与用户偏好(如点击或转化)对齐。这适用于像聊天机器人和个人助理这样开放式应用,因为结果空间不受限制或历史用户自由文本查询丰富。然而,将这些方法应用于旅游搜索存在两个局限。首先,旅行搜索根本受物理库存限制;查询(例如“浪漫海滨别墅”)在巴厘岛可能有大量结果,但在东京很少,因此与用户偏好对齐本身并不依赖于可提供的内容。其次,旅游平台传统上依赖分面搜索界面,没有自由文本查询。这造成了冷启动问题:没有历史查询日志,就没有需求端数据用于对齐,且请求时没有种子查询,建议只能主动从结构化上下文生成。为应对这些挑战,我们提出了SCOUT,一个用于供给感知主动查询建议的自助框架。SCOUT通过用供给侧系统反馈替代缺失的需求端用户反馈来克服数据缺口。它将搜索引擎视为强化学习环境,从生产重排序者的查询列表匹配分数中获得密集奖励,并通过组相对策略优化(GRPO)优化策略。SCOUT在保持多样性的同时提升库存匹配率(IMR@18)12.3%,匹配计算密集型八局三胜策略且边际推断成本为零,并使供给感知建议可部署于实时旅行搜索路径上。
Bellman-Centric Learning: Near-Optimal Regret for Linear Bandits with Memory
贝尔曼中心学习:记忆力强的线性强盗近乎最佳后悔
- Authors: Jingyuan Liu, Huiwen Jia
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05659
- Pdf link: https://arxiv.org/pdf/2610.05659
- Abstract
We study linear bandits with memory, where past actions induce endogenous nonstationarity through an arbitrary known, bounded matrix-valued memory map. To trade off exploration and exploitation while accounting for the memory dynamics, we develop RSM-LinUCB, a Bellman-centric algorithm that learns as in linear bandits and plans as in reinforcement learning. This design admits a novel regret decomposition which separates the memory-induced error from the cumulative reward estimation error along the learner's trajectory. We prove a high-probability regret bound of $\widetilde O\big(dRS(M+1)+\sigma d\sqrt T\big)$, where $T$ is the learning horizon, $d$ is the parameter dimension, $M$ is the memory length, $R$ and $S$ bound the memory-map operator norm and reward-parameter norm, respectively, and $\sigma$ is the sub-Gaussian noise scale. Our results reveal that the multiplicative memory-horizon coupling in prior bounds is not intrinsic: memory only contributes an additive cost, up to logarithmic factors. We also prove a matching minimax lower bound, establishing near-optimality. We further extend the algorithm to generalized linear rewards, preserving this separation with near-optimal memory and leading statistical dependence. Our algorithms outperform the baselines in numerical experiments on synthetic instances and semi-synthetic KV- and semantic-cache tasks.
- 中文摘要
我们研究带有记忆的线性盗贼,其中过去的行为通过已知的有界矩阵值记忆映射诱导内生非平稳性。为了在考虑记忆动态的同时兼顾探索与利用,我们开发了RSM-LinUCB,一种以Bellman为中心的算法,其学习方式如线性强化学,计划则如同强化学习。该设计采用了一种新颖的遗憾分解,将记忆引起的误差与累计奖励估计误差沿学习者轨迹分离。我们证明了一个高概率的后悔界限:$\widetilde O\big(dRS(M+1)+\sigma d\sqrt T\big)$,其中$T$为学习视界,$d$为参数维数,$M$为记忆长度,$R$和$S$分别界定了记忆映射算子范数和奖励参数范数,$\sigma$为亚高斯噪声尺度。我们的结果表明,先验界限中的乘法记忆视界耦合并非内在的:记忆仅贡献对数因子的加法成本。我们还证明了一个匹配的极小极大下界,建立了近似最优性。我们进一步将算法扩展到广义线性奖励,保持这一分离,记忆值近乎最优,并保持领先统计依赖。我们的算法在合成实例和半合成KV缓存和语义缓存任务的数值实验中表现优于基线。
End-to-End Safe Social Navigation via Multi-Task Reinforcement Learning and Probabilistic Perception
通过多任务强化学习和概率感知实现端到端安全社交导航
- Authors: Tommaso Van Der Meer Andrea Garulli, Antonio Giannitrapani, Renato Quartullo, Alberto Vaglio, Alexandre Alahi
- Subjects: Subjects:
Robotics (cs.RO); Human-Computer Interaction (cs.HC)
- Arxiv link: https://arxiv.org/abs/2610.05733
- Pdf link: https://arxiv.org/pdf/2610.05733
- Abstract
Autonomous social navigation requires balancing efficiency, physical safety, and social compliance. Reinforcement Learning (RL) methods provide a viable and effective solution but often rely on unrealistic assumptions, such as the knowledge of humans' position and velocity. In this paper, we introduce JESSI (JAX-based E2E Safe Social Interpretable navigation), a lightweight end-to-end RL framework that maps raw LiDAR scans directly to kinematically feasible control commands. JESSI enhances safety via Dirichlet-parameterized continuous action spaces and deterministic bounding, while an integrated attention-based perception module extracts probabilistic human states for interpretable, socially aware decision-making. Through extensive simulations and real-world deployment on a differential-drive robot, we demonstrate that jointly optimizing the RL policy with a supervised perception signal in a multi-task paradigm enhances social behavior. Ultimately, JESSI is able to balance high navigation success rates and superior social behaviors compared to state-of-the-art baselines.
- 中文摘要
自主社会导航需要在效率、物理安全和社会合规之间取得平衡。强化学习(RL)方法提供了可行且有效的解决方案,但常依赖于不切实际的假设,如对人类位置和速度的认知。本文介绍了JESSI(基于JAX的E2E安全社会可解释导航),这是一个轻量级端到端的RL框架,将原始LiDAR扫描直接映射到运动学上可行的控制命令。JESSI通过狄利克雷参数化的连续动作空间和确定性边界增强安全性,而集成的基于注意力的感知模块则提取概率人类状态,实现可解释且具社会意识的决策。通过在差动驱动机器人上的广泛模拟和实际部署,我们证明在多任务范式中联合优化强化学习策略与监督感知信号,能提升社会行为。最终,JESSI能够在较高的导航成功率和优越的社交行为之间取得平衡,相较于最先进的基线。
Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning
通过多目标强化学习,用四足机器人运输未加固的堆叠有效载荷
- Authors: Nobuo Namura, Masayuki Hiromoto, Kento Uemura, Hironobu Sasaki, Kanata Suzuki
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05819
- Pdf link: https://arxiv.org/pdf/2610.05819
- Abstract
Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.
- 中文摘要
用腿部机器人在崎岖地形上运输未固定有效载荷需要平衡移动性能和有效载荷稳定性,因为剧烈运动即使机器人保持稳定也可能破坏有效载荷稳定性。我们研究了在无边躯干安装板上四足运输未固定叠放箱子,无需专用载荷传感器或主动承载机制。为解决这一权衡,我们提出了有效载荷自适应多目标强化学习用于运输(PAMORT)。PAMORT训练基于偏好向量的多目标基础策略,该向量加权运动组和有效载荷稳定性奖励组,然后训练重量调整器在冻结策略上将该偏好从本体感觉在线调整。在模拟中,尽管仅用两个箱子训练,PAMORT在不同有效载荷配置(包括未见的三盒堆栈)下,整体运输成功率与对应单目标基线相当甚至更好。在Unitree Go2上的实际实验显示,在训练难度或更高坡道上实现零发射转移,八项任务中PAMORT的平均成功率为0.850,基线为0.675。这些结果展示了稳健的无安全有效载荷运输,并实现了移动的在线适应——有效载荷与本体感觉信息的权衡。
Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped
针对欠驱动双足行走的无碰撞运动的层级强化学习
- Authors: Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S.K., Samiran Datta, Abhay Dwivedi, Amit Shukla
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.05855
- Pdf link: https://arxiv.org/pdf/2610.05855
- Abstract
A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, \omega_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A, SAC+RRT, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.
- 中文摘要
双足机器人不能偏离路径以避开障碍物而不扰乱平衡,这种耦合在像这里讨论的双足驱动平台这样欠驱动平台上最为严重,该平台每条腿有四个驱动关节,且无髋部或脚踝滚动。本文提出了一种层级强化学习(HRL)框架,其中高级别(HL)策略观察机器人姿势、36个射线近距离测量、移动障碍物状态和远视线的局部目标,每十步输出一次身体速度指令$(v_x, v_y, \omega_{yaw})$,同时速度条件低级(LL)策略通过PD控制的联合目标追踪每个指令。这两种策略均与软演员-批判者(SAC)联合训练。由于收敛步态与任务无关,它被冻结并由传统规划者在同一指令界面驱动,得到三个受控基线:SAC+A、SAC+RRT 和 SAC+APF。在随机PyBullet环境中,每种方法进行了100次评估试验,拟议方法在静态试验中达到了98.0%,动态试验中达到了88.0%,而规划者混合试验最多达到78.0%和68.0%,路径长度在A*参考值的4%以内,消融结果确认每个观察通道和奖励项对该表现有显著贡献。
Generative-AI for XR Content Transmission in the Metaverse: Potential Approaches, Challenges, and a Generation-Driven Transmission Framework
元宇宙中XR内容传输的生成式人工智能:潜在方法、挑战与代际驱动传输框架
- Authors: Zhe Zhang, Yili Jiang, Xin Wei, Mingkai Chen, Haiwei Dong, Shui Yu
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2610.05888
- Pdf link: https://arxiv.org/pdf/2610.05888
- Abstract
How to efficiently transmit large volumes of Extended Reality (XR) content through current networks has been a major bottleneck in realizing the Metaverse. The recently emerging Generative Artificial Intelligence (GAI) has already revolutionized various technological fields and provides promising solutions to this challenge. In this article, we first demonstrate current networks' bottlenecks for supporting XR content transmission in the Metaverse. Then, we explore the potential approaches and challenges of utilizing GAI to overcome these bottlenecks. To address these challenges, we propose a GAI-based XR content transmission framework which leverages a cloud-edge collaboration architecture. The cloud servers are responsible for storing and rendering the original XR content, while edge servers utilize GAI models to generate essential parts of XR content (e.g., subsequent frames, selected objects, etc.) when network resources are insufficient to transmit them. A Deep Reinforcement Learning (DRL)-based decision module is proposed to solve the decision-making problems. Our case study demonstrates that the proposed GAI-based transmission framework achieves a 2.8-fold increase in normal frame ratio (percentage of frames that meet the quality and latency requirements for XR content transmission) over baseline approaches, underscoring the potential of GAI models to facilitate XR content transmission in the Metaverse.
- 中文摘要
如何通过现有网络高效传输大量扩展现实(XR)内容一直是实现元宇宙的一个主要瓶颈。新兴的生成式人工智能(GAI)已经革新了多个技术领域,并为这一挑战提供了有前景的解决方案。本文首先展示了当前网络在支持元宇宙中XR内容传输时的瓶颈。随后,我们探讨利用GAI克服这些瓶颈的潜在方法和挑战。为应对这些挑战,我们提出了基于GAI的XR内容传输框架,利用云边缘协作架构。云服务器负责存储和渲染原始XR内容,而边缘服务器利用GAI模型在网络资源不足时生成XR内容的关键部分(如后续帧、选定对象等)。提出基于深度强化学习(DRL)的决策模块来解决决策问题。我们的案例研究表明,基于GAI的传输框架在正常帧率(满足XR内容传输质量和延迟要求的帧百分比)上比基线方法提高了2.8倍,凸显了GAI模型在元宇宙中促进XR内容传输的潜力。
Safe Image Generation via Reinforcement Learning
通过强化学习实现安全图像生成
- Authors: Eungyeol Han, Jong-Seok Lee
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.05908
- Pdf link: https://arxiv.org/pdf/2610.05908
- Abstract
Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.
- 中文摘要
最新的文本转图像(T2I)模型实现了显著的视觉图像生成性能,但仍可能生成NSFW(不适合工作场所)内容,包括暴力或露骨图像。现有的安全检查机制主要局限于生成前过滤(如提示级文本分类器)或图像完全合成后施加的事后审核。然而,对抗攻击方法作用范围更广。这种不平衡凸显了在生成过程中介入安全机制的必要性。我们提出了一种代内安全框架,监控去噪轨迹并检测中间表示中新出现的NSFW信号。我们的方法不仅检测NSFW生成,还应用强化学习从NSFW提示生成安全图像。通过将代内检测与可控转向结合,我们的方法即使在生成开始后出现NSFW信号时,也能减轻不安全轨迹。实验结果显示,我们的方法在标准和对抗性评估集上始终优于现有的安全图像生成方法,同时保持感知质量和即时准确性。代码将在接受后发布。
MEND: RL For Flow Models via Proximal Velocity Matching
MEND:通过近距离速度匹配实现流模型的强化学习
- Authors: Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
- Subjects: Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.05954
- Pdf link: https://arxiv.org/pdf/2610.05954
- Abstract
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
- 中文摘要
对流模型的奖励训练后,要么在 KL 惩罚或冻结参考下重新加权模型自身样本,通常需要数千次更新;要么反向传播奖励,移动每个样本时不检查移动是否值其大小。我们引入了 MEND,一种基于近端速度匹配的强化学习方法。MEND 在每个提示组内限制奖励,因此已经得分较好的样本不获得移动。低于上限,它建议沿奖励梯度移动,只有当其上限奖励获得超过二次位移价格时才接受。随后模型回归到所得的速度目标上,没有 KL 项、冻结参考模型或优势权重。在 100 次更新中,MEND 在与基础模型图像相同距离的评估者中,超过了 Flow-GRPO(约 4000 次更新)。在等预算协议下,它在四项训练奖励的每次评估更新中均超过ReFL和DiffusionNFT,分别达到PickScore为24.03、23.92和23.43。300次更新、三次奖励的运行也超过了DiffusionNFT五奖励模型,涵盖其训练的三个奖励。MEND通用且易于采用:适用于任何具有可微分奖励的流程骨干。
Strategic Multi-Agent Learning for Interpretable Action Valuation of All Players in Football
战略多智能体学习,用于对所有球员进行可解释的行动评估
- Authors: Kenjiro Ide, Taiga Someya, Kohei Kawaguchi, Keisuke Fujii
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05961
- Pdf link: https://arxiv.org/pdf/2610.05961
- Abstract
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.
- 中文摘要
评估足球中球员的行为需要考虑22名球员之间的战略互动,包括无球移动和防守位置。现有基于强化学习的方法通常在球队层面汇总决策或独立估计球员价值,导致球员间的战略相互依赖性表现不足。本研究提出了一个受马尔可夫完美均衡(MPE)启发的行动估值框架,适用于所有球员。每个持球被建模为有限视野动态博弈,每个球员作为一个自主的代理,其策略依赖于当前比赛状态。MPE被用作激励性解概念,而非精确均衡。为提升可解释性,我们使用可扩展决策状态(EDMS),并将Q值分解为后继特征基和线性奖励权重向量。价值基底通过线性TD初始化和非线性细化估计。利用95场J1联赛比赛的跟踪和事件数据,我们将拟议表述与独立强化学习基线进行比较。由于两种表述定义了不同目标空间的TD误差,TD MSE仅用于公式内一致性。在EDMS固定的情况下,独立基线在99.21%的评估无球状态中赋予前进运动最高值,而在拟定表述下最频繁的方向则占17.63%。球队层面的平均Q值显示基线与赛季预期目标呈负相关,而对拟议表述的表述则呈弱正相关。定性分析展示了对无球移动和防守站位的上下文依赖估值。总体而言,拟议表述产生了更多上下文敏感的动作排名,尽管比较未能单独隔离MPE启发的部分。
Reachability-Aware Diffusion Policy Optimization
可达性感知扩散策略优化
- Authors: Hikmet Simsir, Kutay Demiray, Ozgur S. Oguz
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.05969
- Pdf link: https://arxiv.org/pdf/2610.05969
- Abstract
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
- 中文摘要
扩散策略为连续控制强化学习提供了表达式动作分布。然而,安全意识的在线扩散策略优化仍未被充分探索,尤其是使用预测性可达性信息而无显式动力学模型的方法。我们提出了可达性感知扩散策略优化(RADPO),这是一种无模型方法,结合了预测性首次命中安全估计与累积成本预算反馈。RADPO学习一个折现的首次命中可达性值,捕捉成本事件的折现风险,赋予较早发生事件更大权重,并利用该信号塑造奖励。一个独立的类似对偶乘数根据实现的情节成本相对于规定预算调整塑形强度。扩散演员通过对奖励批评者评分的候选动作进行加权去噪回归来提升。我们的方法既不需要学习动力学模型,也不需要通过批评者进行动作梯度,也不需要通过反向扩散采样器进行微分。我们建立了可达性值的理论性质,并展示了累积可达性惩罚为未来折现累计成本提供了保守的替代。在十个连续控制安全任务中,RADPO实现了具有竞争力的回报-成本权衡,相较于对比基线,在多个任务上显著减少了约束违规。我们的理论和实证分析支持将可达性与累积预算反馈结合起来,是安全意识扩散政策的可行方法。
Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers
强化学习改进推理教师的转移分层政策提炼
- Authors: Xiaoyu Chen, Bo Shao, Tiangang Zhu, Bintao Wu, Linjun Shou, Fengge Wu, Feng Sun, Wenbiao Ding
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.05974
- Pdf link: https://arxiv.org/pdf/2610.05974
- Abstract
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
- 中文摘要
强化学习可以显著提升推理型教师,但当教师指导较小的政策型学生时,哪些改进得以保留尚不清楚。我们通过比较GRPO前后的教师谱系、多重学生量表、直接GRPO以及若干策略型提纯目标,研究数学推理中的这一问题。核心发现是转移是结构化的而非标量性的:仅靠教师实力不足以使密集提纯具有竞争力,而强化学习改进的教师则能创造有用但依赖指标的学生进步。这促使转移分层策略提炼(TS-OPD)通过学生和教师联合抽样成功筛选培训问题,将习得问题引导至门控前向 KL,将巩固问题导向 gated reverse KL,并添加熵制动以保护抽样覆盖。在主要比较中,TS-OPD是与GRPO改进教师在宏观平均正确度方面最强的学生目标,而pass@K则较为复杂。消融显示,收益来自于路由和令牌门槛,而非跳过问题。这些结果支持了对OPD的迁移意识观点:当监督方向和令牌预算与学生观察到的能力相匹配时,教师越强,而不仅仅是因为教师的终点更强。
How (and How Not) to Use Data Augmentation in VLA Post-Training
如何在VLA训练后使用数据增强(以及如何不使用)
- Authors: Bram Grooten, Joaquin Vanschoren
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.05994
- Pdf link: https://arxiv.org/pdf/2610.05994
- Abstract
Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For $\pi_{0.5}$ and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by $7.8$ and $10.0$ points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.
- 中文摘要
视觉-语言-动作(VLA)模型目前在多种现实机器人任务中表现出强劲表现。然而,它们通常仍缺乏应对大规模视觉外分布转移的泛化能力。用强化学习(RL)对VLA进行后训练已被证明有助于稳健性,但仍有很大改进空间。在本研究中,我们系统地研究了图像增强对VLA后训练的影响。我们发现,在强化学习更新期间仅增强批判模块,同时在推广和更新过程中保持演员输入空白至关重要。对于$\pi_{0.5}$和GR00T N1.5来说,这分别使LIBERO-Plus的非分配成功率提升7.8美元和10.0美元积分,而增强演员则完全崩溃训练。我们研究多种增强类型和优势,并提供提升VLA训练后泛化能力的实用建议。
Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning
自适应专家指导,用于高效策略上的强化学习
- Authors: Daniele Affinita, Ming Xu, Rudolf Reiter, Davide Scaramuzza, Pascal Fua
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06019
- Pdf link: https://arxiv.org/pdf/2610.06019
- Abstract
With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.
- 中文摘要
在大规模并行模拟中,策略上的强化学习方法如PPO已成为许多领域的标准。然而,从零开始学习样本效率低,且无法利用潜在的次优专家存在,如启发式、基于模型的控制器或针对相关任务训练的策略。此类专家通常可用并指导早期训练,但其次优性限制了最终表现。挑战在于如何在专家指导与从奖励中学习之间取得平衡。现有方法通过混合权重、时间表或评估驱动的课程来设定专家影响力。或者,它们通过额外学习组件如批评专家行动或辅助代理进行调整。然而,没有任何方法采用与策略本身相同的策略目标来优化。我们提出一种方法,学习者和专家在每次培训过程中交替控制,专家的控制份额是与策略共同优化的单一可学习参数。学习者在培训初期从专家那里受益,但随着学习者能力提升,其控制份额逐渐减少,最终完全消失。这使得学习者独自行动,优于次优专家。我们利用学习者和基于模型的专家,在两个基准测试中评估了34个任务,涵盖离散和连续动作空间。我们的方法提高了样本效率,优于指导和无指导基线,同时要求极小的超参数变异。随着学习者的进步,专家的控制份额逐渐衰减至零,当专家不再有用时消失。
Scalable Minimal-Change Learning for Controllable Image Editing
可扩展的最小变化学习用于可控图像编辑
- Authors: Shuo Chen, Fengming Huang, Yu Yao, Mingming Gong, Tongliang Liu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06021
- Pdf link: https://arxiv.org/pdf/2610.06021
- Abstract
Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: this https URL
- 中文摘要
图像编辑应仅更改指令指定的属性,保留其他所有属性,但现有方法常常会做出意想不到的更改。我们将这一最小变更原则视为基于指令编辑的优化目标。潜在L1正则化在现代非线性生成器中是输出局部性的良好代理,且通常需要大规模无法实现的监督。我们通过强化学习优化编辑结果。代理视觉语言奖励模型审计每个源图像、指令和编辑图像,识别两种失败类型:未实现的请求更改和非预期的更改。一组级别的评分标准合并并验证这些问题,提供候选编辑间的一致奖励,无需逐指令人工注释。在FLUX.1 Kontext-dev中,ARRO将MinEval、MagicBrush、AnyBench和Emu-Edit的平均编辑分数从5.21提升至5.88。在600个评估样本中,相较于基础编辑器,它减少了8.4%的非目标像素变化。奖励和SFT对照、盲测人类评估以及向OmniGen2的转移提供了补充证据。代码:此 https URL
Grounded Joint-Attention Other-Play for Zero-Shot Coordination
接地、联合注意力、其他对抗以实现零射击协调
- Authors: Giulia Benintendi, Constantin Ruhdorfer, Fabian Kögel, Andreas Bulling
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06025
- Pdf link: https://arxiv.org/pdf/2610.06025
- Abstract
Joint attention - the human ability to share a common visual or cognitive focus with others - enables a meeting of minds that lets us coordinate even with unfamiliar partners. In this work we investigate whether equipping AI agents with a similar mechanism can enable such zero-shot coordination. We introduce Mutual Attention for zero-shot TEaming (MATE): a novel multi-agent reinforcement learning method inspired by human joint attention. MATE encourages agents to coordinate their actions by aligning their visual attention on scene-salient objects during the interaction rather than relying on arbitrary partner-dependent conventions established during training. Unlike symmetry-breaking approaches that merely prevent brittle conventions from emerging, MATE actively promotes coordination through an environment-grounded signal that is naturally shared across partners. We evaluate MATE on three benchmarks: our Card Alignment Game, designed to isolate brittle convention formation, and the more challenging Level-Based Foraging and OvercookedV2 benchmarks. Our experiments consistently show that a joint-attention-inspired signal improves coordination with unknown partners, underlining MATE's potential as a general coordination mechanism that complements and surpasses symmetry-breaking approaches.
- 中文摘要
联合注意力——人类与他人共享共同视觉或认知焦点的能力——促成了思想的融合,使我们即使与陌生的伙伴也能协调。本研究探讨了为人工智能代理配备类似机制是否能实现这种零频率协调。我们介绍了零焦点训练的相互关注(MATE):一种受人类联合注意力启发的新型多代理强化学习方法。MATE鼓励代理在互动过程中将视觉注意力集中在场景显著物体上,而非依赖训练中建立的任意依赖伙伴惯例来协调行动。与仅仅防止脆弱惯例出现的对称破缺方法不同,MATE通过环境基础信号积极促进协调,这种信号在伴侣间自然共享。我们基于三个基准测试评估MATE:设计用于隔离脆弱约定形成的卡片对齐游戏,以及更具挑战性的基于层级采集和过度烹饪V2基准测试。我们的实验持续表明,联合关注启发信号能改善与未知伙伴的协调,强调MATE作为一种通用协调机制的潜力,能够补充甚至超越破缺对称的方法。
Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning
增强对深度强化学习的可转移对抗性攻击
- Authors: Zexin Li, Ruili Yao, Yiming Zeng, Xiaoxue Gao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06083
- Pdf link: https://arxiv.org/pdf/2610.06083
- Abstract
Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
- 中文摘要
大多数对深度强化学习(DRL)的对抗性攻击假设受害者策略具有白盒访问权,但实际上很少成立。本文研究基于转移的黑箱攻击:攻击者对白盒替代代理构建观察扰动,并将其输入未知受害者。我们将攻击表述为每步扰动预算下的返回最小化。首先,我们展示了以每步目标移植可转移图像分类攻击(FGSM、MI-FGSM 和 NI-FGSM)的扰动,扰动虽转移但强度不超过同一预算的随机噪声。随后,我们提出一种轨迹级攻击,通过环境可微模型和温度平滑替代策略,优化沿远离视界的扰动序列,使用相同的优化器。在CartPole-v1中,配备10个DQN和DDQN代理和100对替代-受害者,轨迹级攻击在白盒、跨模型和跨算法设置中优于每步攻击和随机噪声。
I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning
I-BFM:通过无监督强化学习实现的奖励条件强健类人互动
- Authors: Ziqi Han, Yitang Li, Junhan Sun, Fanrong Dong, Yaojie Shen, Lei Ye, Zetong Jing, Yongqi Zhang, Yiming Zhang, Xue Wang, Hao Zhao
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.06129
- Pdf link: https://arxiv.org/pdf/2610.06129
- Abstract
Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.
- 中文摘要
行为基础模型(BFM)最近表明,单一的人形策略可以支持多样化的全身控制,但将这种通用性扩展到物理交互仍然具有挑战性。我们介绍了I-BFM,据我们所知这是首个用于类人-对象交互的BFM。I-BFM不依赖任务特定策略或参考跟踪,而是通过正向-后向表征和无监督强化学习,学习类人、物体及其接触者之间耦合动态的共享表征。在下游任务奖励下,同一策略可以直接基于潜在指令来执行闭环交互,而无需任务特定策略优化。为改善不同时间尺度的交互控制,我们进一步训练策略,同时使用短视野交互目标和长视野目标目标。单一的I-BFM策略执行携带、推送和踢击,同时支持目标达成、动作跟踪、风格控制和长视野任务链。更重要的是,即使与名义执行有较大偏差,I-BFM依然有效:在携带模式下,I-BFM名义成功率为94.3%,机器人倒下后仍保持89.3%,而基于计划的基线为1.3%。Unitree G1的实际实验进一步展示了多样的机车操作行为、交互失败和外部干扰的快速恢复,以及无需任务特定重新训练的任务链。
Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks
基于强化学习的工作负载感知电力传输网络优化
- Authors: Oran Hayes, Maria Pantazi-Kypraiou, Athanasios Tziouvaras, George Stamoulis, Anuj Pathania, Shreejith Shanker, George Floros
- Subjects: Subjects:
Hardware Architecture (cs.AR); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06148
- Pdf link: https://arxiv.org/pdf/2610.06148
- Abstract
Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often resulting in over-provisioned networks and inefficient use of resources. This paper presents a reinforcement learning-based framework for the optimization of workload-aware PDNs. The proposed methodology first generates workload-aware PDNs using architectural power traces obtained from system-level simulations. These power traces are mapped to spatial power density distributions, enabling adaptive allocation of PDN resources according to local current demand. A reinforcement learning agent then performs wire-width optimization to minimize PDN area while maintaining EM and voltage integrity constraints. Electrical and reliability metrics are obtained using SPICE-based circuit analysis and EM lifetime estimation. Experimental evaluation is performed on a dataset of workload-aware PDNs generated from 4-, 8-, and 16-core multiprocessor floorplans using PARSEC and SPLASH-2 benchmark workloads. Furthermore, the proposed Deep Q-Network (DQN)-based optimizer reduces the average normalized PDN area by 47\% while satisfying all EM and IR-drop constraints. Compared to simulated annealing, the proposed approach achieves comparable optimization quality while providing approximately 26$\times$ faster optimization.
- 中文摘要
功率传输网络(PDN)是现代VLSI芯片的关键组成部分,在满足电迁移(EM)和红外降约束的同时,提供稳定的电压水平。传统的PDN设计方法通常依赖最坏情况假设,常导致网络过度配置和资源效率低下。本文提出了基于强化学习的框架,用于优化工作负载感知PDNs。提出的方法首先利用系统级仿真获得的架构功率追踪生成工作负载感知PDN。这些功率跟踪映射到空间功率密度分布,支持根据局部电流需求自适应分配PDN资源。强化学习代理随后执行线宽优化,以最小化PDN面积,同时保持电磁和电压完整性约束。电气和可靠性指标通过基于SPICE的电路分析和电磁寿命估计获得。实验评估基于一个工作负载感知PDN数据集,这些PDN由4核、8核和16核多处理器楼层图生成,使用PARSEC和SPLASH-2基准工作负载。此外,提出的基于深度Q网络(DQN)的优化器在满足所有EM和IR降约束的同时,将平均归一化PDN面积减少了47%。与模拟退火相比,该方法实现了相当的优化质量,同时提供约26美元\倍数的优化速度。
Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
小型语言模型会学会谈判吗?一项针对强化学习培训卖家的受控规模研究
- Authors: Pedro Tabacof, Sagar Joglekar
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06204
- Pdf link: https://arxiv.org/pdf/2610.06204
- Abstract
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.
- 中文摘要
LLM代理开始拥有完整的客户体验。很快,LLM可能分别代表公司和客户进行销售和购买。小型模型在大规模下更具成本效益,但强化学习能否将它们训练成有能力的卖家?我们用GRPO训练了四个Gemma 4检查点(23亿至31亿有效参数),用于双边多议题谈判的程序化效用奖励,并评估同一1152个谈判中的每个臂,对两个培训中未见过的前沿买家进行评估。在相同学习率($10^{-6}$)下,强化学习模型相较基础的收益从23亿美元的+0.001美元升至31亿美元的+0.078美元。每个规模只训练一次,最小的两个检查点采用不同的架构,因此不符合缩放规律。将学习速率提高三倍且训练步骤相同或更少,在所有规模下共享速率提升+0.032美元(23亿美元)至+0.081美元(45亿美元)。在与两个作为卖家运行的前沿模型进行探索性比较中,12B卖家以比两者高出三倍的速率评分训练,尽管其未训练基础得分已高于他们。45亿卖家在该速率下与两者无明显差异,且可安装在一块48GB的GPU上。再用10倍共享速率的23亿臂提升合并分数,但其收益集中在与训练池共享模型族的评估买方。这些结果建议在得出小模型无法学会协商前先调整学习率,并对多个模型家族的买家进行测试。
Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning
受限目标导向平面图生成,结合基于语法的强化学习
- Authors: Nicolas Hochuli, Lorenzo Miele, Kristina Shea, Tino Stankovic
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06244
- Pdf link: https://arxiv.org/pdf/2610.06244
- Abstract
Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during construction. We formulate the generation process as a constrained Markov decision process, where the graph grammar defines the state and action spaces. We further introduce an action projection that maps sampled actions toward state-dependent safe sets, improving constraint satisfaction during training. In contrast to classical graph generators and deep generative models, which typically offer limited goal-directed control or rely on weak constraint satisfaction, our method constructs feasible planar graph embeddings directly during generation. We also introduce a benchmark suite for constrained and goal-directed planar graph generation, together with classical and deep generative baselines. Across all benchmark tasks, our method consistently outperforms baselines while satisfying the formulated constraints.
- 中文摘要
平面图在科学和工程领域的核心应用,但现有生成器在严格的结构和几何可行约束下,对目标导向生成提供了有限的支持。我们提出了一种无数据集的方法,通过将参数图文法与安全强化学习结合生成平面图嵌入,以优化通用任务特定目标,同时满足构建过程中的约束。我们将生成过程表述为受限马尔可夫决策过程,图语法定义状态空间和作用空间。我们还引入了一种动作投影,将采样动作映射到状态依赖的安全集,从而提升训练中的约束满足度。与传统图生成器和深度生成模型通常提供有限的目标导向控制或依赖弱约束满足度不同,我们的方法在生成过程中直接构建可行的平面图嵌入。我们还引入了一套基准测试套件,用于约束和目标导向平面图生成,以及经典和深度生成基线。在所有基准任务中,我们的方法在满足既定约束的同时,始终优于基线。
Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments
通过混合状态深度强化学习在部分可观察的联网车辆环境中实现匝道计量控制
- Authors: Youcef Mehamlia, Nadir Farhi, Meriem Bouali
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2610.06266
- Pdf link: https://arxiv.org/pdf/2610.06266
- Abstract
Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic data. Connected vehicles (CVs) provide vehicle-level observations that can complement aggregate traffic measurements, but their limited penetration produces incomplete microscopic information. This paper proposes a hybrid observation representation combining macroscopic traffic measurements with a two-channel grid encoding observed CV presence and speed. A Dueling Double Deep Q-Network processes these inputs to select ramp-metering green durations. The controller is trained under varying traffic demands and CV penetration rates and evaluated against ALINEA and macroscopic-only DRL variants in SUMO. Across 50 matched evaluation scenarios, the hybrid controller under partial CV visibility reduces the reported total travel time by 11.4 % and mean spillback duration by 84.9 % relative to ALINEA. Evaluating the same trained policy with full CV visibility yields a further travel-time reduction of approximately 1.6 %. Analysis across penetration rates suggests that the performance gap decreases as microscopic observations become more complete. These results support the use of complementary macroscopic and sparse microscopic observations for learning-based ramp metering. The source code implementation of the model is available at: this https URL
- 中文摘要
高速公路匝道合流是拥堵的主要原因,带来显著的经济和环境成本。虽然深度强化学习(DRL)为匝道计量提供了有前景的解决方案,但现有方法主要依赖汇总的宏观数据。连接车辆(CV)提供车辆层面的观测,可以补充总体交通测量,但其有限的渗透率会产生不完整的微观信息。本文提出了一种混合观测表示,结合宏观交通测量与编码观察到CV存在和速度的双通道网格。一个对立的双深度Q网络处理这些输入,选择匝道计量绿灯时长。控制器在不同的交通需求和CV渗透率下进行训练,并在SUMO中与ALINEA及仅宏观的DRL变体进行评估。在50个匹配的评估场景中,混合控制器在部分CV可视性下,相较ALINEA将报告的总行程时间减少11.4%,平均溢出持续时间减少84.9%。在完全CV可视性下评估同一训练策略,进一步减少约1.6%的行程时间。跨渗透率分析表明,随着微观观测的更完整,性能差距会缩小。这些结果支持使用互补的宏观和稀疏显微观测来进行基于学习的斜坡计量。模型的源代码实现可在:此 https URL 获取
GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories
GAMBIT:学习规划连续多机器人轨迹
- Authors: Rishabh Jain, Akmaral Moldagalieva, Lorenzo Magnino, Michael Amir, Keisuke Okumura, Ajay Shankar, Wolfgang Hönig, Amanda Prorok
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2610.06290
- Pdf link: https://arxiv.org/pdf/2610.06290
- Abstract
GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.
- 中文摘要
GAMBIT是一种开局国际象棋走法,玩家牺牲一个棋子,通常是兵,以便在游戏后期获得位置优势。类似地,在多机器人协调中,单个机器人可能需要放弃局部最大化奖励的行为,以提升整体团队表现。这种自我牺牲行为在人工设计的启发式中难以捕捉,尤其是在密集且交互丰富的环境中。本研究聚焦双积分器连续动力学,研究如何在多机器人轨迹执行中学习此类协调启发式的运动原语。我们的框架GAMBIT首先通过模仿学习协调运动原语选择,随后通过强化学习微调策略。我们还引入了一种带有备份轨迹的安全展开机制,确保始终无碰撞执行。实验表明,GAMBIT在包括集中式运动规划器和去中心化被动规划器在内的多种基线中表现显著优越,同时展现出强大的可扩展性。特别是,它协调超过一千台在连续域中规划延迟低于几百毫秒的机器人。
VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
VepAgent:通过工具增强强化学习桥接因果过渡,用于视频事件预测
- Authors: Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang, Fangfang Lin, Xinyi Long, Yin Wang, Zijian Song, Rui Lan, Shi Qiu, Boyuan Pan, Yang Luo, Yuyin Zhou
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.06293
- Pdf link: https://arxiv.org/pdf/2610.06293
- Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
- 中文摘要
多模态大型语言模型(MLLMs)在视频理解方面展现出显著潜力,但其对回顾性总结和以文本为中心的先验,常常限制了其在视频事件预测(VEP)中弥合未观察到因果转移的能力。为此,我们提出了VepAgent,一种将因果转移推理与工具增强强化学习(RL)整合的智能框架,实现了稳健的VEP过程。与以往被动地从历史依赖预测未来轨迹的方法不同,我们的方法明确建模了从终端观察状态到未来事件的逻辑进程。具体来说,我们首先构建了futurebench-4K,这是一个高质量的思维链数据集,用于监督微调(SFT),通过结构化未观察到的中间状态推断,有效弥合了因果逻辑的鸿沟。随后,我们开发了整合状态追踪、帧检索和区域放大的诊断工具库,使智能体能够动态地利用外部工具增强推理,恢复缺失的时空证据并解决推理过程中的视觉歧义。此外,我们提出了一种综合奖励机制,共同优化预测准确性、因果一致性和可靠先验,迫使智能体依赖真实的视觉基础,而非表面文本相似性。对FutureBench和NEPBench数据集的广泛评估表明,我们的方法实现了最先进的性能,显著优于大型MLLMs,并验证了我们代理性、面向未来的推理范式的实证有效性。
Graph Neural Network-Driven Deep Reinforcement Learning for Scalable RIS Allocation
图神经网络驱动深度强化学习,实现可扩展RIS分配
- Authors: Martin Mark Zan, Stefan Schwarz
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2610.06295
- Pdf link: https://arxiv.org/pdf/2610.06295
- Abstract
Reconfigurable Intelligent Surfaces (RISs) offer a promising paradigm to mitigate blockage and extend millimeter-wave coverage in 6G multi-cell networks. However, dynamically allocating shared RIS infrastructure across competing base stations is an NP-hard problem posing severe scalability bottlenecks. In this paper, we propose a scalable framework combining Graph Neural Networks (GNNs) with Deep Reinforcement Learning (DRL) for dynamic shared RIS orchestration. By formulating allocation as a Markov Decision Process, we introduce a physical topology sparsification strategy that prunes dense channel matrices into a sparse tripartite graph. This pruning reduces edge density by 77% and removes representation noise, thereby improving global coverage probability while reducing computational complexity. Our relational message-passing architecture naturally generalizes to arbitrary network dimensions without model retraining. Furthermore, structural ablation studies reveal that physical path loss localizes surface dependencies, enabling a highly efficient localized graph design with linear computational scaling. Extensive simulations in dense urban environments demonstrate that under strict infrastructure budget constraints, the proposed GNN-DRL framework consistently outperforms greedy heuristic baselines by up to ~12% in coverage while delivering faster inference speed via GPU acceleration.
- 中文摘要
可重构智能表面(RIS)为缓解阻塞和扩展6G多小区网络中的毫米波覆盖提供了有前景的范式。然而,动态分配共享RIS基础设施在竞争基站之间是一个NP难问题,带来了严重的可扩展瓶颈。本文提出了一个结合图神经网络(GNN)与深度强化学习(DRL)的可扩展框架,用于动态共享RIS编排。通过将分配表述为马尔可夫决策过程,我们引入了物理拓扑稀疏化策略,将密集信道矩阵修剪成稀疏的三分图。这种剪枝降低了77%的边密度,消除了表示噪声,从而提高了全局覆盖概率,同时降低了计算复杂度。我们的关系式消息传递架构自然地推广到任意网络维度,无需模型重新训练。此外,结构消融研究表明,物理路径损耗能局部化表面依赖,实现线性计算缩放的高效局部图设计。在密集城市环境中的大量模拟表明,在严格的基础设施预算约束下,所提出的GNN-DRL框架在覆盖率上持续优于贪婪启发式基线多达12%,同时通过GPU加速实现更快的推理速度。
RollPlace: Improving Macro Placement via Monte Carlo Rollout Search
RollPlace:通过蒙特卡洛推送搜索提升宏置位置
- Authors: Qi Zhou, Guojun Liu, Guangzhi Qi, Ming Lu, Jiechu Liu, Zhongli Liu, Jianqun Yang, Xingji Li
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06316
- Pdf link: https://arxiv.org/pdf/2610.06316
- Abstract
The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, they predominantly focus on generating optimal layouts in a single attempt, often producing solutions that require subsequent refinement. To address this limitation, we propose RollPlace, a novel and generalized macro placement framework. RollPlace adopts a two-stage optimization strategy: generating initial placement solutions via machine learning methods or heuristic-based strategies, and refining these layouts efficiently by adjusting specific macros derived from the initial stage. This strategy circumvents the sequential generation constraints inherent in traditional RL-based placement methods. Furthermore, RollPlace seamlessly integrates Monte Carlo Tree Search (MCTS) to balance exploration and exploitation, and employs a rollout mechanism for efficient local search. Extensive experiments on the ISPD 2005 benchmark demonstrate that RollPlace outperforms state-of-the-art methods. Additionally, end-to-end experimental results based on OpenROAD across 19 benchmarks show that RollPlace excels in multiple metrics. The proposed framework offers a robust and scalable solution for addressing the growing complexity of modern chip design challenges.
- 中文摘要
强化学习(RL)在电子设计自动化(EDA)中的应用,特别是芯片放置方面,近年来引起了广泛关注。虽然现有基于机器学习的方法取得了显著进展,但它们主要专注于一次性生成最优布局,常常产生需要后续细化的解决方案。为解决这一限制,我们提出了RollPlace,一种新颖且通用的宏置框架。RollPlace采用两阶段优化策略:通过机器学习方法或启发式策略生成初始布局解,并通过调整初始阶段衍生的特定宏高效优化布局。该策略绕过了传统基于强化学习的部署方法固有的顺序生成限制。此外,RollPlace无缝集成了蒙特卡洛树搜索(MCTS),以平衡探索与利用,并采用滚动机制实现高效的局部搜索。基于ISPD 2005基准测试的大量实验表明,RollPlace优于最先进的方法。此外,基于OpenROAD的19个基准测试端到端实验结果显示,RollPlace在多项指标上表现出色。该框架为应对日益复杂化的现代芯片设计挑战提供了稳健且可扩展的解决方案。
CRAFTER: Causality-based Self-adaptation for Autonomous IoT Systems
CRAFTER:基于因果律的自主物联网系统的自我适应
- Authors: Houssam Hajj Hassan (IP Paris,SAMOVAR), Ajay Kattepur, Denis Conan (IP Paris,TSP - INF,ACMES-SAMOVAR), Georgios Bouloukakis (IP Paris,TSP - INF,ACMES-SAMOVAR,ECE)
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06320
- Pdf link: https://arxiv.org/pdf/2610.06320
- Abstract
This paper presents CRAFTER, an automated framework for designing and deploying self-adaptive IoT systems using Causal Reinforcement Learning (CRL). As IoT devices increasingly populate pervasive computing spaces, smart environments are enabled with advanced monitoring and interactive services. The dynamic nature of these environments, such as fluctuating workloads and evolving application demands, poses significant challenges in maintaining consistent Quality of Service (QoS) levels of IoT applications. While existing self-adaptation techniques offer adaptive capabilities, they are often designed to deal with specific application domains, hindering the design of self-adaptive solutions that can be re-used across multiple IoT verticals. In addition, there is a lack of automated pipelines that act on identifying key performance drivers to take effective adaptation decisions. CRAFTER addresses these issues by using Causality as a formal framework for performance analysis of IoT systems. CRAFTER generates causal graphs to uncover dependencies among system components and guide adaptation decisions based on cause-effect relationships. Then, adaptation agents can leverage this knowledge to take more effective adaptation decisions in dynamic situations. Our experimental evaluation demonstrates how CRAFTER enables deriving causal graphs spanning diverse IoT use cases. Furthermore, we showcase how CRAFTER improves self-adaptation performance by 25% compared to state-of-the-art Reinforcement Learning-based approaches.
- 中文摘要
本文介绍了CRAFTER,一个利用因果强化学习(CRL)设计和部署自适应物联网系统的自动化框架。随着物联网设备日益普及的计算空间,智能环境通过先进的监控和交互服务得以实现。这些环境的动态特性,如工作负载波动和应用需求演变,给维持物联网应用服务质量(QoS)水平带来重大挑战。虽然现有的自适应技术具备自适应能力,但通常设计用于特定应用领域,阻碍了可跨多个物联网垂直领域重复使用的自适应解决方案设计。此外,缺乏自动化流水线识别关键性能驱动因素以做出有效适应决策。CRAFTER通过使用因果关系作为物联网系统性能分析的形式框架来解决这些问题。CRAFTER生成因果图,揭示系统组件间的依赖关系,并基于因果关系指导适应决策。随后,适应代理可以利用这些知识在动态情境下做出更有效的适应决策。我们的实验评估展示了CRAFTER如何能够推导跨越多种物联网用例的因果图。此外,我们展示了CRAFTER如何将基于强化的自我适应性能提升25%,相比最先进的基于强化学习的方法。
Dynamic Minimax Regret Optimization for Robust LLM Post-Training
动态极大后悔优化,用于稳健的大型语言模型后训练
- Authors: Chengbo Zang, Haoyu Dong, Mehmet Kerem Turkcan, Gil Zussman, Zoran Kostic, Javad Ghaderi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06329
- Pdf link: https://arxiv.org/pdf/2610.06329
- Abstract
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an optimizer updates the model parameters using stochastic gradients from the selected source. We focus on the practically restrictive setting where source losses evolve with model training but historical data are not re-evaluated, requiring the sampler to track instantaneous worst-sources from stale partial feedback. We propose DUCB-OGD, a simple and scalable algorithm that couples a Discounted Upper-Confidence-Bound sampler with an Online Gradient Descent optimizer. The sampler maintains exponential moving average loss estimates and confidence radii based on discounted effective sample sizes, avoiding costly re-evaluation of past data or intrusive changes to standard training pipelines. For $K$ data sources and $T$ training steps, we prove that DUCB-OGD achieves a dynamic minimax regret of $\tilde{O}(K^{1/4}T^{3/4})$, which is optimal up to logarithmic factors for the undiscounted objective under our feedback model. Extensive experiments across supervised fine-tuning, preference optimization, and reinforcement learning show that DUCB-OGD integrates seamlessly into modern LLM training pipelines and improves worst-group robustness with negligible computational overhead compared with standard sampling baselines.
- 中文摘要
现代大型语言模型训练越来越依赖跨越不同领域、任务、偏好分布和难度等级的异构数据源。我们研究了在即时迷你批仅bandit反馈下,组分布稳健LLM训练后的动态极小极大遗憾。该框架将训练视为双人采样器优化过程:采样器利用bandit反馈自适应地从数据源中选择,优化器则利用随机梯度更新模型参数。我们关注一种实际限制性设定,即源损失随模型训练演变,但历史数据不被重新评估,要求采样器从陈旧的部分反馈中追踪即时最差源。我们提出了DUCB-OGD,一种简单且可扩展的算法,将折扣上置信度界限采样器与在线梯度下降优化器结合。采样器基于折现有效样本量维持指数移动平均损耗估计和置信半径,避免了对过去数据的昂贵重新评估或对标准训练流程进行侵入性修改。对于$K美元数据源和$T美元训练步骤,我们证明DUCB-OGD实现了动态极小极大遗憾值$\tilde{O}(K^{1/4}T^{3/4})$,在我们的反馈模型下,该指标在对数因子下最优。涵盖监督微调、偏好优化和强化学习的广泛实验表明,DUCB-OGD无缝集成于现代LLM训练流水线,且与标准采样基线相比,计算开销可忽略不计,提升最恶组鲁棒性。
MeSD: Multi-Evidence Self-Distillation for VideoLLM
MeSD:视频LLM的多证据自我提炼
- Authors: Weijie Zhu, Han Fang, Hanyu Fu, Yuzhe Zhang, Xin Wei, Zhaoyan Pan, Feiran Liu, Xunjie Jin, Hongbo Sun, Zhiyu Lin, Tianyi Gao, Tianyi Ding, Ye Yuan, Zhongjiang He, Hao Sun, Zhiheng Wu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06342
- Pdf link: https://arxiv.org/pdf/2610.06342
- Abstract
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.
- 中文摘要
虽然带有可验证奖励的强化学习为视频大型语言模型提供了可靠的结果监督,但序列级奖励提供的代币级指导有限。策略上自我蒸馏通过让自学者在特权信息上获得密集的代币级监督来解决这一限制。然而,在单一教师语境中汇聚异质证据会掩盖跨证据的共识和冲突。另一个挑战在于确定教师指导应优化基于奖励的更新,还是对失败轨迹提供纠正性监督。为解决这些问题,我们提出了MeSD,一种面向视频LLM的多证据自我蒸馏框架。MeSD构建了三个具有共享参数的证据条件教师,使用真实答案作为通用语义语境,同时分别纳入时间和空间证据。在相同学生生成的前缀下,MeSD评估了与答案教师相关的证据特定偏好,并将教师共同偏好与教师特定残差的门控差异融合。此外,MeSD引入了验证引导优化,将轨迹分类为成功、失败或不确定。对于成功和不确定轨迹,MeSD在保持奖励衍生符号的同时,细化了代币层面的优势幅度。对于包含所需证据的已验证失败轨迹,MeSD应用失败条件蒸馏,并对融合分布进行逆KL修正。在多个视频基准测试上的实验显示,相较于强化学习和自蒸馏基线,实现了持续的提升。
Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering
本体概念重叠作为训练信号:基于知识的强化学习用于临床问答
- Authors: Aditya Tanna, Abhishek Jindal
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06360
- Pdf link: https://arxiv.org/pdf/2610.06360
- Abstract
Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1) gives a graded, externally specified reward computed without a model in the loop. We combine it inside GRPO with an entropy-normalised LLM judge, which covers the safety and evidence axes overlap cannot see, and a small consistency penalty on padding and repetition that keeps early-training samples scorable. This three-term composite improves over SFT on Phi-3-mini (3.8B) over MedQA by 2.9% on EM (0.700 vs 0.680) and 39% on Token-F1 (0.202 vs 0.145); on Llama-3.2-3B the corresponding gains are 14% on EM and 35% on Token-F1. We report Token-F1 as the primary metric because it credits partially-correct clinical content that EM discards at this open-generation scale. Main-table results are means over 3 seeds with standard deviations below 0.005. The method transfers to PubMedQA, where training on the PubMedQA train set with the same composite reward improves Token-F1 over SFT by 22% on Phi-3-mini and 17% on Llama-3.2-3B without retuning. A reward ablation on Phi-3, varying the judge-ontology split at a fixed consistency weight, attributes 3 EM points to the ontology term, the contribution that catches entity substitutions the judge cannot. Three negative findings constrain the design: DPO under random negatives underperforms SFT for strong-prior models but helps the weakest-prior one; PPO under a sparse neural reward diverges; GRPO with KL-in-loss collapses at 7B.
- 中文摘要
语言模型的强化学习训练后依赖两种奖励设计:人类偏好(RLHF、DPO)和二元验证器(RLVR)。临床问题回答不符合这两者。接近正确答案仅差一个替换实体,且无可执行检查决定临床正确性。我们从维护的受控词汇中实例化软验证器:UMLS 概念唯一标识符重叠(通过 scispaCy,集合级 F1)给出一个分级、外部指定的奖励,无需循环中模型计算。我们将它与熵归一化的 LLM 评判结合,覆盖重叠轴无法看到的安全性和证据,以及对填充和重复的少量一致性惩罚,保持早期训练样本可评分。该三项综合指标在Phi-3-mini(3.8B)上较SFT提升2.9%(0.700对0.680),Token-F1提升39%(0.202对0.145);Llama-3.2-3B对应EM提升14%,Token-F1提升35%。我们将Token-F1作为主要指标,因为它将EM在开放生成尺度下部分正确临床内容归功于此。主表结果为3个种子均值,标准差低于0.005。该方法可转移至PubMedQA,在PubMedQA列集上,使用相同复合奖励的训练,Token-F1在Phi-3-mini上提升22%,在Llama-3.2-3B上提升17%,且未重新调谐。对Phi-3进行奖励消融,在固定一致性权重下改变法官-本体分裂,为本体项赋予3个EM点,即识别评判无法捕捉的实体替代的贡献。设计受到三个负面限制:在随机负条件下,DPO在强先验模型中表现逊于SFT,但对最弱先验模型有利;在稀疏神经奖励下,PPO发散;带有KL损失的GRPO在7B崩溃。
Visual Swarm Navigation via Deep Reinforcement Learning and Evolutionary Hybrid Design
通过深度强化学习和进化混合设计实现可视化群体导航
- Authors: Álvaro Díez (Department of Computer Science and Artificial Intelligence, University of Alicante), Fidel Aznar (Department of Computer Science and Artificial Intelligence, University of Alicante)
- Subjects: Subjects:
Robotics (cs.RO); Multiagent Systems (cs.MA); Neural and Evolutionary Computing (cs.NE)
- Arxiv link: https://arxiv.org/abs/2610.06400
- Pdf link: https://arxiv.org/pdf/2610.06400
- Abstract
Swarm robotics presents a robust and cost-effective paradigm for advanced automation in complex, dynamic environments, such as those encountered in search and rescue or environmental monitoring. A fundamental challenge for this field is the data-driven design of decentralized controllers capable of generating emergent collective behaviors. This paper proposes a novel, AI-driven hybrid methodology for the automatic synthesis of swarm robotic controllers for autonomous visual navigation. This approach synergistically combines multi-agent reinforcement learning with neuro-evolutionary strategies, specifically leveraging implementations of the cross-entropy method and the covariance matrix adaptation evolution strategy to optimize a pre-trained individual navigation policy. The underlying deep architecture is engineered for low-cost, resource-constrained platforms, utilizing a compact neural network that relies exclusively on monocular camera imagery. This vision-based design emphasizes computational and energy efficiency, a critical requirement for practical swarm deployments. Experiments, performed in a high-fidelity physics simulator, demonstrate that the resulting controllers enable robust and scalable collective exploration of diverse indoor environments. The controller trained using our cross-entropy method achieves superior exploration coverage, visiting 36.20% more regions compared to the covariance matrix adaptation evolution strategy. Critically, our best vision-based policy achieves exploration performance statistically comparable to traditional methods relying on more expensive distance sensors, while delivering a significant 31.40% average reduction in energy consumption. These findings validate an effective and economically viable autonomous control system, establishing a path for deploying highly efficient collective intelligence in real-world engineering applications.
- 中文摘要
群体机器人为复杂、动态环境中的先进自动化提供了稳健且经济高效的范式,如搜救或环境监测。该领域的根本挑战是数据驱动设计能够生成涌现集体行为的去中心化控制器。本文提出了一种新颖的人工智能驱动混合方法,用于自动合成群体机器人控制器,实现自主视觉导航。该方法协同结合多智能体强化学习与神经进化策略,特别是利用交叉熵法和协方差矩阵适应进化策略,优化预训练的个体导航策略。其底层深度架构设计为低成本、资源受限的平台,采用仅依赖单目摄像头图像的紧凑神经网络。这种基于视觉的设计强调计算和能效,是实际群体部署的关键需求。在高保真物理模拟器中进行的实验表明,所产生的控制器能够实现对多样室内环境的稳健且可扩展的集体探索。使用我们交叉熵方法训练的控制器实现了更优越的探索覆盖范围,访问区域比协方差矩阵适应演化策略多36.20%。关键是,我们最优的基于视觉的策略在统计上实现了与依赖更昂贵距离传感器的传统方法相当的探索性能,同时实现了显著的平均能耗降低31.40%。这些发现验证了一种有效且经济可行的自主控制系统,为在现实工程应用中部署高效集体智能奠定了路径。
The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
辅助困境:通过多回合强化学习学习教学
- Authors: Jakub Macina, Manu Kapur, Mrinmaya Sachan
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06446
- Pdf link: https://arxiv.org/pdf/2610.06446
- Abstract
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
- 中文摘要
训练用来回答问题的大型语言模型(LLMs)天生教学能力较差。对模拟学生进行强化学习(RL)是改善教学法的有前景方法,但现有RL训练导师会奖励学生在辅导题上的成功,且导师的话仍置于上下文中。奖励最容易提高的方式是告诉学生答案,而需要调整惩罚来减少讲述。借鉴学习科学,我们引入了掩蔽近迁移测试:学生在辅导问题的一个看不见变体上进行测试,导师的言语被掩盖,因此奖励只能通过学生自己回合的写作提升。这抑制了学生的认知分担,允许连续惩罚被两个二元奖励门(导师回答的事实正确性,无解答交接)取代。一个省略的消融表明,单靠学习-增益奖励并不区分教学与讲述:门减少了解的交接,而测试后的近迁移则改善了域外转移。利用这些奖励设计,我们开发了Eduardo,一种多回合的强化学习方案,用于训练LLM导师,并用它训练两个不同LLM架构的4B、9B、14B和27B模型。我们的后训练Eduardo-27B模型在MathTutorBench上的Gemini-3.1-Pro和TutorMoments上的Claude Opus 4.8在思考标记上少2.4-6.2倍,这对互动辅导非常重要。虽然未在奖励中点名,但该模型在推动辩护教师行动的使用量上增加了一倍多,同时支持度下降(例如布置独立作业)的效果超出单一问题对话集,被训练淘汰。我们将培训环境开源,包含8671个问题的近迁移数据集,并训练模型以供进一步开发。
GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
GPlaceRL:一个用于详细定位的开源图强化学习框架
- Authors: Pavlos Stoikos, Foteini Oikonomou, Christos Poulos, Maria Pantazi-Kypriou, Athanasios Tziouvaras, Christos Anagnostopoulos, Georgios Karakonstantis, George Floros
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.06489
- Pdf link: https://arxiv.org/pdf/2610.06489
- Abstract
Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning, macro placement, or global placement, while detailed placement refinement remains relatively unexplored. In this paper, we present GPlaceRL, an open-source graph reinforcement learning framework for detailed placement refinement. GPlaceRL represents legalized placements as graphs and provides a modular environment for studying graph encoders, policy architectures, reward formulations, and local placement actions. To demonstrate the capabilities of GPlaceRL, we conduct a systematic evaluation of proximal policy optimization (PPO) policies with graph attention network (GAT) encoders in a per-design optimization setting. Across five placement benchmarks, the best greedy evaluation results achieve HPWL improvements ranging from $3.27\%$ to $32.87\%$. The results highlight the importance of compact GAT architectures and flexible local action spaces for placement optimization. Overall, GPlaceRL provides a reproducible and extensible framework for systematic research on RL-based detailed placement refinement.
- 中文摘要
强化学习(RL)已成为一种有前景的布局优化方法,尤其是在与捕获电路连接性的图神经网络(GNN)结合时。然而,大多数基于学习的布局方法侧重于楼层规划、宏观布局或全局布局,而详细布局细化仍相对较少被探索。本文介绍了GPlaceRL,一个开源的图强化学习框架,用于详细布局细化。GPlaceRL将合法的布局表示为图,并为研究图编码器、策略架构、奖励表述和局部布局动作提供了模块化环境。为展示GPlaceRL的能力,我们系统评估了基于设计优化的图形注意力网络(GAT)编码器的近端策略优化(PPO)策略。在五个配置基准测试中,最佳贪婪评估结果实现了从3.27美元到32.87%美元不等的HPWL改进。结果强调了紧凑的GAT架构和灵活的局部动作空间在布局优化中的重要性。总体而言,GPlaceRL为基于强化学习的详细布局优化系统研究提供了可重复且可扩展的框架。
MIRT: Transformers for Truthful Generative Auctions with Whole-feed Permutation Externalities
MIRT:适用于真实生成拍卖的全给置换外部性转换器
- Authors: Ali Elahi, Ermis Soumalias, Jason Cheuk Nam Liang, Daniel Yao, Michael J. Curry
- Subjects: Subjects:
Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06559
- Pdf link: https://arxiv.org/pdf/2610.06559
- Abstract
Modern online platforms commonly rank ads and organic content separately before blending them into a feed displayed to the user, overlooking externalities: an item's click-through rate depends on its surrounding content, not only on its own position. Recent learning-based feed generation mechanisms model some of these cross-type interactions to globally optimize for the whole feed's welfare. However, these approaches either fix the ordering of organic content, or lack exact strategyproofness guarantees for bidders. To combat these shortfalls, we introduce the Maximal-in-Range Transformer (MIRT) mechanism class, which uses a transformer to generate a range of candidate feeds that jointly order ads and organic content, and selects the welfare-maximizing feed in the range. However, there is a tension: strategyproofness requires the generated range to be bid-independent, even though a candidate feed's welfare depends linearly on the bids. Our key technical contribution is a reinforcement learning approach that incorporates both candidate generation and bid-aware selection into training, enabling a bid-independent transformer to learn to generate high-welfare ranges by accounting for both individual feed quality and the collective quality of the range. Additionally, we bound the pseudo-dimension of the MIRT class under hard attention, showing that near-optimal expected welfare is learnable with sample complexity polynomial in the transformer size and only logarithmic in the range size. Empirically, MIRT outperforms the previous non-strategyproof state-of-the-art feed models while remaining exactly strategyproof. Our results show that transformer-based auctions can deliver externality-aware whole-feed optimization without sacrificing exact incentive compatibility, removing a major obstacle to their practical deployment.
- 中文摘要
现代在线平台通常会先将广告和自然内容分别排名,然后将它们混合到用户面前的推送中,忽略了外部性:一条条目的点击率取决于其周围内容,而不仅仅是其自身位置。近期基于学习的推送生成机制模拟了部分跨类型交互,以全局优化整个推送的福利度。然而,这些方法要么固定了自然内容的排序,要么缺乏竞标者的精确策略性保证。为解决这些不足,我们引入了“最大范围内变换器”(MIRT)机制类,利用变换器生成一系列候选推送,这些推流共同排序广告和自然内容,并在该范围内选择最大化福利的推送。然而,存在一种矛盾:策略无效性要求生成的范围与竞标无关,尽管候选推送的福利度线性依赖于竞标。我们的关键技术贡献是一种强化学习方法,将候选人生成和竞标感知选择结合训练,使竞标无关的变换器能够通过考虑单个馈源质量和范围的整体质量,学习生成高福利范围。此外,我们对MIRT类的伪维度进行了严格约束,表明近似最优的期望福利在变压器大小中为样本复杂度多项式时是可学习的,而在范围规模中仅为对数。从实证角度看,MIRT优于之前非策略防护的先进馈源模型,同时保持完全策略安全。我们的结果表明,基于变换器的拍卖能够实现外部性感知的全馈优化,同时不牺牲精确激励兼容性,消除了其实际部署的主要障碍。
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches
LoGRA:利用低阶梯度草图扩展大型语言模型强化学习
- Authors: Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.06647
- Pdf link: https://arxiv.org/pdf/2610.06647
- Abstract
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7\% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{this https URL}{Molt library}.
- 中文摘要
强化学习(RL)极大地提升了大型语言模型(LLM)的能力,但其内存需求仍是更广泛采用的障碍。我们引入了LoGRA,这是一种通过保留低秩梯度草图中有用学习信号来减少内存的强化学习后训练方法。这些紧凑的表示支持模型更新和高效的策略同步。为防止过大更新干扰学习,我们用预测的KL步进控制来补充梯度压缩,后者在每次更新前估计策略变化并相应调整幅度。在推理任务中,LoGRA在不牺牲性能的情况下,平均训练内存减少了高达45.7%。它还支持在单一八GPU节点上稳定训练一个27B参数模型,持续超过1100步,而密集的Adam内存耗尽,使此前内存难以实现的强化学习变得可行。代码可在\href{this https URL}{Molt library}中获得。
Considering Context: When World Models Need Context Encoders
考虑上下文:当世界模型需要上下文编码器时
- Authors: Oleg Smirnov, Sofiane Ennadir, John Pertoft, Bjartur Hjaltason, Sara Karimi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06651
- Pdf link: https://arxiv.org/pdf/2610.06651
- Abstract
Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that quantity into a history-recoverable part, a residual requiring the true context, and the deficit added by a finite model. We classify context-aware algorithms by the predictive risk their conditioning set can target and demonstrate across environments of increasing identification difficulty that the headroom does not follow the MDP class. The same task under different priors leaves predictive headroom in one setting and nothing distinguishable from zero in another, where the agent's behavior implicitly identifies the context and any benefit of such a mechanism cannot be attributed to missing information. Where headroom persists, the learned state exposes it only partially, and adding the true context still lowers the risk. Our contribution is a practical criterion for matching contextual mechanisms to the information available to them, estimated from the ordinary trained agent without a reference policy.
- 中文摘要
基于模型的强化学习中的泛化方法通常假设代理无法从自身经验中恢复支配环境动态的潜在上下文,因此该上下文会向外部提供。我们用\emph{predictive sufficiency}形式化并检验这一假设,该假设量化了对上下文访问对下一步预测的贡献,并将该量分为可恢复历史的部分、需要真实上下文的残差和有限模型所增加的缺口。我们根据其条件反射集能够针对的预测风险进行分类,并在识别难度不断增加的环境中证明,headroom不遵循MDP类。同一任务在不同先验下会在不同环境中留下预测余量,而在另一种环境中则无可区分的余量,而在另一种情况下,代理的行为隐含识别了上下文,且此类机制的任何益处都不能归因于缺失的信息。当余裕存在时,学习到的状态仅部分暴露,加入真实上下文仍能降低风险。我们的贡献是将上下文机制与其可用信息匹配的实用标准,这些信息由无参考策略的普通训练代理估算。
Reward Stealing Attack on Large Language Models
对大型语言模型的奖励窃取攻击
- Authors: Jiaming Qian, Pengyang Zhou, Jiahe Xu, Chaochao Chen
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.06670
- Pdf link: https://arxiv.org/pdf/2610.06670
- Abstract
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at this https URL.
- 中文摘要
对大型语言模型(LLMs)的对抗性攻击旨在诱导有害内容。然而,现有方法存在较高的计算成本或严格的模型配对依赖,限制了其可扩展性和可转移性。我们提出了奖励窃取攻击(ReSA),这是一种针对LLM对齐背后潜在安全奖励的对抗性攻击框架。ReSA采用最大熵逆强化学习,仅从对齐模型的行为中恢复代理奖励模型。提取的奖励随后在推理时被反转,以推导出对抗策略,并通过奖励引导解码机制高效实现。实验表明,单个回收的奖励会在提示和多样模型间泛化,揭示根本的对齐漏洞,使ReSA在有效性和可转移性上显著优于现有攻击。该代码可在此 https URL 获取。
MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks
MedPrune:医疗VQA任务中的拓扑高效多模态多代理通信演进
- Authors: Jiuheng Wan, Runze Li, Chen Chen, Tingyuan Hu, Daiyang Yu, Yimin Jing, Taolin Zhang, Richang Hong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.06695
- Pdf link: https://arxiv.org/pdf/2610.06695
- Abstract
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.
- 中文摘要
虽然医学多模态大型语言模型(Med-MLLM)推动了医疗视觉问答(VQA),但现有受临床工作流程启发的多智能体框架仍存在交互模式和冗余通信拓扑导致的过高计算开销。本文提出了MedPrune,一种高效的医疗多模态多智能体协作框架,动态修剪通信拓扑中的节点和边,以提升推理能力和令牌效率。具体来说,我们首先将诊断过程表述为异构通信图,节点代表各科室的专业代理,边缘捕捉部门内外的交互。基于该图,我们引入了两种稀疏化机制以实现自适应协作进化:(1)异构节点稀释,通过强化学习驱动的拓扑优化消除与当前多模态问题无关的任务相关专业代理;(2)异构边缘稀疏化,通过联合优化任务性能和拓扑复杂性,选择性保留最具诊断意义的部门内外连接。在全套和少数样本训练设置下的大量医学VQA实验证明,MedPrune超越多智能体基线,并以强对抗鲁棒性提升令牌效率。
Improving Diversity in LLM Short Story Generation
提升LLM短篇故事生成的多样性
- Authors: Zahra Solati Dehkordi, Vasileios Lampos
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.06729
- Pdf link: https://arxiv.org/pdf/2610.06729
- Abstract
Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diversity across these dimensions, we introduce DivLM, an LLM post-training framework consisting of two phases. First, we perform continued pre-training on a creative writing corpus and restore instruction-following capabilities using weight residuals. We then apply reinforcement learning with a custom, composite reward function that jointly maximizes diversity across the targeted narrative dimensions while maintaining response quality. Our empirical results on two LLM families show that DivLM increases diversity metrics by more than 9% on average compared to alternative approaches, while preserving instruction following, overall response quality, and similarity to human outputs.
- 中文摘要
大型语言模型(LLMs)能够生成准确的回答,但这些模型缺乏多样性。我们尝试在创意短篇故事生成中解决这个问题。借助既定的写作惯例和已知的LLM局限性,我们针对体裁、语气、风格和命名实体的变异。为促进这些维度的多样性,我们引入了DivLM,一种由两个阶段组成的LLM后训练框架。首先,我们对创意写作语料库进行持续预训练,利用权重残差恢复指令跟随能力。然后我们应用定制的复合奖励函数进行强化学习,在保持响应质量的同时最大化目标叙事维度的多样性。我们对两个LLM家族的实证结果显示,DivLM相比其他方法平均提升多样性指标超过9%,同时保持指令跟随、整体响应质量和与人类输出的相似性。
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
CLIFT:Web代理培训和测试时间缩放的共形自我验证
- Authors: Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu, Ran Xu, Zeyuan Chen
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06829
- Pdf link: https://arxiv.org/pdf/2610.06829
- Abstract
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
- 中文摘要
开源网络代理现在足够强大,能够执行真实的浏览器任务,但用强化学习训练它们仍依赖于薄弱的监督:二元任务的成功率过于稀疏,无法分配学分,而前沿语言模型评判者在每个步骤调用成本过高,且不能假设部署时已可用。我们引入CLIFT,这是一种围绕共形自我验证构建的训练和测试时间缩放方法。在训练过程中,代理回答关于自身部署的自然语言验证问题;组合共形认证者只保留与训练时间评判一致的URL条件证据的问题信号,通过极性感知提升分配签名信任权重,并将验证者得分与每步奖励合并,且不减去评审基线。测试时,同一认证银行被冻结并作为共形轨迹选择(CTS)的结构化证据再利用:代理采集贪婪的推广和一次或多次多样化的重复尝试,自我验证者总结每个URL追踪,保守多数票规则决定是否从现有机构切换,无需外部裁判。这一单一机制支持三种设置。在WebArena Infinity上,CLIFT在开源网络代理中达到了最先进的性能。在VisualWebArena上,使用开放模型训练的银行在测试时转移到GPT-5.5,并在规范框架下达到最先进性能。在在线Mind2Web上,无需对基准进行培训,通过翻译认证题库即可提升在线网络代理的零样本评估能力。这些结果共同使得共形自验证成为将昂贵的评判反馈转化为可重复使用的训练信号和无评判测试时间的扩展信号。
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
MemPilot:为LLM代理编排按需多模态内存管理
- Authors: Haozhen Zhang, Haodong Yue, Quanyu Long, Jianzhu Bao, Qingyuan Liu, Tao Feng, Bohan Liu, Weida Liang, Wenya Wang
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06830
- Pdf link: https://arxiv.org/pdf/2610.06830
- Abstract
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.
- 中文摘要
内存已成为LLM代理生态系统的核心,支持信息的保存和跨交互重用。然而,大多数现有代理内存系统采用与查询无关的方式构建内存,这可能导致不必要的预处理成本,并丢弃后来证明至关重要的细节。近期研究开始将内存处理转向运行时适配,但通常专注于特定操作或固定处理方案,导致对性能、成本和延迟的灵活控制尚未被充分探索。为应对这一挑战,我们提出了\textbf{MemPilot},一个灵活框架,可在不同性能-成本-延迟偏好下协调按需内存管理。具体来说,我们通过强化学习优化多步LLM策略,迭代选择从查询无关内存检索,或将多模态历史的查询特定管理委托给异构LLM和VLM之间。该策略联合控制证据数量、策划指令、模型选择和可视化访问,实现运行时计算的细粒度分配。为优化该政策,我们通过分别估计每个目标的优势后再汇总,调整了目标层面优势解耦。此外,我们引入了基于前缀的边际效用估计,用于多步推广的细粒度信用分配。五个多模态代理-内存基准测试的实验显示,优化偏好在性能-成本-延迟方面存在有利权衡,偏好扫除带来的前沿比现有权衡意识基线更广。
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
迈向正确完成的循环模型,第二部分:固定点的重新思考
- Authors: Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.06833
- Pdf link: https://arxiv.org/pdf/2610.06833
- Abstract
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.
- 中文摘要
每次循环语言模型的重复都会增加训练、解码、预填充和强化学习(RL)的成本。递归状态越接近固定点,路径的重要性就越小。这使得训练中的截断反向传播成为可能;几乎不损失准确度的终端键值(KV)解码;预填充速度高达1.79倍的精简学生;以及从保存的滚动状态计算梯度的RL更新,速度是回放轨迹反向传播速度的2倍。因此,我们改进了塑造这些固定点的训练两个组成部分:深度先验和输入注入。固定深度训练破坏了KV共享,而Huginn的广义深度先验支持共享,但比共享所需的更稀释目标深度的监督;我们通过预测反馈学习先验,并有一个熵项使其保持宽广。现有的注入方案允许状态沿输入的分量放大或抵消注入;我们用正交注入去除该分量。从100M到1.6B参数,学习到的先验和正交注入在各尺度上相较于Huginn的先验和现有注入方案,分别降低了混淆度。在1.6B时,学习先验且KV缓存小3倍,与固定深度训练与全缓存的下游平均值相匹配。
Base Models Can Reason By Taking a Cue From Training Data
基础模型可以通过从训练数据中获取提示来推理
- Authors: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.06851
- Pdf link: https://arxiv.org/pdf/2610.06851
- Abstract
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
- 中文摘要
本文研究训练数据如何将基础模型响应开头的标记与随后的推理行为之间建立关联。首先,我们证明修正特定的起始标记提示使基础模型在数学和编码方面的性能能够与其强化学习(RL)训练的对应模型竞争。例如,提示“.\n\nOkay”将Olmo-3-7B的MATH-500 pass@1准确率从42%提升到78%,而“Alright”则将Qwen3-14B的准确率从72%提升到87%。其次,强化学习使这些提示更有可能出现,而修正它们则能恢复其相较基础模型的大部分性能提升。第三,我们将标记提示的推理效应追溯到训练数据。我们进行因果数据干预,将任意词语(如“鸡”)转化为有效的推理线索,或消除已有线索的效果。类似的编辑使提示指令“Think duck duck goose”与“Think by Step Step Think”在引发推理方面同样有效。我们还发现不同线索诱导的隐藏状态表示与训练集中的不同文档类型相关。最后,我们通过语言模型安全性的案例研究扩展了代币线索的研究,发现不同线索引发了对应不同训练数据类型的拒绝和顺从行为。
Keyword: diffusion policy
The Unexpired Plan: A Free Monitor for Accelerated Diffusion Policies
未过期计划:加速扩散政策的免费监测
- Authors: Yi Zhao, Sebastian Scherer
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.05747
- Pdf link: https://arxiv.org/pdf/2610.05747
- Abstract
Training-free acceleration of a diffusion policy is accepted when an internal similarity signal reports that the shortcut changed nothing. We price every monitor a control loop can afford in closed-loop success rather than feature distance, over $114$ accelerator configurations and four policy families: two accelerators each clear their own gate's bar and log every reuse as certified, while on one task one finishes every episode and the other none. A monitor is two designs, not one --- the statistic it reads, and what it does when that statistic fires. Published gates re-arm after every rejection, and under that response even an oracle handed every forward pass and the exact local action error loses twenty points; absorbing the same statistic costs one point, and most of its speed. What makes a response that never forgets affordable is a statistic that rarely fires, and what it must measure is deviation from the policy the accelerator replaced --- which a chunked policy has already paid for, its last plan not yet expired and free to read. Guarding every call this way cuts per-call compute by $1.55$--$3.09\times$, where any reference-requiring check at the same coverage would have to stop accelerating altogether. On three of our four families the schedule alone already holds the pre-stated $\pm2$-point margin. What the monitor is measurably worth shows in three places: on the fourth family, where it rescues the candidate selection landed on; on four configurations it did not select; and on a contact-rich fifth family, chosen where the schedule was expected to fail and run after every design choice was frozen, where no unmonitored arm at its speed holds the margin and the monitored one does.
- 中文摘要
当内部相似性信号报告捷径没有改变时,则接受无训练的扩散策略加速。我们对控制环能承受的每个监控器定价为闭环成功而非特征距离,超过114美元加速器配置和四类策略:两个加速器各自清除自己的门条,并将每次重用记录为认证,而在一个任务中一个完成所有集数,另一个则不完成。一个监视器由两个设计组成,而非其读取统计数据及其触发时的反应---一个。发布的网关在每次拒绝后重新激活,在该响应下,即使是预言机处理了每一次前向传递和准确的局部动作错误,也会损失20分;吸收相同的统计数据需要1分,且大部分速度会下降。使得永不忘记的响应经济实惠的是很少触发的统计数据,而它必须衡量的是加速器已付费、上一个计划尚未到期且可自由阅读的政策---偏离。以这种方式保护每通电话可减少每次通话计算1.55美元——3.09美元乘以加倍,而同一覆盖范围的推荐查询则必须完全停止加速。在我们四个家族中有三个,仅该计划就已保持预先设定的$\pm2点余裕。监视器的可测量价值在三个地方显示:第四个家族,它救援候选选择的选项;四个未选择配置;以及接触较多的第五个家族,计划预期失败且在所有设计选择冻结后运行,且没有未监控的机械臂保持其速度的余裕,而被监控的机械臂则保持了边际。
Demonstration-Calibrated Port-Hamiltonian Retuning for Manipulation Policies
演示校准的Port-Hamiltonian操作策略重调
- Authors: Yulong Yang, Fan Wu, Christine Allen-Blanchette, Amit Chakraborty
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.05755
- Pdf link: https://arxiv.org/pdf/2610.05755
- Abstract
Diffusion and VLA policies for manipulation are often deployed through downstream impedance controllers. The stiffness and damping gains of these controllers affect task success, yet are commonly inherited from data collection rather than selected for the deployed policy. Although empirical gain sweeps can improve performance, they require repeated evaluation rollouts. We introduce PHRetune, an offline method that derives controller gains for a frozen policy without evaluation rollouts or gain search. Our approach learns a port-Hamiltonian model from demonstrations to estimate the effort and energy associated with the policy's predicted actions. The policy is applied to recorded demonstration observations, and its predictions are assessed against demonstration-derived effort and energy budgets. From this comparison, we derive a single gain scale in closed form, adjusting the downstream controller while preserving the policy and its action representation. The gains are fixed before evaluation, without requiring a prior manipulator model, task rewards, policy retraining, or additional runtime computation. Across LIBERO suites, PHRetune improves Diffusion Policy success by up to 9.4 percentage points, with the derived gains achieving the highest observed success rates in empirical gain sweeps. On all four real-world manipulation tasks, PHRetuned Diffusion Policy outperforms the nominal policy, alternative gain-tuning methods, and a policy-retraining baseline. The same procedure improves success with SmolVLA and OpenVLA-OFT on every task, while reducing acceleration and jerk for both VLA backbones.
- 中文摘要
扩散和VLA策略的操作通常通过下游阻抗控制器部署。这些控制器的刚性和阻尼增益影响任务成功,但通常是从数据收集中继承而来,而非部署策略中选择的。尽管实证增益扫频可以提升性能,但需要反复的评估推广。我们引入PHRetune,这是一种离线方法,可在冻结策略中推导控制器增益,无需评估部署或增益搜索。我们的方法通过演示学习端口哈密顿模型,估算策略预测动作所需的工作量和能量。该策略应用于录制的演示观测,其预测与演示导出的努力和能量预算进行评估。通过该比较,我们推导出一个封闭形式的单一增益尺度,调整下游控制器,同时保持策略及其动作表示。收益在评估前固定,无需事先操作员模型、任务奖励、策略重训练或额外的运行时计算。在LIBERO套件中,PHRetune可将扩散策略成功率提升高达9.4个百分点,衍生收益在经验增益扫荡中实现最高成功率。在所有四种真实操作任务中,PHRetuned扩散策略的表现优于名义策略、替代增益调整方法和策略重训基线。同一过程提升了SmolVLA和OpenVLA-OFT在所有任务中的成功率,同时减少了两个VLA骨干的加速和抖动。
Reachability-Aware Diffusion Policy Optimization
可达性感知扩散策略优化
- Authors: Hikmet Simsir, Kutay Demiray, Ozgur S. Oguz
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.05969
- Pdf link: https://arxiv.org/pdf/2610.05969
- Abstract
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
- 中文摘要
扩散策略为连续控制强化学习提供了表达式动作分布。然而,安全意识的在线扩散策略优化仍未被充分探索,尤其是使用预测性可达性信息而无显式动力学模型的方法。我们提出了可达性感知扩散策略优化(RADPO),这是一种无模型方法,结合了预测性首次命中安全估计与累积成本预算反馈。RADPO学习一个折现的首次命中可达性值,捕捉成本事件的折现风险,赋予较早发生事件更大权重,并利用该信号塑造奖励。一个独立的类似对偶乘数根据实现的情节成本相对于规定预算调整塑形强度。扩散演员通过对奖励批评者评分的候选动作进行加权去噪回归来提升。我们的方法既不需要学习动力学模型,也不需要通过批评者进行动作梯度,也不需要通过反向扩散采样器进行微分。我们建立了可达性值的理论性质,并展示了累积可达性惩罚为未来折现累计成本提供了保守的替代。在十个连续控制安全任务中,RADPO实现了具有竞争力的回报-成本权衡,相较于对比基线,在多个任务上显著减少了约束违规。我们的理论和实证分析支持将可达性与累积预算反馈结合起来,是安全意识扩散政策的可行方法。
Robotizing Human Videos with Physically Consistent Interactions
机器人化具有物理一致性交互的人类视频
- Authors: Ching-Lam Cheng, Shengfeng He, Bin Zhu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.06137
- Pdf link: https://arxiv.org/pdf/2610.06137
- Abstract
Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.
- 中文摘要
人类视频提供了可扩展的操作数据,但人手与机器人操作器之间的身体差距限制了其直接使用。现有的视频编辑方法用渲染机器人替代手部,但不准确的交互重建和合成可能导致抓取不一致和机器人-物体遮挡不合理。我们从两个互补的物理方面解决这些失败:交互几何和场景可见性。首先,交互感知接触重建模块结合手部物体分割与网格级接触预测,恢复密集的三维接触,然后将其转换为平行颚抓握器的时间稳定抓取。其次,深度感知合成模块利用场景和机器人深度强制执行物理一致的机器人-物体遮挡。最终视频以机器人兼容的形式保留了人类演示的交互结构,并与机器人演示共同训练。利用相同的人类视频和机器人数据,我们与仅机器人训练及原始 Masquerade 流水线进行比较。在四个机器人双生任务和两个扩散策略视觉编码器中,我们的方法实现了最高的平均成功率,尤其是在分布外场景变化下提升显著。实际部署进一步表明,当任务几何可观察时,所提协同训练方法提升了对视觉干扰的鲁棒性,而深度敏感抓握的性能仍受限于单摄像头配置。