生成时间: 2026-09-29 00:32:17 (UTC+8); Arxiv 发布时间: 2026-09-28 20:00 EDT (2026-09-29 08:00 UTC+8)
今天共有 32 篇相关文章
Keyword: reinforcement learning
From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
从弱数据到强策略:Q-目标实现可证实的上下文强化学习
- Authors: Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.30391
- Pdf link: https://arxiv.org/pdf/2609.30391
- Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.
- 中文摘要
现有的上下文强化学习方法主要通过监督行为预测目标预训练Transformers。这使任务能够从上下文推断,但学习策略高度依赖离线动作的质量:当轨迹弱或次优时,模仿本身就变成有偏的学习信号。我们提出了Q目标预训练变换器(QTPT),它保留了上下文条件的Transformer架构,但用Bellman式Q目标目标替代行为克隆。因此,QTPT学习在上下文中使用奖励和转换来估计动作值,而不仅仅是模仿行为策略。我们在随机线性盗垒和有限视野MDP中理论分析QTPT,显示出比监督预训练更强的数据质量鲁棒性。在实证上,QTPT在受控RL基准测试中表现优于监督行为预测,且数据为随机或次优,我们考察了D4RL Kitchen和AntMaze的扩展。补充实验评估了骨干鲁棒性、元强化学习比较、任务一致性上下文以及无支持动作价值高估。这些比较区分了Q目标预训练的优势与离线覆盖的剩余限制。
Privacy-Preserving Prompted Policy Search for Robotic Control
保护隐私的提示策略 机器人控制搜索
- Authors: Ali Irshayyid, Feng Lin, Chong Li, Jun Chen
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.30554
- Pdf link: https://arxiv.org/pdf/2609.30554
- Abstract
Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy parameters and rewards history to cloud-based LLM APIs, exposing proprietary control strategies to third-party service providers. To address this issue, this paper introduces Privacy-Preserving Prompted Policy Search (PP-ProPS), a framework that enables LLM-guided policy optimization while keeping policy and environmental parameters confidential. PP-ProPS encodes policy parameters and reward values using secret client-side transformations before they are included in each API request, ensuring that the LLM provider observes only encoded policy parameters and scaled reward information. Furthermore, unlike Vanilla ProPS, the proposed framework does not require the true optimal episodic return to be known or disclosed to the LLM. Beyond protecting the optimization data, PP-ProPS improves the search process in two ways. First, it provides the LLM with individual reward components instead of only a single total return, offering more informative feedback about each candidate policy. Second, it uses a bounded history that prevents the prompt from growing indefinitely, improving search with high-dimensional policies and supporting the use of open-weight LLMs. The proposed PP-ProPS is evaluated on both continuous and discrete control problems spanning Multi-Joint dynamics with Contact (MuJoCo) locomotion, classic control, highway driving, and robotic arm manipulation. Compared to Vanilla ProPS, the proposed PP-ProPS outperforms ProPS in seven of the ten evaluated tasks, and surpasses conventional RL methods including PPO, SAC, and TRPO, in five of the six tasks.
- 中文摘要
大型语言模型(LLM)最近展示了作为强化学习(RL)上下文策略优化器的有前景能力,使策略搜索能够同时由数值奖励信号和自然语言推理驱动。然而,实际部署此类方法需要将原始策略参数和奖励历史传输到基于云的LLM API,从而向第三方服务提供商暴露专有控制策略。为解决这一问题,本文介绍了隐私保护提示策略搜索(PP-ProPS),该框架在保持政策和环境参数机密的前提下实现LLM引导的策略优化。PP-ProPS通过秘密客户端转换编码策略参数和奖励值,确保LLM提供者仅观察编码的策略参数和可扩展的奖励信息。此外,与原版ProPS不同,所提框架不要求将真正的最优情节回报信息告知或披露给LLM。除了保护优化数据外,PP-ProPS还通过两个方面改进了搜索过程。首先,它为LLM提供单独的奖励组成部分,而不仅仅是单一的总回报,从而对每个候选策略提供更多信息。其次,它采用有界历史,防止提示无限增长,通过高维策略提升搜索效果,并支持开放权重LLM的使用。所提PP-ProPS在连续和离散控制问题上评估,涵盖多关节动力学、接触(MuJoCo)运动、经典控制、高速公路驾驶和机械臂操作。与原版ProPS相比,PP-ProPS在十个评估任务中有七个表现优于ProPS,并在六个任务中有五个超过了包括PPO、SAC和TRPO在内的传统强化学习方法。
SoGuDiff: Socially Guided Diffusion for Steerable, Norm-Grounded Robot Navigation
SoGuDiff:社会引导扩散,用于可转向、规范基础的机器人导航
- Authors: Christian Schaible, Haoran Ji, Yash Vardhan Pant, Stephen L. Smith
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30560
- Pdf link: https://arxiv.org/pdf/2609.30560
- Abstract
Beyond collision avoidance, socially competent robot navigation requires adherence to implicit social conventions that vary across contexts, cultures, and deployment requirements. Many conventional navigation policies learn a single normative behavior, either through reinforcement learning against a fixed reward function or imitation of human demonstrations, exposing no interface for adjusting that conduct at runtime. We present a diffusion-based navigation framework whose social behavior can be tuned at deployment: a desired style is specified, such as how closely the robot passes, which side it yields to, or how much it defers to groups, and the planner adapts accordingly. Continuous style axes can be followed independently or composed, spanning a behavioral space rather than discrete, primitive-based specifications. A feasibility projection layer separates learned social behavior from kinematic feasibility and collision avoidance. A single-axis sweep illustrates a tradeoff curve that strictly dominates the evaluated fixed-behavior baseline configurations, and stylistic differences are replicated in real-world demonstrations.
- 中文摘要
除了碰撞避免,具备社会能力的机器人导航还要求遵守隐含的社会惯例,这些惯例会因上下文、文化和部署需求而异。许多传统导航策略学习单一规范行为,要么通过对固定奖励函数的强化学习,要么模仿人类演示,不暴露用于在运行时调整该行为的接口。我们提出了一个基于扩散的导航框架,其社会行为可在部署时调整:指定期望的风格,如机器人通过的距离、向哪一侧让路或向群体让路的程度,规划者则相应调整。连续风格轴可以独立或组合,跨越行为空间,而非离散的基于原始的规范。可行性预测层将学习的社会行为与运动学可行性和碰撞避免区分开来。单轴扫描展示了一条权衡曲线,严格支配所评估的固定行为基线配置,风格差异在实际演示中得以复制。
MOCHA: Multi-Objective Co-Design using Hypernetwork Architectures
MOCHA:利用超网络架构进行多目标协同设计
- Authors: Varun Madabushi, Neil Janwani, Maegan Tucker
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30570
- Pdf link: https://arxiv.org/pdf/2609.30570
- Abstract
In this work, we present MOCHA, the first, to our knowledge, reinforcement learning based approach to computing a family of Pareto-optimal policies across the design space of a robot using a single network. Specifically, MOCHA leverages the hypernetwork architecture to learn a network that produces specialized network parameters optimized for a given objective and parameterized robot design; we term this a multi-objective design hypernetwork (MDH). We demonstrate the capabilities of MDHs to represent a complex family of design-dependent strategies on two distinct robot morphologies, each with six design dimensions and across 2-3 objectives. Moreover, we propose an approach for efficiently producing a Design Pareto set using evolutionary search of the learned policy network, generating the optimal design-policy combination for each objective prioritization. Lastly, we provide an efficient method for computing generalist robot designs which achieve the best cumulative performance across the entire set of objectives.
- 中文摘要
在本研究中,我们提出了MOCHA,这是据我们所知,基于强化学习的首个方法,用于利用单一网络计算机器人设计空间中一系列帕累托最优策略。具体来说,MOCHA利用超网络架构学习一个能够生成针对特定目标和参数化机器人设计优化的专用网络参数的网络;我们称之为多目标设计超网络(MDH)。我们展示了MDH在两种不同机器人形态上表示复杂设计依赖策略的能力,每个机器人形态具有六个设计维度,跨越2-3个目标。此外,我们还提出了一种通过进化搜索所学策略网络高效生成设计帕累托集的方法,生成每个目标优先级的最优设计-策略组合。最后,我们提供了一种高效的方法,用于计算通用机器人设计,使其在整个目标集合中实现最佳累积性能。
HuGo: LLMs as Whole-Body Policy Code Designers for Humanoid Loco-Manipulation
HuGo:大型语言模型作为人形机车操控的整体政策代码设计者
- Authors: Seoyeon Choi, Shizhao Ye, Nicholas Bui, Aayushi Shrivastava, Kanghyun Ryu, Dhruva Tirumala, Markus Wulfmeier, Negar Mehr
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30594
- Pdf link: https://arxiv.org/pdf/2609.30594
- Abstract
For humanoids to be useful in everyday environments, they must perform a wide range of tasks that couple locomotion and manipulation. Existing approaches commonly acquire a loco-manipulation policy through reward engineering or demonstrations followed by task-specific training, making it costly to scale to new tasks. In this work, we propose a hierarchical approach to humanoid loco-manipulation that eliminates these per-task requirements. HuGo, Humanoid policy code Generation, uses a Large Language Model (LLM) to generate executable, closed-loop high-level policy code from a task description on top of a frozen low-level whole-body policy. Given the task, observation, and command specifications, the LLM constructs the task logic in code. HuGo then refines the policy from its rollouts using numerical trajectories and selected video frames to produce feedback and targeted code updates. Across five simulation tasks, using two different low-level policies, HuGo substantially outperforms a high-level reinforcement learning baseline and approaches the performance of a demonstration-based baseline. We achieve this level of performance without task-specific reward design or demonstration collection. We further demonstrate zero-shot transfer of simulation-generated policies to hardware and show that applying the same refinement loop to real-world rollouts can further improve transfer performance without expert demonstrations or policy retraining. Project website is this https URL
- 中文摘要
为了使类人机器人在日常环境中有用,必须执行将移动与操作结合的广泛任务。现有方法通常通过奖励工程或演示,随后进行任务特定培训,获得机车操作策略,这使得扩展新任务的成本较高。在本研究中,我们提出了一种层级式的人形机车操作方法,消除了这些每个任务的要求。HuGo,人形策略代码生成,利用大型语言模型(LLM)从任务描述生成可执行的闭环高级策略代码,基于冻结的低层级整体策略。LLM基于任务、观察和命令规格,构建代码中的任务逻辑。HuGo随后通过数值轨迹和选定视频帧从部署中细化策略,生成反馈和有针对性的代码更新。在五个模拟任务中,使用两种不同的低层策略,HuGo的表现显著优于高级强化学习基线,并接近基于演示的基线的性能。我们无需任务特定奖励设计或演示收集即可实现这一性能水平。我们还进一步演示了模拟生成策略的零样本转移到硬件,并展示了将相同的精炼循环应用于真实世界推广,可以进一步提升传输性能,无需专家演示或策略重新训练。项目网站为 https URL
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
基于深度强化学习的入侵检测系统,具有可解释性的概率性鲁棒性驱动的普遍对抗扰动
- Authors: Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou, Carsten Maple
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.30605
- Pdf link: https://arxiv.org/pdf/2609.30605
- Abstract
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the universality objective of UAPs, i.e., PR quantifies the prevalence of misclassification in the input space, making it a natural signal for guiding UAP generation. Hence, we propose PR-based UAP, which represents the first integration of an explicit PR-driven objective into generating UAPs against DRL-based IDS. Building on this formulation, we introduce PX-UAP, which leverages explainable artificial intelligence (XAI) to guide perturbation shaping under realistic domain constraints, and provides a rigorous theoretical analysis of its design. Extensive experiments demonstrate that PX-UAP consistently outperforms state-of-the-art UAP methods in attack effectiveness.
- 中文摘要
深度强化学习(DRL)使动态网络环境中的自适应入侵检测成为可能,但也使入侵检测系统(IDS)暴露于对抗性威胁,如通用对抗扰动(UAP),这些扰动应用单一输入无关扰动以降低流量检测性能。概率鲁棒性(PR)作为事后评估指标,提供了一种原则性的、群体级的对抗影响衡量指标,概念上与UAP的普遍性目标相符,即PR量化输入空间中错误分类的普遍性,使其成为引导UAP生成的自然信号。因此,我们提出基于PR的UAP,这是首次将显式PR驱动目标整合到针对基于DRL的IDS生成UAP的案例。基于这一表述,我们介绍了PX-UAP,利用可解释人工智能(XAI)在现实的域约束下引导微扰整形,并对其设计进行了严谨的理论分析。大量实验表明,PX-UAP在攻击效能上始终优于最先进的UAP方法。
MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization
MVAgent:通过一致的条件构建和镜头级策略优化实现多代理视频生成
- Authors: Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun, Chenyang Lyu, Longyue Wang, Weihua Luo
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.30609
- Pdf link: https://arxiv.org/pdf/2609.30609
- Abstract
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoint, a Spatial Grounding agent samples views from generated camera-traversal clips and anchors each shot to the view matching its framing. As generated shots drift from the plan, an Observer records how each shot ends in a continuity memory, from which a Transition agent builds character action and spatial references for the next shot. An Orchestrator composes these inputs into each request. Since a request reveals its effect only after rendering, we train it by agentic reinforcement learning with Trunk-GDPO, which compares rendered candidates at every shot rather than once per video and continues the best as the trunk. With generator and judges frozen, MVAgent attains the highest cross-shot consistency and narrative-planning quality among the compared methods on ViMax-Bench and is preferred over the strongest agentic baseline in human evaluation.
- 中文摘要
多镜头代理视频生成需要角色外观一致、不同摄像机角度的空间布局稳定,以及镜头间的连续角色状态。当每个镜头都是对冻结生成器的独立请求时,重复文本不会决定外观、布局或状态。因此,我们将问题重新构造为条件构造,并呈现MVAgent,一个多代理管道,代理通过类型化的条件输入协作。由于环境图像显示一个视角,空间基础代理从生成的摄像机穿越片段中采样视角,并将每个镜头锚定到与其构图相符的视角。随着生成镜头偏离计划,观察者会在连续性记忆中记录每个镜头的结尾,过渡代理从中构建角色动作和空间参考,为下一镜头构建。编排器将这些输入组合到每个请求中。由于请求只有在渲染后才显现效果,我们通过Trunk-GDPO进行智能强化学习训练,该方法在每个镜头中比较渲染候选对象,而非每个视频一次,并且作为主干持续进行最佳表现。在生成器和评委冻结后,MVAgent在ViMax-Bench上的所有比较方法中实现了最高的交叉镜头一致性和叙事规划质量,并且在人类评估中优于最强的代理基线。
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
OpenHail:一个以活动为驱动的电动网约车车队控制体育馆环境
- Authors: Tommaso Schettini, Nicholas D. Kullman, Jorge E. Mendoza
- Subjects: Subjects:
Machine Learning (cs.LG); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2609.30628
- Pdf link: https://arxiv.org/pdf/2609.30628
- Abstract
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs for training and evaluation. For electric fleets, this environment must also capture the interaction among stochastic demand, vehicle operations, and capacitated charging infrastructure. We present OpenHail, an open-source Gymnasium environment for joint control of electric ride-hailing fleets. Its fixed-size observation--action interface exposes request assignment, repositioning, and charging to a single policy. The event-driven simulator represents requests with pickup deadlines, vehicle job queues, battery dynamics, and finite-capacity charging facilities with first-in--first-out queues. A configurable decision-epoch mechanism separates internal simulator events from policy interactions, supporting event-driven, periodic, hybrid, and policy-requested control within the same operational model. The software provides seeded instances, feasible-action utilities, evaluation tools, operational metrics, and baseline policies. The source code is available at this https URL.
- 中文摘要
近年来,机器学习政策对网约车车队控制的关注日益增加。尤其是强化学习,需要结构化的仿真环境,明确观察、动作、奖励和决策阶段以供训练和评估。对于电动车队,该环境还必须捕捉随机需求、车辆运营和电容充电基础设施之间的相互作用。我们介绍OpenHail,一个开源的Gymnasium环境,用于联合控制电动网约车车队。其固定规模的观察-动作界面将请求分配、重新定位和充电暴露为单一策略。事件驱动模拟器表示带有取车截止日期、车辆作业队列、电池动态和有限容量充电设施的请求,采用先进先出队列。可配置的决策时代机制将内部模拟器事件与策略交互分离,支持同一运营模型内的事件驱动、周期性、混合和策略请求控制。该软件提供种子实例、可行操作工具、评估工具、运营指标和基线策略。源代码可在该 https URL 获取。
Learning Vision-Based Agile Gap Traversal: Differentiable Simulation with a Warm-Started Critic
基于愿景的敏捷差距穿越学习:带着热启批评者的可微分仿真
- Authors: Nuthasith Gerdpratoom, Tianchen Sun, Yichao Gao, Lin Zhao
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.30696
- Pdf link: https://arxiv.org/pdf/2609.30696
- Abstract
Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.
- 中文摘要
对于自主四旋翼机来说,穿越狭窄间隙具有挑战性,尤其是当控制指令直接来自高维视觉观察时。现有端到端方法通常依赖行为克隆或通过可微分仿真进行全滚动反向传播(BPTT),这可能限制策略性能或产生高额训练成本。我们提出了一个两阶段强化学习框架,用于更高效的以自我为中心的视觉运动间隙穿越策略训练,利用准分析策略梯度(QPG)通过可微仿真和批评热启动。该框架利用QPG通过视觉渲染避免反向传播,降低计算和内存成本,同时提升样本效率。第一阶段,专家行为者和批评者通过特权观察(包括间隙几何)进行训练。与以往的间隙穿越方法不同,我们利用QPG的训练无需沿优化的参考轨迹重置代理。第二阶段,视觉策略使用来自两个以自我为中心的摄像头和低维观测的二元间隙掩码进行训练,而其特权批评者则从第一阶段开始热启动。这相比于冷启动批评者或全展开BPTT,显著提升了训练效率和穿越成功率。我们的框架在系统参数变化时无需重新训练专家行为者,这使得跨无人机平台的泛化比基于动作监督的先进视觉间隙穿越方法更高效。所学的视觉策略还可推广到具有未见形状的间隙。大量真实实验进一步展示了使用在线渲染的二元遮罩的稳健间隙遍历。除了间隙遍历外,所提出的框架是通用的,且可扩展到其他视觉运动机器人学习任务。
MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning
MM-VeriAgent:学习使用大量工具通过强化学习验证多模态错误信息
- Authors: Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2609.30698
- Pdf link: https://arxiv.org/pdf/2609.30698
- Abstract
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training this http URL on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.
- 中文摘要
现实世界的多模态错误信息常涉及混合伪造源,需要针对样本的检测策略。现有的工具增强方法依赖预定义的工作流程或推理时间规划,限制了适应性或增加推理成本。为解决这个问题,我们引入了 \textbf{MM-VeriAgent},它学习如何用工具验证混合源多模态错误信息。我们首先构建了 \textbf{MM-VeriTools},一个专门用于虚假信息检测代理的工具包。通过对混合源检测所需的子任务进行各种候选模型和方法的基准测试,我们挑选出最强的文本、可视化和跨模态伪造分析工具,并将它们封装为可调用工具,并统一接口。在该工具包之上,我们通过强化学习训练 LVLM 代理,教它如何利用这些工具更好地解决混合源检测问题。由于许多工具是专门模型,每次推出时的在线执行严重限制了强化学习效率,我们进一步引入了 \textbf{Tool-Execution Cache},它预先执行候选工具调用,并在训练时重用其缓存输出。这保留了多步部署,同时减少了在线工具执行,显著提升了训练。MMFakeBench 上的 http URL 在推理时无需显式工具搜索,就在基础模型上展示了显著的准确性提升。消融和效率分析进一步验证了所学工具使用策略,并显示工具执行缓存在训练过程中减少了在线工具执行。
WeEnv: The Environment for Agentic Reinforcement Learning at WeChat
WeEnv:微信中的能动强化学习环境
- Authors: Yang Yu, Jing Lei, Shaoxun Zeng, Xinyu Gao, Jindi Shi, Ci Lei, Junjie Zhang
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2609.30766
- Pdf link: https://arxiv.org/pdf/2609.30766
- Abstract
Agentic reinforcement learning (RL) differs from conventional RL in that every task executes inside a complex environment, e.g., a virtual machine or a container. We find that agentic RL pays a heavy environment tax: a large share of the iteration time goes to the environment rather than to learning. The root cause is the lack of a full-lifecycle solution to environment management. We present WeEnv, which manages environments across packaging, initialization, and provisioning. WeEnv packages components as independently published layer groups and composes them at initialization, so that updating a component republishes one small group rather than every artifact containing it. To speed up environment initialization, WeEnv launches environments instantly and fetches contents on demand. During task execution, WeEnv provisions CPU and memory elastically, adjusting each environment's quota from its observed usage to fit the varying demands. WeEnv reduces the initialization by 5.6-14.2x over E2B, Docker, and AgentENV, cutting its share of the iteration time from up to 53.4% to 9.1%. WeEnv is deployed for agentic RL at WeChat.
- 中文摘要
代理强化学习(RL)不同于传统强化学习,每个任务都在复杂环境中执行,例如虚拟机或容器。我们发现代理强化学习承担了沉重的环境税:大量迭代时间用于环境,而非学习。根本原因是缺乏完整的环境管理生命周期解决方案。我们介绍WeEnv,它管理环境打包、初始化和配置。WeEnv将组件打包为独立发布的层组,并在初始化时组合,因此更新组件时会重新发布一个小组,而非所有包含该组的工件。为加快环境初始化,WeEnv即时启动环境并按需获取内容。在任务执行过程中,WeEnv会弹性地配置CPU和内存,根据观察到的使用情况调整每个环境的配额以满足不同需求。WeEnv 相比 E2B、Docker 和 AgentENV 将初始化时间缩短了 5.6-14.2 倍,将其在迭代中的份额从高达 53.4% 缩短到 9.1%。WeEnv 已部署为微信的代理强化学习平台。
Learning Chance-Constrained MDPs with Bellman Distributional Certificates
学习机会受限的MDP与贝尔曼分配证书
- Authors: Chenbei Lu, Hongyu Yi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.30856
- Pdf link: https://arxiv.org/pdf/2609.30856
- Abstract
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
- 中文摘要
安全强化学习(RL)通常强制执行期望成本约束,但这种期望安全可能无法控制罕见的高成本轨迹的概率。机会约束MDP(CCMDPs)对概率层级要求更强,但由于偶然约束非凸且依赖于整个轨迹而非Bellman线性期望,因此被广泛认为更难。本文揭示,这种计算难度并不一定意味着更高的统计价格。对于表式折现CCMDPs,支持固定有界继继者并可访问认证的规划预言机,我们建立了基于模型的上界,下界可匹配至对数项。技术上,我们的核心思想是\emph{Bellman分布证书},它在策略选择前构造约束违规概率的Bellman递归。该证书可在候选策略间重复使用;结合共享的逐行反KL置信集,它能提供策略均匀轨迹-KL转移,无需对策略或时间预算Bellman表进行并集约束。对于随机策略,我们给出一个模型无方差约简策略梯度算法,具有有限样本的期望KKT剩余保证,并且对每个被接受策略进行独立验证。合成CCMDP和IEEE 14总线能量存储控制基准测试的数值实验展示了所提算法的安全性和机制行为。
VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL
VLaRL:用仿真训练的潜条件残留强化学习增强视觉-语言-行动模型
- Authors: Namiko Saito, Kinam Kim, Heecheol Kim, Katsushi Ikeuchi, Yasuyuki Matsushita
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30868
- Pdf link: https://arxiv.org/pdf/2609.30868
- Abstract
Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.
- 中文摘要
视觉-语言-动作(VLA)模型提供广泛的指令条件操作行为,但在接触丰富交互中,其物理执行可能不精确。残余强化学习(RL)可以在保持VLA冻结的同时纠正此类错误,但真实机器人RL成本高昂且安全性至关重要。我们提出了VLA潜在条件RL(VLaRL),使冻结VLA的残余RL能够在模拟中训练并部署到真实机器人上,无需真实现实的RL或在线适配。关键挑战是尽管模拟与现实之间存在视觉差距,仍能转移已学到的残余策略。VLaRL不要求像素级视觉对应,而是利用VLA内部的视觉语言潜在表征来条件残差控制,并作为模拟到现实的传输接口,并学习一种轻量级映射器,将仿真衍生的潜在映射映射器转化为真实潜在分布。在四种接触丰富操作任务和两条VLA骨干链中,VLaRL提升了所有任务-骨干组合的实际成功率,而受控消融则展示了潜条件和潜在比对在传递模拟训练残差控制中的重要性。
QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking
QReason:以查询为中心的解耦思维链,促进文书重新排序
- Authors: Yang Zhang, Wenhan Liu, Qiannan Zhu, Mingming Li, Yuanfei Huang
- Subjects: Subjects:
Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2609.30904
- Pdf link: https://arxiv.org/pdf/2609.30904
- Abstract
Passage reranking plays a crucial role in information retrieval by refining the ordering of candidate passages to better reflect relevance. Existing listwise LLM rerankers with Chain-of-Thought (CoT) reasoning can handle complex queries effectively, but they suffer from substantial redundancy and high latency due to sliding-window strategies, which repeatedly generate highly similar CoTs. To address this, we propose QReason, a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment. Specifically, QReason introduces a dedicated rewriter that generates a ranking-oriented reasoning query once, capturing the query's core intent while avoiding redundant reasoning, and then reuses it across all windows with a non-reasoning reranker. The rewriter is trained via a two-stage process that first uses supervised fine-tuning with relevant-passage guidance through semantic evidence to produce deeply grounded, query-focused CoTs. It then applies reinforcement learning to align CoT generation with both the inference-time setting and the reranking objective, optimizing listwise metrics and passage-level discrimination to produce reusable reasoning chains for reranking. Experiments on the BRIGHT benchmark demonstrate that QReason significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.
- 中文摘要
文章重排序在信息检索中起着关键作用,通过优化候选文章的排序,更好地反映相关性。现有的列表式LLM排序器通过Chain-of-Thought(CoT)推理能够有效处理复杂查询,但由于滑动窗口策略,存在大量冗余和高延迟,这些策略反复生成高度相似的CoT。为此,我们提出了QReason,一个解耦框架,将以查询为中心的推理与窗口特定文章相关性评估区分开来。具体来说,QReason引入了专用重写工具,生成一次以排名为导向的推理查询,捕捉查询的核心意图,避免冗余推理,然后通过非推理重新排序器在所有窗口中重复使用该查询。重写器通过两阶段过程进行训练,首先通过语义证据进行监督微调,辅以相关段落指导,生成深度扎根、以查询为中心的CoT。随后应用强化学习,使CoT生成与推理时间设置和重排序目标对齐,优化列表级指标和文章级辨别,生成可重复使用的重排序推理链。BRIGHT 基准测试的实验表明,QReason显著减少了冗余推理,排名性能与强推理重排序器相当甚至更好,并且优于现有查询重写模型。
ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
ToolSearcher:通过强化学习大规模优化工具选择
- Authors: Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.30906
- Pdf link: https://arxiv.org/pdf/2609.30906
- Abstract
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.
- 中文摘要
大型语言模型(LLM)在自然语言处理方面表现出色,但在与外部环境交互方面表现不佳。工具学习为将LLM扩展为可操作的代理提供了有前景的方式,工具选择是成功使用工具的关键前提。现有工作通常假设工具集较小或预定义,导致大规模工具选择尚未被充分探索。现实中的数据库包含大量多样的工具,使得LLM在上下文长度约束下难以有效搜索、区分和组合工具。我们将大规模工具选择视为代理强化学习的新挑战,强调现有基于知识的强化学习方法在选择工具时不足以兼顾兼容性。为应对这一挑战,我们提出了ToolSearcher,一种用于大规模工具选择中高效多回合搜索和细粒度优化的新型强化学习框架。具体来说,我们引入了类别约束工具辨别,以提升模型区分功能相似工具的能力;事件级搜索建模,显式优化多回合搜索中目标工具的发现;以及轨迹对齐的信用分配,为搜索选择过程的不同阶段提供细粒度的奖励信号。大规模工具选择基准测试的广泛实验表明,ToolSearcher 在涉及迭代搜索和复杂工具组合的复杂环境中,始终优于一组强基线。
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
PORL:针对工作坊排班问题的预训练离线强化学习
- Authors: Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.30948
- Pdf link: https://arxiv.org/pdf/2609.30948
- Abstract
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.
- 中文摘要
作业车间调度问题(JSSP)是工业优化中的一个基础组合优化问题。本研究引入了预训练离线强化学习(PORL),这是一种结合基于模拟的在线预训练与对生产特定数据的离线微调相结合的混合方法。通过在线交互进行强化学习使得探索通用调度策略成为可能,但通常依赖于仿真环境,可能存在仿真与现实之间的差距。相比之下,离线强化学习通过从历史数据中学习避免与环境的直接交互,但其性能受到数据集质量和覆盖率的强烈影响。PORL结合了两种范式的优势,先通过在线交互学习通用调度策略,随后离线将其调整至目标分布。引入基于KL发散的策略约束,限制微调过程中偏离预训练策略的情况。该方法在JSSP实例中评估,数据分布偏移,数据集由启发式、噪声专家和随机行为策略生成。结果显示,PORL始终比独立离线强化学习和考虑的一般调度基线实现更低的最优性差距。此外,随着数据集质量下降,PORL相较独立离线强化学习的优势增强,显示出对可用离线数据质量和覆盖率的敏感度降低。结果表明,离线调整预训练策略是工业调度环境中一种有前景的方法,尤其是在直接在线探索不切实际的情况下。
MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
MVVBench:视觉语言模型中四维推理的基准测试
- Authors: Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.30952
- Pdf link: https://arxiv.org/pdf/2609.30952
- Abstract
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
- 中文摘要
多视角视频理解需要整合多条且常常不重叠的摄像机流的空间和时间证据:跟踪视点间的实体切换,时间对齐事件,并推理潜在的4D连续性,而非单一可见帧。我们介绍MVVBench,基于真实多摄像机数据集构建的多视角视频推理基准。问题在视角和时间轴上均为单眼模糊:每个问题在指定输入集中的任何单一视角都无法回答,大多数问题在任何单一时刻也无法回答。每个问题只有通过跨视角和跨时间的联合推理,才能独一无二地解决。MVVBench涵盖多样的动态场景,探索六项能力:隐式/显式属性识别、隐式/显式相对距离、相对摄像机姿态和构图计数,配合人工QA和严格验证。除了基准测试外,我们还广泛分析当前视觉语言模型何时及为何成功或失败,描述了时间误定位、交叉视角身份断裂和脆弱多跳推理导致的错误。随后,我们研究了释放潜在多视角能力---任务特定思维链支架和结构化交叉视角证据聚合的推理时间引发策略---实现了无需再训练即可取得显著进步。最后,我们提出了带有可验证奖励的强化学习初步证据,表明基础模型中可能引发潜在的多视角能力,指出训练时间方法为未来研究的有前景方向。MVVBench共同提供了对4D多视角推理的严谨评估,并为未来实现可靠具身感知奠定基础。
Robust Successor Features
稳健的继任功能
- Authors: Erik Nikulski, Yamen Habib, Vicenç Gomez, Anders Jonsson, Rubén Moreno-Bote, Javier Segovia-Aguas
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.31016
- Pdf link: https://arxiv.org/pdf/2609.31016
- Abstract
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneously from several articles in the field of operations research. In Robust RL, the transition kernel is unknown, and the goal is to maximize the expected reward under this uncertainty. Our work unifies these two paradigms through robust successor features, which generalize across both the reward function and the transition kernel, under the assumption that tasks are linear Markov Decision Processes. We derive a bound on Generalized Policy Improvement (GPI) that explicitly quantifies how performance degrades with the mismatch between transition kernels, recovering existing successor-feature guarantees when dynamics are shared. Finally, the generalization capabilities of robust successor features are validated on several grid-based benchmarks and compared to previous alternatives that focus solely on either the reward or the transition kernel.
- 中文摘要
强化学习中的泛化指的是在智能体接受不同任务训练后,能够在未见任务中执行接近最优策略的能力。基于后继表示的开创性工作及对函数近似的进一步调整,强化学习中的转移传统上专注于推广到仅奖励函数不同的任务。后继表示引入十年后,强健强化学习同时从运筹学领域的多篇文章中出现。在强健强化学习中,过渡核未知,目标是在该不确定性下最大化期望奖励。我们的工作通过稳健继承特征统一了这两种范式,这些特征在奖励函数和转移核上均可推广,前提是任务是线性马尔可夫决策过程。我们推导出了通用策略改进(GPI)的界限,明确量化了性能如何随着过渡核之间的不匹配而下降,并在动态共享时恢复现有的继任特征保证。最后,稳健继任特征的泛化能力在多个基于网格的基准测试上得到了验证,并与之前仅关注奖励或过渡核的替代方案进行了比较。
Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control
高速精度:液压挖掘机控制的高效在线模型基础钢筋学习
- Authors: Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2609.31025
- Pdf link: https://arxiv.org/pdf/2609.31025
- Abstract
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.
- 中文摘要
对于具有复杂驱动动力学的机器人来说,精确且高速的控制依然具有挑战性。直接在硬件上学习还受到现实世界交互成本的限制。我们提出了一个在线基于模型的强化学习框架,从零开始学习概率动力学系绵模型,用于基于采样的模型预测控制。精确门控轮廓目标对路径精度的进度奖励为条件,优先考虑精度而非速度。在数据驱动挖掘机模拟器中,该框架的采样效率高于评估的基于模型的强化学习基线。我们通过直接在一台11.5吨的Menzi Muck M445液压挖掘机上学习,无需演示或模拟预训练,验证了该框架。经过20分钟的交互,控制器达到了与之前基于100-150分钟数据训练的学习控制器相当的跟踪精度。40分钟后,高速运行时平均路径误差达亚厘米。
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
抽象阶梯的上下:语言代理的基于代码的技能
- Authors: Bartłomiej Cupiał, Jens Tuyls, Maciej Wołczyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Miłoś, Karthik R. Narasimhan
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31076
- Pdf link: https://arxiv.org/pdf/2609.31076
- Abstract
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.
- 中文摘要
语言代理在需要长时间低级动作序列的环境中难以行动和学习。基于代码的抽象可以通过让这些代理调用可重复使用的技能而非反复选择单个动作,从而提高生产力。代码处理重复的局部决策,而语言模型决定使用哪些技能以及如何组合它们。然而,抽象存在漏洞,超出技能能力范围的情境可能需要回归原始动作。基于生产力与灵活性之间的权衡,我们系统地研究基于代码的动作抽象如何影响语言代理的性能、推理成本和学习。我们在NetHack中研究这一点,这是一个具有挑战性的长视野游戏环境,使用CodeHack——我们基于代码的自然语言描述技能库。我们利用该库比较仅使用原语的代理与仅使用语义技能或结合原语的代理。我们在三种环境中评估这些代理:零机会提示、监督微调和强化学习。在NetHack的广泛零机会评估中,我们发现技能与原语相比,游戏进程几乎是三倍,且每集推理成本降低了86%。将技能与原语结合保留了大部分优势,同时保留了回归低级动作的路径。最后,在强化学习中,我们发现基于技能的代理学习速度显著快于对原语行动的代理,在同一训练预算下,地下城平均提升幅度为7.2倍。这些结果表明,提供的技能库可以提升性能、效率和学习能力,同时保留原语在库不足时提供灵活性。我们与训练和评估代码一起发布了CodeHack。
Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning
监控越狱:无编码推理规避思维链监控
- Authors: Julian Schulz
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31121
- Pdf link: https://arxiv.org/pdf/2609.31121
- Abstract
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing them when a monitor detects reasoning about the side task. Surprisingly, models learn to evade monitors without encoding their reasoning. Instead, they learn to phrase and format their chains of thought such that monitors fail to flag side task reasoning, while the reasoning remains completely transparent to human readers. We call this phenomenon monitor jailbreaking. We find that monitor jailbreaking arises across different model sizes, monitors, and tasks. Jailbreaks generalize to monitors not seen during training, including both less and more capable monitors, and transfer across different monitor prompts. While jailbreaking strategies appear simple, manually replicating them does not reliably fool monitors. Finally, we show that paraphrasing is an effective defense: paraphrasing a jailbroken CoT allows the same monitor to correctly flag it, while still allowing the model to perform both tasks.
- 中文摘要
思维链(CoT)监控是一种有前景的安全技术,能够在模型行动前检测出问题推理。一个关键问题是编码推理,即模型以监控者和人类无法解读的方式隐藏其真实推理。CoT监控器在强化学习过程中的优化压力被认为是此类行为的可能驱动因素。我们通过训练推理模型执行主任务和副任务,同时在监视者检测到对支线任务的推理时进行惩罚来研究这一点。令人惊讶的是,模型学会了在不编码其推理的情况下规避监视者。相反,它们学会了措辞和格式化思维链,使得监视者无法标记副任务推理,而推理对人类读者则完全透明。我们将此现象称为监控越狱。我们发现监控越狱发生在不同模型大小、监视器和任务中。越狱会泛化到训练中未见的监视器,包括能力较弱和更强的监视器,并在不同的监视提示间转移。虽然越狱策略看似简单,但手动复制它们并不能可靠地欺骗监视器。最后,我们证明了意译是一种有效的防御:转述越狱的CoT可以让同一监视器正确标记它,同时模型仍能执行这两种任务。
JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
JevAdvBench:校准决策模型中强化学习的基准与黑盒攻击
- Authors: Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang
- Subjects: Subjects:
Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2609.31142
- Pdf link: https://arxiv.org/pdf/2609.31142
- Abstract
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input. Project website: this https URL
- 中文摘要
用校准决策强化学习(RLCD)训练的模型,如Jev,会以概率、选择或评分回答关于输入状态的类型问题,软件在不读者的情况下对答案进行反应。其鲁棒性尚未被测量:对抗性基准测试对模型生成或执行的内容进行评分,而类型化模型不生成任何信息,即使经过操作也返回一个结构良好的答案。测量也很难,因为相同的请求可能返回不同的答案,大多数可用标签来自模型本身,API会在视野之外预处理每个请求。我们的核心思想是将每个被攻击的决策与模型自身的干净决策进行评分,而非标签,并将其与相同重运行引起的变化进行判读。基于此,我们引入了JevAdvBench,据我们所知这是RLCD模型的首个对抗基准测试,包含812个输入问题,涵盖66个场景,以及一套包含9,744种单次编辑变体的黑箱攻击套件,每个变体编辑请求的一部分,计费输入令牌确认编辑已到达模型。在jev-1.13.0中,重述保持在重跑基线的1.2个百分点以内,模式外字段永远不会到达模型。相比之下,附加在状态上的一个未经验证的意见会翻转12.1%的决策,统计上与最强注入命令(10.1%)持平,且38%的自信答案低于0.8置信阈值,从而被引导至人工审核。因此,基于RLCD模型构建的应用应将状态视为不可信输入,这点有争议。项目网站:此 https URL
Deep Reinforcement Learning for Misbehavior Detection Under Partially Observable V2X Data
在部分可观测的V2X数据下进行不良行为检测的深度强化学习
- Authors: Roshan Sedar, Charalampos Kalalas
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2609.31217
- Pdf link: https://arxiv.org/pdf/2609.31217
- Abstract
Misbehavior detection in vehicle-to-everything (V2X) systems is essential for ensuring the semantic correctness of exchanged messages and preventing the dissemination of falsified information. Existing data-centric misbehavior detection approaches largely rely on statistical validation or supervised machine learning models under the implicit assumption of fully observable V2X streams. In practice, however, vehicular environments are inherently partially observable due to hardware failures, intermittent connectivity, and environmental occlusions. Moreover, missingness itself can be strategically exploited by adversaries to evade detection. In this paper, we study misbehavior detection under incomplete V2X observations and propose a deep reinforcement learning (DRL)-based detection framework that learns adaptive policies with incomplete data. We further introduce an adversarial threat model in which attackers exploit or deliberately induce missingness to evade detection, including evasion via natural occlusions and adversarial feature suppression. Extensive experiments conducted on the VeReMi dataset under various missingness patterns demonstrate that DRL significantly outperforms a powerful XGBoost baseline under natural partial observability. However, results also reveal a critical vulnerability: DRL policies can be highly susceptible to evasion attacks that strategically exploit natural missingness. In contrast, DRL exhibits more gradual degradation under direct feature suppression compared to static tree-based models.
- 中文摘要
车辆对一切(V2X)系统中的不良行为检测对于确保交换消息的语义正确性和防止伪造信息传播至关重要。现有以数据为中心的异常行为检测方法主要依赖统计验证或监督机器学习模型,隐含假设V2X流完全可观测。然而,实际上,由于硬件故障、间歇性连接和环境遮挡,车辆环境本质上部分可观察。此外,缺失本身也可能被对手策略性地利用以规避检测。本文研究了在不完整V2X观测下的异常行为检测,并提出了基于深度强化学习(DRL)的检测框架,在数据不完整时学习自适应策略。我们进一步引入了一种对抗性威胁模型,攻击者利用或故意诱导缺失以规避检测,包括通过自然遮挡和对抗特征抑制来规避。在VeReMi数据集上进行的大量实验在多种缺失模式下表明,DRL在自然部分可观测性下显著优于强大的XGBoost基线。然而,结果也揭示了一个关键漏洞:DRL策略极易受到利用自然缺失的规避攻击。相比之下,DRL在直接特征抑制下表现得比静态基于树的模型更为缓慢。
MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
MA-WAM:多智能体世界行动模型用于测试时间规划
- Authors: Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31281
- Pdf link: https://arxiv.org/pdf/2609.31281
- Abstract
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simultaneous actions of multiple agents. We propose Multi-Agent World-Action Model (MA-WAM), a test-time planning framework that enables a frozen multi-agent flow policy to evaluate futures of candidate joint actions. To our knowledge, MA-WAM is the first test-time world-model planner for multi-agent flow policies. MA-WAM predicts the consequences of each joint action according to cross-agent dependencies and enables efficient candidate scoring. Across 30 offline multi-agent reinforcement learning (MARL) settings on MAMuJoCo, SMAC, and MPE, MA-WAM achieves mean relative gains of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for 2.5% of the measured generation-and-scoring time.
- 中文摘要
多智能体协作任务需要不同智能体同时执行联合行动,每个智能体的行动会影响其他智能体的观察和响应。因此,需要一个世界模型来预测所有智能体联合行动后团队的回报。一个简单的扩展是直接将单智能体世界模型应用于每个智能体的行动,以逐步预测团队返回。然而,这种扩展未能捕捉多个智能体同时行动之间的依赖关系。我们提出了多智能体世界行动模型(MA-WAM),这是一种测试时间规划框架,使冻结的多智能体流策略能够评估候选联合行动的未来。据我们所知,MA-WAM是首个多智能体流策略的测试时间世界模型规划器。MA-WAM根据跨智能体依赖关系预测每个联合行动的后果,并实现高效的候选评分。在 MAMuJoCo、SMAC 和 MPE 的 30 个离线多智能体强化学习(MARL)设置中,MA-WAM 在直接执行时平均获得 22.0% 的相对收益,在均匀动作选择时获得 25.6%。根据 A100 GPU 的标准评估协议,MA-WAM 增加了 12.1 毫秒,占测量生成和评分时间的 2.5%。
G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
G2MAF:多智能体流策略的测试时间梯度指导
- Authors: Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2609.31286
- Pdf link: https://arxiv.org/pdf/2609.31286
- Abstract
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi Agent Flow (G2MAF), a refinement framework for optimizing joint policies at test-time. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate all agents' corrections while keeping the action both feasible and close to the frozen policy proposal. Across 24 MPE and SMAC settings, its canonical variant improves 20 frozen settings, with mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by about 6% only.
- 中文摘要
离线多智能体强化学习(MARL)从固定数据集学习合作策略,无需进一步的环境交互,且已学习的策略在部署时被冻结。这种冻结策略通常提出单一联合动作,并在部署时直接执行。然而,这种一次性部署往往承诺一个次优方案,即使更好的邻近替代方案与行为数据保持一致。为解决这个问题,我们提出了梯度引导多智能体流(G2MAF),这是一种用于测试时优化联合策略的精炼框架。G2MAF应用一个全局归一化的预测批评梯度来指导和协调所有代理的修正,同时保持动作可行且接近冻结策略提案。在24种MPE和SMAC设置中,其典型变体改进了20个冻结设置,MPE平均相对提升为9.2%,SMAC平均提升为8.9%,模型推理延迟仅增加了约6%。
Dynamic Sampling for Telemetry in Microservices: A Reinforcement Learning and Entropy-Based Approach
微服务遥测的动态采样:一种强化学习与基于熵的方法
- Authors: Renan Martins Alves, Jéferson Campos Nobre, Juliano Araujo Wickboldt
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2609.31292
- Pdf link: https://arxiv.org/pdf/2609.31292
- Abstract
Microservices architectures are increasingly deployed in cloud-based distributed environments, making application development and maintenance more dynamic, but also increasing the complexity of troubleshooting and observability. Distributed tracing tools are therefore essential for request analysis and debugging, despite introducing additional overhead that can be amplified by excessive and inefficient data collection. This article proposes RADAR (Reinforcement Learning Agent for Dynamic And Relevant trace sampling), an agent that combines reinforcement learning with a data entropy assessment to achieve more efficient capture of traces relevant to system monitoring, based on the OpenTelemetry standard. RADAR tests different sampling rules to discover which combination is most efficient. A test environment simulating a minimalist online store with several microservices distributed across a Kubernetes cluster served as the basis for the experiments, which evaluated the agent's convergence and the system's performance in terms of resource consumption and collected data quality. Results showed that RADAR reduced network bandwidth consumption by 97.4% and CPU usage by 99.0% compared to full data collection, also outperforming a fixed-rate sampling baseline. Beyond these resource savings, the approach preserved observability of critical scenarios, retaining approximately 85.6% of rare trace patterns and increasing the average entropy of the stored information by approximately 25%, validating the feasibility of using entropy to orchestrate telemetry autonomously and efficiently.
- 中文摘要
微服务架构越来越多地部署在基于云的分布式环境中,使应用开发和维护更加动态,但也增加了故障排除和可观测性的复杂性。因此,分布式追踪工具对于请求分析和调试至关重要,尽管会带来额外开销,且可能因过度且低效的数据收集而加剧。本文提出了RADAR(动态且相关跟踪采样强化学习代理),这是一种结合强化学习与数据熵评估相结合的代理,以实现更高效的系统监控相关跟踪捕获,基于OpenTelemetry标准。RADAR测试不同的采样规则,以发现哪种组合最高效。以模拟一个极简在线商店、分布在Kubernetes集群上的多个微服务的测试环境作为实验基础,评估了智能体的收敛性以及系统在资源消耗和收集数据质量方面的性能。结果显示,RADAR相比全数据采集降低了97.4%的网络带宽消耗和99.0%的CPU使用率,同时优于固定速率采样基线。除了这些资源节省外,该方法还保留了关键场景的可观测性,保留了约85.6%的稀有踪迹模式,并将存储信息的平均熵提高了约25%,验证了利用熵自主高效编排遥测的可行性。
See to Reach, Feel to Grasp: Learning A Blind Grasp Reflex for Anthropomorphic Robotic Hands
看到够到,感觉到抓握:学习拟人化机器人手的盲抓反射
- Authors: Alexander Alexiev, Tzu-Yuan Lin, Sang Min Kim, Ho Jae Lee, Yonghyeon Lee, Sangbae Kim
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.31323
- Pdf link: https://arxiv.org/pdf/2609.31323
- Abstract
In this work we study if a robotic hand using proprioception alone can grasp diverse objects with no visual observation. We present a modular dexterous grasping architecture that separates global arm motion from local contact control. An independently controlled arm guides the hand toward the object, while a reinforcement learning policy grasps and stabilizes it using only hand proprioceptive feedback. We call this \textit{a blind grasp reflex}: grasping without images, object poses, or geometric observations. A learned stable-grasp score determines when the object is securely held, allowing the arm to begin post-grasp manipulation. This separation makes grasping a reusable hand-level skill that can be combined with independently designed arm controllers for various manipulation tasks. Experiments in simulation and on hardware demonstrate robust blind grasping across diverse objects and seamless composition with a range of arm controllers. Moreover, despite never observing contact geometry, the learned grasp score closely aligns with an independent physics-based measure of grasp stability. The resulting approach follows a simple principle: see to reach, feel to grasp. Project page: this https URL.
- 中文摘要
本研究中,我们研究仅使用本体感觉的机器人手是否能在无视觉观察的情况下抓住多样的物体。我们提出了一种模块化的灵活抓取架构,将全局手臂运动与局部接触控制分离。独立控制的手臂引导手向物体移动,而强化学习策略仅通过手部本体感受反馈来抓握和稳定。我们称之为 \textit{盲抓反射}:无图像、物体姿态或几何观察的抓取。学习到的稳定抓取评分决定物体何时被牢牢握住,使手臂能够开始抓取后的操作。这种分离使抓握成为可重复使用的手级技能,可以与独立设计的手臂控制器结合,用于各种操作任务。模拟和硬件实验展示了在不同物体间的稳健盲抓,以及多种手臂控制器的无缝组合。此外,尽管从未观察接触几何,但所学的抓握分数与基于物理的独立抓握稳定性测量高度吻合。最终方法遵循一个简单原则:见以伸手,触感握持。项目页面:此 https URL。
Adaptive Switching Between Leader-Based and Leaderless BFT Protocols
基于领导者与无领导者的BFT协议之间的自适应切换
- Authors: Sudip Bhujel, Yue Li, Ning Zhang, Y. Thomas Hou, Wenjing Lou, Yang Xiao
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2609.31388
- Pdf link: https://arxiv.org/pdf/2609.31388
- Abstract
Byzantine fault-tolerant (BFT) protocols are known for providing operational consistency and resilience in distributed systems. However, evolving network conditions, often driven by the network's inherent dynamism or adversarial influence, make it suboptimal to rely on a static protocol at all times. Existing BFT protocol adaptation solutions switch only among leader-based protocols and coordinate each switch through a separate consensus round, leaving them ineffective at handling severe asynchrony or situations in which an adaptive adversary targets the network's leader. We propose BFTide, a protocol adaptation architecture that enables a BFT system to intelligently and swiftly switch to a suitable protocol as network conditions shift. BFTide integrates a novel protocol switching layer that embeds protocol transition logic into the ongoing BFT operation, enabling safe and low-overhead transitions between partially synchronous leader-based protocols and asynchronous leaderless protocols. It further incorporates an offline-trained reinforcement learning policy that allows nodes to propose protocols at runtime based on observed system metrics. Experimental results show that BFTide reduces transaction latency under adverse network conditions compared with static BFT protocols and the state-of-the-art BFT protocol adaptation scheme BFTBrain (NSDI'25), while maintaining comparable throughput. The switching layer adds a modest 10-21% overhead to median latency when idle and requires no separate consensus round per switch.
- 中文摘要
拜占庭容错(BFT)协议以在分布式系统中提供操作一致性和韧性而闻名。然而,网络环境的变化,通常由网络固有的动态性或对抗性影响驱动,使得始终依赖静态协议并不理想。现有的BFT协议适配解决方案仅在基于领导者的协议之间切换,并通过单独的共识轮次协调每个交换机,因此在处理严重异步或自适应对手针对网络领导者的情况时效果有限。我们提出了BFTide,一种协议适配架构,使BFT系统能够智能且迅速地切换到合适的协议,随着网络状况的变化。BFTide集成了一种新型协议交换层,将协议转换逻辑嵌入BFT的持续运行中,实现部分同步基于领导者协议与异步无领导者协议之间的安全且低开销的过渡。它还结合了离线训练的强化学习策略,允许节点在运行时基于观察到的系统指标提出协议。实验结果显示,BFTide 在不利的网络条件下相比静态 BFT 协议和最先进的 BFT 协议适配方案 BFTBrain(NSDI'25)能降低交易延迟,同时保持相当的吞吐量。交换层在空闲时为中位延迟增加了 10-21% 的适度开销,且每个交换机无需单独的共识轮。
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
HySTAR:合作多智能体强化学习中稳定学分分配的锚定超图
- Authors: Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.31531
- Pdf link: https://arxiv.org/pdf/2609.31531
- Abstract
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a MAPPO-based framework that separates adaptive representation learning from a temporally consistent high-order value-decomposition basis. HySTAR anchors an overlapping sparse hypergraph as a uniformly covered decomposition scaffold, uses a spatiotemporal encoder to represent physical and task-dependent interactions, and combines temporal and structural relevance to construct agent-specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE demonstrate consistent improvements over MAPPO-style, value-factorization, and dynamic-grouping baselines. On the hardest SMAC settings, HySTAR achieves relative gains of 16.7\% over MAPPO and 15.6\% over HYGMA, ranks first on all six GRF scenarios, reduces Traffic Junction convergence epochs by up to 40.2\% relative to MAGIC, and obtains the highest MPE episode rewards. Controlled topology, agent-death, neighborhood, and parameter analyses support the benefit of anchoring the decomposition scaffold while adapting the propagated representations.
- 中文摘要
在部分可观察性和共享奖励下的合作多智能体强化学习需要将团队结果分配给单个智能体和高阶联盟。MAPPO风格的批评者将联合行为压缩为一个全局值,而动态重建分组拓扑的批评者则随着交互或主动智能体的演变,将映射从智能体和联盟转变为价值组件。我们将这种不一致称为结构目标漂移。我们介绍了HySTAR,这是一个基于MAPPO的框架,将自适应表示学习与时间一致的高阶价值分解基分开。HySTAR锚定了重叠稀疏超图,作为统一覆盖的分解支架,使用时空编码器表示物理和任务依赖的交互,结合时间和结构相关性构建智能体特定优势。在SMAC、GRF、Traffic Junction和MPE上的实验显示,相较于MAPPO风格、价值因式分解和动态分组基线,实现了持续的提升。在最难的SMAC设置下,HySTAR相较于MAPPO提升16.7%,较HYGMA提升15.6%,在所有六个GRF场景中均排名第一,相较MAGIC将交通交汇收敛时间缩短多达40.2%,并获得最高的MPE集数奖励。受控拓扑、代理死亡、邻域和参数分析支持锚定分解支架的优势,同时调整传播表示。
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
学会停止而不学会停止:自我监督自信训练提升推理效率
- Authors: Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2609.31619
- Pdf link: https://arxiv.org/pdf/2609.31619
- Abstract
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.
- 中文摘要
推理模型通常产生非常长的推理迹,导致推理计算成本高昂。现有方法通常通过推理时间的早期停止机制或明确鼓励训练中的推理时间缩短,例如带长度惩罚的强化学习来提高效率。我们证明,显著的效率提升可以来自另一种监督方式:\textit{confidence}。通过自监督过程,我们对推理模型进行微调,仅用600个训练问题预测其在推理轨迹中间点对答案的信心。信心仅作为训练目标使用:丢失不包含推理长度、效率或停止的目标。在推理阶段,微调模型使用标准生成过程,没有信心诱导或早期停止机制。尽管如此,自监督置信微调使推理更高效,在Gemma、Qwen、Nemotron和GPT-OSS模型的数学、科学和编码推理基准测试中,生成的代币在匹配准确率下可减少多达25%,效率提升与明确优化短推理的方法相当。推理事件分析进一步表明,置信监督在很大程度上保留了基础模型的高层次推理组合,而非选择性地抑制特定行为。我们的结果表明,高效的推理可能是学习元认知信号的下游结果,而非直接优化。
Keyword: diffusion policy
Policy-Calibrated DAgger: Offline Calibrated Noise Injection for Imitation Learning
策略校准DAgger:用于模仿学习的离线校准噪声注入
- Authors: Jenny Wang, George Kantor
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30462
- Pdf link: https://arxiv.org/pdf/2609.30462
- Abstract
Policies trained with imitation learning can accumulate errors over time, causing the robot to drift outside the training distribution. Existing methods mitigate this covariate shift by collecting additional data where the policy fails or is likely to fail. The first places the robot in unsafe conditions and the second requires choosing an appropriate noise distribution to collect new expert demonstrations under that noise. We propose Policy-Calibrated DAgger, a method that makes use of the properties of recent generative policies to estimate the policy's noise offline by using its own predicted action distribution. We measure a diffusion policy's spread of predicted actions at observations along the expert trajectory and measure its closed-loop error relative to a recorded trajectory. To address issues with measuring error in a multimodal action space, we guide the policy towards the trajectory during closed-loop control through partial denoising, and use properties of a diffusion model to unnormalize the measured error as if we did not guide it. We experiment in a scenario where a robot is tasked to reach an engine lever in a cluttered and narrow environment and show results in a 3D photorealistic simulator and a 2D planar reacher environment. We show that our method surpasses policies trained with dataset aggregation without noising and matches the performance of the best noise level in hindsight, without requiring a sweep over noise levels.
- 中文摘要
通过模仿学习训练的策略可能随着时间累积错误,导致机器人偏离训练分布。现有方法通过收集策略失败或可能失败的额外数据来缓解这种协变量偏移。第一种方法使机器人处于不安全状态,第二种则需要选择合适的噪声分布以在该噪声下收集新的专家演示。我们提出了策略校准DAgger方法,利用近期生成策略的特性,利用其自身的预测动作分布来离线估计策略噪声。我们测量扩散策略在专家轨迹沿观测处预测动作的分布,并测量其相对于已记录轨迹的闭环误差。为解决多模态作用空间中测量误差的问题,我们通过部分去噪引导策略向闭环控制轨迹方向移动,并利用扩散模型的特性使测量误差如同未被引导而非归一化。我们在一个场景中实验:机器人被要求在狭窄杂乱的环境中触及发动机操纵杆,并在3D写实模拟器和2D平面助景环境中展示结果。我们证明,我们的方法超越了用数据集聚合训练的策略而不产生噪声,且事后看来能匹配最佳噪声水平的性能,无需对噪声水平进行扫荡。
Aerial Manipulation in the Wild with Onboard Perception, Policy Learning, and Whole-Body Control
野外空中操控,结合机载感知、策略学习和全身控制
- Authors: Yuanzhu Zhan, Yufei Jiang, Zemu Zhang, Junyi Geng
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2609.30521
- Pdf link: https://arxiv.org/pdf/2609.30521
- Abstract
Aerial manipulation in outdoor environments remains challenging due to the simultaneous requirements of reliable state estimation, stable aerial motion, and precise manipulation under external disturbances. In this work, we present a real-world outdoor aerial manipulation framework that integrates imitation learning, onboard LiDAR-inertial state estimation, and whole-body model predictive control. A Diffusion Policy is trained from manipulation demonstrations to generate desired end-effector motions from onboard observations. These learned commands are executed by a whole-body MPC that jointly coordinates the aerial platform and manipulator to realize the desired end-effector trajectory. To eliminate reliance on external motion-capture infrastructure, the platform employs onboard LiDAR-inertial odometry for state estimation during outdoor operation. We validate the complete framework on a physical aerial manipulator and demonstrate successful execution of outdoor manipulation tasks. The experimental results show that demonstration-driven manipulation policies can be effectively integrated with onboard state estimation and model-based whole-body control to enable aerial manipulation beyond controlled indoor environments.
- 中文摘要
户外空中操作依然充满挑战,因为需要同时要求可靠的状态估计、稳定的空中运动以及在外部干扰下的精确操作。本研究提出了一个现实世界的户外空中操作框架,整合了模拟学习、机载激光雷达惯性状态估计和全身模型预测控制。通过操作演示训练出扩散策略,从机载观测中生成期望的终端执行器运动。这些学习到的指令由全体MPC执行,该MPC联合协调空中平台和机械臂,实现期望的终端执行器轨迹。为消除对外部动作捕捉基础设施的依赖,平台在户外运行时采用了机载激光雷达惯性里程估计状态。我们在物理空中机械臂上验证了完整框架,并展示了户外操作任务的成功执行。实验结果表明,演示驱动的操作策略可以有效整合于机载状态估计和基于模型的全身控制,从而实现超越受控室内环境的空中操作。