生成时间: 2026-10-09 23:05:19 (UTC+8); Arxiv 发布时间: 2026-10-09 20:00 EDT (2026-10-10 08:00 UTC+8)
今天共有 66 篇相关文章
Keyword: reinforcement learning
Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
教机器人狗新把戏:通过强化与模仿学习结合对抗性任务选择,实现多样化四足动物技能
- Authors: Lemon Foxmere, Anthony Furman, Yizheng Du, Oliver Chang, Leilani Gilpin, Steve McGuire
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.10601
- Pdf link: https://arxiv.org/pdf/2610.10601
- Abstract
Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional stages train a student policy with a multi-teacher distillation setup that uses a combined RL and Imitation Learning (IL) objective under an adversarial task selection process that focuses training on the worst-performing task. With this method, we train a student policy that performs 22 tasks using 8 teachers. Evaluations show our method preserves motion quality and tracks commands more accurately than PPO and distill-then-finetune baselines, and in some cases generalizes to new tasks without explicit training. Finally, we demonstrate real-world robustness by deploying the resulting policy on a Unitree B1 quadruped. Video: this https URL
- 中文摘要
强化学习(RL)使腿部机器人能够在单一任务环境中执行多种技能。然而,农场机器人或太空探索等应用需要多样技能,如移动、挖掘或近距离测量。由于多任务学习中样本效率低和任务间梯度冲突等挑战,训练端到端策略仍难以解决这一问题。我们提出了一种三阶段方法,训练单一策略执行行走、挖掘和跳跃等不同任务,并将其组合成新颖的行为,如爬行。首先,利用强化学习对狭义任务进行多教师策略训练。随后,另外两个阶段通过多教师提炼框架训练学生策略,该方案结合强化学习和模仿学习(IL)目标,通过对抗性任务选择过程,专注于表现最差的任务。通过该方法,我们训练了一个学生策略,该策略由8位教师执行22项任务。评估显示,我们的方法比PPO和提取后微调基线更准确地保持了运动质量,并能更准确地跟踪命令,有时甚至无需明确培训即可推广到新任务。最后,我们通过在Unitree B1四足动物上部署策略,展示了现实世界的鲁棒性。视频:此 https URL
Sample-Efficiency of Kolmogorov-Arnold Networks
Kolmogorov-Arnold 网络的样本效率
- Authors: Kevin Riehl, Shaimaa K. El-Baklish, Fan Wu, Anastasios Kouvelas
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10627
- Pdf link: https://arxiv.org/pdf/2610.10627
- Abstract
Deep reinforcement learning has achieved substantial performance gains over classical control approaches. Yet, a central challenge to learning in real-world applications is acquiring costly samples. Kolmogorov-Arnold Networks are a recently proposed architecture that can learn physical relationships in control problems effectively, with significantly higher parameter efficiency and interpretability when compared to Multi-Layer-Perceptron architectures. In this work, we systematically study sample-efficiency using computational experiments, covering the Feynman dataset and the Gymnasium RL benchmark. The results show that similar performance can be achieved with 40% fewer samples using the Kolmogorov-Arnold architecture, and that relative performance improvements up to 50% occur during the training process. The observed gains are robust to varying levels of noise in rewards. These results highlight the potential of the Kolmogorov-Arnold architectures for more sample-efficient reinforcement learning. Code: this https URL
- 中文摘要
深度强化学习在性能上相较于经典控制方法取得了显著提升。然而,现实应用中学习的一个核心挑战是获取昂贵的样本。Kolmogorov-Arnold 网络是一种新提出的架构,能够有效学习控制问题中的物理关系,参数效率和可解释性显著高于多层感知器架构。本研究通过计算实验系统研究样本效率,涵盖费曼数据集和 Gymnasium RL 基准。结果显示,使用 Kolmogorov-Arnold 架构,样本数量减少 40% 即可实现类似性能,且在训练过程中相对性能提升可达 50%。观察到的提升对不同噪声水平的奖励表现出稳健性。这些结果凸显了 Kolmogorov-Arnold 架构在提升样本效率强化学习方面的潜力。代码:此 https URL
Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction
医疗代币用于诊断预测的覆盖意识推理
- Authors: Kaisong Zhang, Haotian Fang, Junmeng Zhou, Hang Lv, Yulan Pan, Yanchao Tan
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10641
- Pdf link: https://arxiv.org/pdf/2610.10641
- Abstract
Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language. However, reinforcement learning for LLM reasoning commonly rewards each trajectory according to the correctness of its final answer. In next-visit diagnosis prediction, multiple diagnoses can be simultaneously valid, but independently rewarding one diagnosis per trajectory does not distinguish repeated hits from coverage of different diagnoses. The policy can therefore concentrate on a few correct diagnoses, leaving others uncovered. Meanwhile, LLM tokenizers can split ICD codes into several generic tokens with limited clinical meaning, requiring multiple decoding steps to predict each diagnosis and hindering reasoning over a large disease vocabulary. To address both challenges, we propose CARing, a framework that represents diagnoses with compositional Semantic IDs (SIDs) and optimizes reasoning trajectories for multi-label coverage. Concretely, we first encode ontology-enriched disease semantics into compact SIDs through residual quantization, and ground the resulting SID tokens in natural language and longitudinal EHR contexts through multi-task alignment and reasoning-enriched training to unlock transferable LLM reasoning. CARing further improves unordered multi-label prediction through a coverage reward for reinforcement learning and multi-positive supervision. At inference time, the model supports both efficient direct constrained decoding and multi-chain reasoning with rank fusion. On MIMIC-III and MIMIC-IV, CARing exceeds all EHR-trained baselines in weighted F1 and attains the highest top-k recall at every reported cutoff, including R@30 of 46.04% and 46.52% in reasoning mode. Our codes and logs are available at this https URL.
- 中文摘要
大型语言模型(LLMs)因其能够整合纵向临床证据并用自然语言推理,从而在下次就诊诊断预测方面具有良好潜力。然而,LLM推理的强化学习通常会根据最终答案的正确性奖励每个路径。在下次就诊诊断预测中,多个诊断可以同时有效,但独立奖励每个轨迹中的一个诊断,并不能区分重复命中与覆盖不同诊断。因此,该政策可以专注于少数正确诊断,而忽略其他诊断。与此同时,LLM分词器可以将ICD代码拆分为多个具有有限临床意义的通用标记,需要多次解码步骤来预测每个诊断,并阻碍在庞大疾病词汇中进行推理。为应对这两个挑战,我们提出了CARing框架,该框架表示具有组合语义ID(SID)的诊断,并优化推理轨迹以实现多标签覆盖。具体来说,我们首先通过残差量化将本体丰富的疾病语义编码为紧凑的SID,并通过多任务对齐和推理丰富训练将所得SID代币置于自然语言和纵向EHR语境中,以解锁可转移的LLM推理。CARing通过强化学习和多正向监督的覆盖奖励,进一步提升了无序多标签预测。在推理阶段,模型支持高效的直接约束解码和带秩融合的多链推理。在MIMIC-III和MIMIC-IV测试中,CARing在加权F1中超过所有EHR训练基线,并在每个报告的截止点都达到最高的Top-K回忆,包括推理模式下R@30 46.04%和46.52%。我们的代码和日志可在此 https URL 获取。
Safe Learning of Adaptive Control Policies for Remote Patient Monitoring
远程患者监测自适应控制政策的安全学习
- Authors: Ramanan Tamizholi, Siddharth Chandak, Isha Thapa, Nicholas Bambos, David Scheinker
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.10720
- Pdf link: https://arxiv.org/pdf/2610.10720
- Abstract
Remote Patient Monitoring (RPM) enables continuous observation of patients in their daily environments, improving both health outcomes and quality of life. A key challenge in RPM is determining the optimal monitoring intensity, while balancing patient safety and monitoring costs. This problem is further complicated when system parameters, such as transition probabilities and costs, are initially unknown. We develop a learning-based control framework that estimates these parameters and adapts the monitoring policy in real time. The proposed approach is an online model-based reinforcement learning algorithm tailored to RPM, with patient safety explicitly prioritized during exploration. We provide theoretical guarantees on safety and convergence to the optimal policy. Simulation results show that the algorithm converges to the optimal threshold-based policy, maintains low treatment costs, and reduces the risk of patients reaching critical health states.
- 中文摘要
远程患者监测(RPM)使患者能够在日常环境中持续观察,从而提升健康结局和生活质量。RPM的一个关键挑战是确定最佳监测强度,同时平衡患者安全与监测成本。当系统参数如过渡概率和成本最初未知时,问题更加复杂。我们开发了一个基于学习的控制框架,能够实时估计这些参数并调整监测策略。所提出的方法是一种基于在线模型的强化学习算法,专为RPM量身定制,在探索过程中明确优先考虑患者安全。我们提供了安全性和趋同于最优策略的理论保证。模拟结果显示,算法能够收敛到基于阈值的最优策略,保持低治疗成本,并降低患者达到危急健康状态的风险。
Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
对话任务对表格数据的消歧:泄漏感知表述、基准套件与训练
- Authors: Nafiseh Ghoroghchian, Luis Scoccola, Tina Sedaghat, Omid Vaheb, Hannah Chen, Dino D'Agostino, Keyvan Golestan
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.10740
- Pdf link: https://arxiv.org/pdf/2610.10740
- Abstract
Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
- 中文摘要
对表格数据进行对话式任务消歧,通过对话解决用户预期任务的缺失信息,然后再生成表或数据库的解决方案。现有的评估和培训缺乏泄漏意识基础。任务成功结合了智能体的消歧义和解决方案生成能力,也可以反映oracle泄漏,即用户模拟器揭示的超出真实用户的信息。现有数据集也缺乏对歧义和访问边界的共享表示。我们引入了模糊可验证任务的概念,形式化歧义和解析,将代理分解为请求策略和解决方案策略,将环境分解为预言机和验证者。该框架提供了基线和指标,用于分别评估任务消歧,并与解决方案生成分开,正式定义oracle泄露,无判判泄漏诊断,以及请求策略的训练目标。我们用AmbiTab实现了文本转SQL框架,AmbiTab是一个基准测试套件,统一了六个模糊数据集,统一表示代理、预言机和验证者可访问的内容。我们评估澄清策略和预言机泄露,并用强化学习训练提问策略。受过训练的提问者提升了六个数据集的消歧义指标,五个数据集的任务成功率提升,我们的泄露诊断衡量训练对预言机泄露的影响。
Adaptive E2E Monitoring Path Selection with Deep Reinforcement Learning in Multi-domain Optical Networks
多域光网络中的自适应端对端E监测路径选择与深度强化学习
- Authors: Soham Choudhury, Martin Bojinov, Jason P. Jue, Genya Ishigaki
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI)
- Arxiv link: https://arxiv.org/abs/2610.10753
- Pdf link: https://arxiv.org/pdf/2610.10753
- Abstract
Network slicing over multi-domain optical networks enables resource isolation for high-bandwidth, virtually-dedicated networks. However, guaranteeing end-to-end (E2E) performance remains challenging when domains withhold internal topology and performance metrics from external entities. Under this limited visibility, a control function (slice coordinator) needs to select E2E monitoring paths to detect and localize link failures. This paper formulates two complementary optimization problems, the initial and progressive monitoring path selection problems, which together realize closed-loop autonomous monitoring at the slice coordinator level. The initial problem maximizes failure coverage without prior information, while the progressive problem concentrates additional paths around suspected failure locations identified from prior monitoring. We propose a deep reinforcement learning (DRL) algorithm that adaptively selects an optimal set of monitoring paths for each phase. Experiments using the GNPy optical network simulator demonstrate that our approach outperforms baseline algorithms for both problems and achieves near-optimal localization in the initial selection, suggesting the feasibility of AI-driven closed-loop failure localization in future multi-domain optical network architectures.
- 中文摘要
多域光网络上的网络切片实现了高带宽、虚拟专用网络的资源隔离。然而,当域对外部实体隐瞒内部拓扑和性能指标时,确保端到端(E2E)性能仍然具有挑战性。在这种有限的可见性下,控制函数(切片协调器)需要选择端对端监控路径以检测和定位链路故障。本文提出了两个互补的优化问题:初始监测路径选择问题和渐进监测路径选择问题,这两者共同实现片协调者级别的闭环自主监控。初始问题在无先行信息的情况下最大化故障覆盖率,而渐进问题则围绕先前监控识别的疑似故障位置集中更多路径。我们提出了一种深度强化学习(DRL)算法,能够自适应地为每个阶段选择最优监控路径集。使用GNPy光网络模拟器的实验表明,我们的方法在两个问题上都优于基线算法,并在初始选择中实现了近乎最优的定位,这暗示了未来多域光网络架构中AI驱动闭环故障定位的可行性。
Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
能源转型中价值创造的战略投资决策:强化学习方法
- Authors: Yasaman Cheraghi (1), Reidar B. Bratvold (1), Aojie Hong (2), Ressi B. Muhammad (1), Sergey Alyaev (3) ((1) Department of Energy and Petroleum Engineering, University of Stavanger, Norway, (2) Independent Researcher, Stavanger, Norway, (3) NORCE Norwegian Research Centre, Bergen, Norway)
- Subjects: Subjects:
Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.10768
- Pdf link: https://arxiv.org/pdf/2610.10768
- Abstract
The global challenge of climate change has driven significant steps to reduce CO2 emissions, guided by international agreements like the Paris Agreement of 2015. Acting too slowly could result in future losses and reputational damage, while moving too quickly could jeopardize shareholder value due to the marginal profitability or potential losses due to technology immaturity of many renewable projects. To navigate this complex transition, energy companies must adopt Sequential Decision Making (SDM) strategies to maximize value creation from decision flexibility under uncertainties. To support this, we developed a custom simulation environment to model the dynamic energy landscape up to 2050. Building on this, we designed a multi-criteria SDM framework that explores various decision strategies related to different portfolios for allocating funds across three sectors: oil & gas, renewables, and CO2 reduction. It aims to maximize value during the transition while accounting for uncertainties in productions, energy prices, and costs. This framework has three objectives: maximizing profit, minimizing CO2 social costs, and enhancing competitive advantage in the renewable energy sector. This research evaluates the use of Reinforcement Learning (RL) to identify optimal investment policies within the defined SDM framework. The agent's sequential decisions shape a virtual dynamic environment by influencing key variables such as oil and gas production, renewable energy output, CO2 emissions, and revenues. Through repeated interaction, the RL algorithm explores the state space and learns an optimal policy under uncertainty. We benchmark the RL strategy against a set of manually defined baseline policies and find it consistently outperforms them in adaptability and long-term value creation.
- 中文摘要
气候变化这一全球挑战推动了大量减少二氧化碳排放的措施,这些措施受到2015年巴黎协定等国际协议的指导。行动过慢可能导致未来损失和声誉损害;过快则可能因边际盈利能力或因技术不成熟而危及股东价值。为应对这一复杂转型,能源公司必须采用顺序决策(SDM)策略,以最大化在不确定性下通过决策灵活性创造价值。为此,我们开发了定制模拟环境,以模拟2050年前的动态能源格局。基于此,我们设计了一个多标准SDM框架,探讨与石油天然气、可再生能源和二氧化碳减排三大行业不同投资组合相关的多种决策策略。该框架旨在在转型期间最大化价值,同时考虑生产、能源价格和成本的不确定性。该框架有三个目标:最大化利润、最小化二氧化碳社会成本,以及增强可再生能源领域的竞争优势。本研究评估了强化学习(RL)在定义的SDM框架内识别最优投资政策的应用。代理的连续决策通过影响关键变量如石油和天然气产量、可再生能源产出、二氧化碳排放和收入,塑造虚拟动态环境。通过反复交互,RL算法探索状态空间,在不确定性下学习最优政策。我们将强化学习策略与一组手动定义的基线政策进行基准对比,发现其在适应性和长期价值创造方面始终优于这些策略。
VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
VICO:视觉环境为视觉语言模型推理而共同演进
- Authors: Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10782
- Pdf link: https://arxiv.org/pdf/2610.10782
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
- 中文摘要
带有可验证奖励的强化学习(RLVR)已成为训练后视觉语言模型(VLM)的标准配方,但它通常假设训练环境是静态的。随着演员的提升,固定任务逐渐脱离其学习前沿:许多任务变得琐碎,另一些则无法解决;学习信号因此崩溃。我们主张VLM的后训练应与演员一起演进视觉环境,而不仅仅是演员本身。我们提出了VICO,一种共进化框架,演员与环境即重写者(EnvRewriter)共同训练:EnvRewriter编辑可验证的图像侧结构,如场景图、图表表或受保护区域掩码,并重新渲染以生成标签有效的训练样本,其难度通过基于通过率的奖励根据演员当前能力校准。该循环持续将任务难度与演员能力重新对齐,无需额外人工注释。在九个涵盖数学推理和视觉基础理解的多模态基准测试中,VICO-8B 在域外任务上相较基础模型提升高达 +5.0%,在最强的自我演化和文本编辑共演化基线上分别提升 +4.3% 和 +8.4%,并且在使用减少 16-160 倍的标注样本后,仍可与图表专用的 RLVR 方法相当。通过从人工标记监督转向图像编辑共进化,VICO 为视觉推理提供了超越静态语料库 RLVR 的可扩展路径。
NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
NavGPT-3:在分层导航运行时中利用上下文
- Authors: Gengze Zhou, Yicong Hong, Jiazhao Zhang, Xunyi Zhao, Jian Zhou, Zixing Lei, Zun Wang, Chongyang Zhao, Xionghui Chen, Stephen Gould, Anton van den Hengel, Qi Wu
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10787
- Pdf link: https://arxiv.org/pdf/2610.10787
- Abstract
Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.
- 中文摘要
使用长视野智能体强化学习训练的语言模型可以通过推理泛化知识,表达精确动作,并在多步中追求目标,提高了具身智能体理解和决策的上限。然而,物理交互仍属于行动策略领域,提供密集且低延迟的控制。我们介绍NavGPT-3,一种连接两种模型的机束,其上构建了类似操作系统的运行时:推理、行动和监控作为线程运行,拥有自己的上下文、工具和权限,运行时则调度并决定哪个线程控制机器人的运动,使机器人能够通过中断和线程切换来应对突发的现实事件。在其下方,我们的动作策略NavGPT VLA,基于1928万个样本训练,按场景变化比例分配视觉标记;仅其8B模型在R2R-CE上就达到74.51 SR,并以78.19 SR领先RxR-CE。在完整的线带下,NavGPT-3在R2R-CE(81.51 SR)上奠定了最前沿,并首次将自主智能体带入人类水平:在RxR-CE上,它在成功率(90.43对90.4 SR)和路径忠实度(78.47对77.7 nDTW)上与人类追随者匹配,每集约需1分22秒,而人类约需3分钟。我们全面简化了机束设计及两个模型之间的交互,展示了工具和动作策略如何塑造从语言-模型推理到物理控制的路径:当NavGPT VLA执行该路径时,推理循环缩短,系统的最小反应时间从每个语言-模型决策3-19秒降至每动作-策略步骤(1-2 Hz)0.5-1秒。这些结果表明,设计该具象接口对于连接前沿语言-模型智能与低层物理控制至关重要。我们将发布所有模型、代码和评估记录。
On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
上班时间:迈向准时且高效的时间预算AI代理
- Authors: Aaron Wang, Neelabh Madan, Vlad Sobal, Matthew Trager, Michael Kleinman, Elman Mansimov, Wei Xia, Stefano Soatto
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10833
- Pdf link: https://arxiv.org/pdf/2610.10833
- Abstract
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively. We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance. In the simplest setting, where the budget is stated only in the prompt, agents fail to translate the stated budget into controlled use of time. These failures arise from gaps in time awareness, since the harness provides no timing feedback, but also because they cannot reliably anticipate the duration of actions, and do not have a learned mapping from available time to an appropriate strategy. We investigate two complementary classes of interventions: harness-based mechanisms that expose timing information and enforce deadlines, and reinforcement learning with budget-aware rewards. Injecting timing information through the harness substantially improves budget adherence for Qwen3.6-27B without measurable loss in performance, while enforcement hooks tighten adherence further. RL with GRPO achieves near-perfect budget adherence on Zork I and generalizes to held-out budgets not seen during training, but does not improve task performance over the untrained harness on MLE-Bench. Once agents are made to respect the budget, they still fail to use additional time to improve task performance. RL-trained policies learn when to stop but often fill extra time with repeated actions, and GRPO training on multiple budgets tends to collapse toward the strategy learned for the shortest budget. Our results reveal a gap between time adherence and productive time allocation, which remains a central challenge for budget-conditioned agents.
- 中文摘要
我们研究小型LLM代理是否能在明确的墙壁时间预算下有效运行,同时尊重分配的运行时间并有效利用可用时间。我们在MLE-Bench Lite和Qwen3-4B的五项竞赛中评估Qwen3.6-27B,分别是Zork I(Jericho)上的两个代理基准测试,额外的计算时间可以显著提升性能。在最简单的情境下,预算仅在提示中说明,代理未能将预算转化为时间的受控使用。这些失败源于时间意识缺口,因为工具束不提供时间反馈,也因他们无法可靠预测动作持续时间,且没有从可用时间到适当策略的学习映射。我们研究了两类互补干预:基于利用机制,揭示时间信息并执行截止日期,以及基于预算的强化学习,提供预算感知奖励。通过束带注入计时信息,显著提升Qwen3.6-27B的预算依从性,且性能无可测量损失,而执行钩则进一步收紧了遵从性。带GRPO的强化学习在Zork I上实现了近乎完美的预算遵循,并推广到培训中未见的预留预算,但并未优于MLE-Bench上未受训练的约束器提升任务表现。即使代理被要求遵守预算,他们仍未能利用额外时间提升任务表现。强化学习训练的策略学会何时停止,但常常用重复行动填补额外时间,而多预算下的GRPO训练往往倾向于最短预算所学策略。我们的结果揭示了时间遵循与有效时间分配之间的差距,这仍是预算条件代理面临的核心挑战。
SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages
SPLIT-RL:分阶段感知-语言推理训练,具索赔级优势
- Authors: Raja Kumar, Rajat Koner, Ritwick Chaudhry, Zhuowei Li, Nishant Sankaran, Yifan Xing
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.10889
- Pdf link: https://arxiv.org/pdf/2610.10889
- Abstract
Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further introduce Claim-Level Advantage (CLA-GRPO), which decomposes VR-phase rollouts into atomic visual claims and provides a fine-grained advantage at claim level based on visual-type group formation. Although trained in two phases, trained policy is evaluated like GRPO model, with a single CoT call at inference time. Under this protocol, SPLIT-RL improves average accuracy over GRPO by 1.4-6.1 points across Qwen3-VL models from 2B to 30B-A3B and InternVL3.5-8B. Evaluating each capability using an oracle based diagnostic shows that answer-only GRPO leaves perception unchanged, whereas SPLIT-RL improves both VR and LR.
- 中文摘要
视觉语言(VL)推理需要一个模型既能从图像中提取相关且准确的信息(视觉推理,VR),又能从中推断答案(语言推理,LR)。带有可验证奖励的强化学习通常通过单一思维链训练两者,并以最终答案奖励为基础。这使得每个CoT代币在序列层面享有相同的优势,但无法区分能力特定的错误。我们提出了SPLIT-RL,一种分阶段的训练后方法,在不相交阶段训练VR和LR。由于一组的部署在一次一个能力上不同,群体相对优势将其隔离开来,每个阶段通过阶段特定奖励进行优化。我们进一步引入了索赔层优势(CLA-GRPO),它将VR阶段的推广分解为原子级的视觉索赔,并在基于视觉类型群体形成的索赔层面提供细粒度优势。虽然训练分两阶段,但训练策略的评估方式类似于GRPO模型,推断时只需一次CoT调用。在该协议下,SPLIT-RL在Qwen3-VL模型(2B至30B-A3B)和InternVL3.5-8B中,平均准确率提升了1.4-6.1个百分点。使用基于预言机的诊断评估每个能力表明,仅回答的GRPO保持感知不变,而SPLIT-RL则同时提升了VR和LR。
Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory
基于强化学习和博弈论的自适应多判别器WGAN框架,用于资源受限车辆互联网
- Authors: Farhoud Jafari Kaleibar, Amr M. Zaki, Marin Litoiu
- Subjects: Subjects:
Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
- Arxiv link: https://arxiv.org/abs/2610.10926
- Pdf link: https://arxiv.org/pdf/2610.10926
- Abstract
Managing machine learning workloads as a network service introduces a resource-orchestration problem distinct from conventional model training; which nodes should be allocated to a task, how communication and computation budgets should be divided among them, and how service quality should be sustained as connectivity and node availability change with mobility. Deploying Generative Adversarial Networks (GANs) in Internet of Vehicles (IoV) environments is a demanding instance of this problem; resource constraints, dynamic network topologies, and competing optimization objectives mean that traditional GAN architectures cannot simultaneously achieve high accuracy, efficient resource use, low delay, and low communication overhead. This paper introduces an adaptive multi-discriminator Wasserstein GAN (MD-WGAN) framework that integrates reinforcement learning with game-theoretic coordination to address these challenges jointly. In our framework, roadside units host generators paired with Deep Q-Network (DQN) agents that select discriminator subsets and manage distributed training across mobile vehicular nodes, while a game-theoretic coordination step allocates training epochs between generators and discriminators. A unified optimization objective ties adversarial learning quality to resource efficiency, communication overhead, and latency under vehicular constraints, allowing the framework to continuously adapt its training behavior as network conditions change. Evaluation on real-world NGSIM trajectory data shows that the framework attains prediction accuracy comparable to state-of-the-art GAN baselines - the lowest RMSE (1.029) and MAE (0.894) among all evaluated methods - while markedly improving resource efficiency: average CPU utilization is reduced by roughly 28% and mean memory usage by roughly 6%, at competitive communication overhead and latency.
- 中文摘要
将机器学习工作负载作为网络服务管理引入了与传统模型训练不同的资源编排问题;应分配哪些节点执行任务,通信和计算预算如何分配,以及随着移动性变化,连接性和节点可用性的变化如何维持服务质量。在车联(IoV)环境中部署生成对抗网络(GAN)是该问题的一个严峻实例;资源限制、动态网络拓扑和竞争的优化目标意味着传统GAN架构无法同时实现高准确性、高效资源使用、低延迟和低通信开销。本文介绍了一个自适应多判别器Wasserstein GAN框架,将强化学习与博弈论协调结合,共同应对这些挑战。在我们的框架中,路边单位与深度Q网络(DQN)代理配合,后者选择判别器子集并管理移动车辆节点间的分布式训练,同时博弈论协调步骤在生成器和判别器之间分配训练纪元。统一的优化目标将对抗学习质量与资源效率、通信开销和车辆约束下的延迟挂钩,使框架能够随着网络条件变化持续调整训练行为。对现实世界NGSIM轨迹数据的评估显示,该框架的预测准确度可与最先进的GAN基线相媲美——在所有评估方法中RMSE(1.029)和MAE(0.894)均为最低——同时显著提升资源效率:平均CPU使用率约降低28%,平均内存使用率约6%,在竞争性通信开销和延迟下降低。
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
目标条件强化学习的世界模型政策仲裁器
- Authors: Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart, Igor Gilitschenski
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10932
- Pdf link: https://arxiv.org/pdf/2610.10932
- Abstract
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.
- 中文摘要
离线目标条件强化学习(GCRL)产生了多样的目标达成算法,但没有单一算法能在环境、目标甚至同一任务的不同阶段中表现最佳。我们不只部署表现最好的策略,而是询问一组固定的目标条件策略是否可以作为组合使用,在每个状态决定哪个策略应采取行动。在每个状态下选择策略并不简单。策略自身的价值函数无法直接比较:它们可能使用不同的尺度,有些策略没有价值函数。我们需要根据每个策略可能达到的状态来判断,尽管我们一次只能执行一个策略。我们还需要避免频繁切换以导致控制不稳定。为应对这些挑战,我们引入了世界模型策略仲裁器(WMPA),这是一个测试时间框架,给定一组冻结策略作为输入,在学习到的状态空间世界模型中部署每个冻结策略,评估一个共享的目标条件值函数的想象未来,并在下一轮仲裁(策略选择)前执行评分最高的策略,时间短。WMPA假设访问一组冻结的目标条件策略,无需策略再训练或任务特权知识。根据官方OGBench评估协议,涵盖18个基于状态的数据集,涵盖迷宫导航以及立方体、场景和谜题操作,WMPA将宏观平均成功率从每个数据集中选定的最佳策略的44%提升至58%,并在12个数据集上实现统计学显著提升。这些提升包括立方双杀+33个百分点和场景玩法+36个百分点。
StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
StoreBench:一个用于评估和培训自主运营代理的实时商务环境
- Authors: Daksh Raghuvanshi, Ved Vedere, Yifan Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10942
- Pdf link: https://arxiv.org/pdf/2610.10942
- Abstract
Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.
- 中文摘要
强化学习环境现在已成为提升大型语言模型(LLM)能力的主要杠杆,但大多数代理基准仍保持不变:世界仅在代理行动时移动,奖励是终极判决,通行条则任意设定。我们介绍StoreBench,一种实时商业环境,代理在生产级商业后台运营一家中型在线服装店,测试长期规划和经济判断,在不确定性下进行测试。客户全天候订购,供应商重新定价并失败,市场冲击部分甚至无预警到来。代理通过与人类操作员相同的29种商户工具行动,采用窗口式运营预算,模拟时间取决于所采取的操作,因此模型延迟无法影响模拟时间。通过阈值根据脚本锚策略校准,奖励针对一系列奖励黑客进行硬化,每集在动作序列下均以相同方式重放。我们在30至45天的11个场景和完整模拟年中评估了七个前沿大型语言模型,涵盖三个世界种子,推理努力匹配。没有模型平均能匹配脚本化智能分流策略:最佳的DeepSeek-V4-Pro通过49%的任务种子单元,而启发式的97%。使用相同工具和预算的人类专家得分优于所有模型(平均合成0.708对0.700)。在Claude代码框架下的完整模拟一年内,大多数模型表现出显著的性能提升。在GRPO训练后运行中,Qwen3.5-27B仅训练了五个不相交任务,其在未完成评估任务中的平均复合值从0.136提升至0.373。我们发布了五个训练拆分任务示例、十条样本轨迹以及评分和验证工具;为保持基准不受污染,完整环境和评估套件被保留。
Learning Multi-Step Query Rewriting via Corpus Feedback for Conversational Search
学习通过语料库反馈进行多步骤查询重写以实现会话搜索
- Authors: João Coelho, Hong Wang, Jie Yuan, Zhuoer Wang, Samson Koelle, Wei Niu
- Subjects: Subjects:
Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2610.10955
- Pdf link: https://arxiv.org/pdf/2610.10955
- Abstract
Conversational Query Rewriting (CQR) turns a context dependent user turn into a standalone query for a retriever, and most methods do this in a single step from the dialogue history before retrieving once. The rewrite is therefore fixed before any corpus evidence is available to correct its reference resolution or its vocabulary. We recast CQR as a sequential retrieval problem: an agent rewrites the current turn, retrieves, and conditions its next rewrite on the returned passages. The agent acts in a typed space of three rewriting operations, resolving conversational intent into a standalone query, generating lexical reformulations, or synthesizing pseudo-documents for document-to-document matching, together with a stop action that ends the episode. We train the policy with supervised fine-tuning followed by reinforcement learning against a single retrieval-quality reward, using no human rewrite annotations. Across TopiOCQA and QReCC, the agent outperforms several retrieval-aligned baselines, while remaining effective across retrieval backends and generalizing to the CAsT benchmarks without additional training. Further analysis shows that, through retrieval-reward optimization alone, the learned policy develops a behavior of grounding pseudo-documents in passages retrieved by earlier steps, substantially improving retrieval.
- 中文摘要
对话式查询重写(CQR)将上下文依赖用户的切换转换为检索器的独立查询,大多数方法在检索前只需一步从对话历史完成。因此,重写在任何语料库证据可用以纠正其引用解析或词汇之前就已被修复。我们将CQR重新定义为顺序检索问题:代理重写当前回合,检索并对返回的段落进行条件。代理在一个由三个重写操作组成的类型化空间中行动,将会话意图解析为独立查询,生成词汇重构,或合成伪文档以进行文档间匹配,并配合停止动作结束该集。我们通过监督微调训练策略,随后针对单个检索质量奖励进行强化学习,不使用人工重写注释。在 TopiOCQA 和 QReCC 中,该代理表现优于多个检索对齐基线,同时在检索后端保持有效,且无需额外训练即可推广至 CAsT 基准。进一步分析显示,仅通过检索-奖励优化,该策略就能发展出将伪文档基于前一步检索的段落的行为,显著提升检索率。
ATLAS: Adaptive TDA-guided Landscape-Aware Transistor Sizing
ATLAS:自适应TDA导向景观感知晶体管尺寸
- Authors: Youngmin Oh, Jihwan Won, Yuntae Park, Bosun Hwang, Suwan Kim
- Subjects: Subjects:
Hardware Architecture (cs.AR)
- Arxiv link: https://arxiv.org/abs/2610.10985
- Pdf link: https://arxiv.org/pdf/2610.10985
- Abstract
Analog transistor sizing, finding design parameters that simultaneously satisfy multiple performance specifications, is a labor-intensive bottleneck in circuit design. To support analog circuit experts, various automation methods have been proposed, including Bayesian optimization (BO), reinforcement learning (RL), and others. Yet existing methods are oblivious to the topological structure of a feasible design space, which can fragment into disconnected regions due to operating-regime transitions, conflicting specification trade-offs, and nonconvex device physics. This topological blindness causes the optimizer to converge within a single feasible region while missing others that may contain superior designs. To address this limitation, we propose ATLAS. a BO framework utilizing Topological Data Analysis (TDA). At each iteration, a Mapper graph is constructed over a surrogate-predicted feasible region to estimate connected regions, enabling topology-aware exploration from the very first iteration without any observed feasible points. A topological sensitivity score classifies candidates as bridge, frontier, or interior points, injecting a targeted exploration bonus into the acquisition function. Experiments on four analog circuit benchmarks in the GF180 and SKY130 processes demonstrate that \coin finds feasible designs with significantly fewer simulations than baselines, including RL and BO methods. To the best of our knowledge, this is the first work to apply topological data analysis to analog circuit design automation. The official implementation is publicly available on this https URL.
- 中文摘要
模拟晶体管尺寸调整,即寻找同时满足多个性能规格的设计参数,是电路设计中的劳动密集型瓶颈。为支持模拟电路专家,提出了多种自动化方法,包括贝叶斯优化(BO)、强化学习(RL)等。然而,现有方法对可行设计空间的拓扑结构缺乏了解,由于工作区间转变、规范权衡冲突和非凸器件物理,空间可能分裂成不相连的区域。这种拓扑盲点导致优化器收敛在单一可行区域内,而缺少可能包含更优设计的其他区域。为解决这一限制,我们提出了ATLAS。一个利用拓扑数据分析(TDA)的BO框架。每次迭代时,在代理预测的可行区域上构建映射图以估计连通区域,使得从第一迭代起即可进行拓扑感知的探索,无需任何可观测点。拓扑敏感度评分将候选点分类为桥接点、前沿点或内部点,为获取函数注入有针对性的探索加成。在GF180和SKY130工艺中对四个模拟电路基准测试的实验表明,'Pr-Coin's在模拟次数远少于基线的情况下,找到了可行设计,包括强化学习和BO方法。据我们所知,这是首个将拓扑数据分析应用于模拟电路设计自动化的工作。官方实现可在此https URL公开获取。
Measuring and Mitigating Solution Mode Collapse in RLVR
测量与缓解RLVR中解模坍缩
- Authors: Liv G. d'Aliberti, Marwa Abdulhai, Sofiia Druchyna, Peter Henderson, Manoel Horta Ribeiro
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11064
- Pdf link: https://arxiv.org/pdf/2610.11064
- Abstract
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.
- 中文摘要
语言模型(LM)通常可以用多种方式回答同一问题,但带可验证奖励的强化学习(RLVR)对模型生成的正确答案并不介意。无论解答是熟悉答案的第千个副本,还是模型从未生成过的答案,都将获得相同的奖励。然而,让模型在训练过程中保留多个正确解法具有潜在价值。例如,多模式可能为用户提供选择,并提供解决问题策略,从而提升整体模型性能。这里,我们介绍ModeBench,这是一个多解任务基准测试,验证者会返回正确性和发现的模式。然后我们用ModeBench测量RLVR训练后解多样性的变化情况。我们发现RLVR训练后即使准确率保持或提升,概率仍集中于更少的正确模式,而且前沿模型已经高度集中。随后我们介绍了我们的解决方案Re:Max,它在每个发现模式中存储一个经过验证的样本,并均匀训练这些存储模式。因此,一次找到的解被实践的频率与反复找到的解一样多。在三种模型尺度、两个强化学习目标和更难的任务构建中,重放不仅提高了策略成功的频率,也提高了其成功方式的多样性。
DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks
DaCe-DT:通过自适应提示和轨迹修正的数据中心离线多任务强化学习,针对异构任务
- Authors: Xinfei Wang, Shanchen Pang, Chenhao Zhang, Shudong Wang, Wenhao Ji, Haiyuan Gui, Meng Han, Xiaojian Liao
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.11085
- Pdf link: https://arxiv.org/pdf/2610.11085
- Abstract
Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
- 中文摘要
离线多任务强化学习(离线MTRL)高度依赖于预先收集数据的质量和分布。然而,现有方法主要侧重于算法优化,较少强调数据层面的改进以提升学习能力和泛化性能。本文从数据角度揭示了限制离线MTRL性能的三个关键瓶颈:(i)在不同任务复杂性下提示长度的低效利用,(ii)随机抽样提示段的语义无关,(iii)由碎片化且不连续轨迹引发的误导性监督。为应对这些挑战,我们提出了DaCe-DT框架,这是一个稳健的离线MTRL框架,设计上对异构任务复杂性和数据质量不敏感,具备长度门控提示掩蔽(LGPM)、检索增强提示构建(RAPC)和值适应返回校准(VARC)。这些机制共同使DaCe-DT能够实现以数据为中心的提示适应和轨迹优化,从而实现了在异构离线数据和任务中实现多任务泛化和稳定策略学习。Meta-World上的实验结果显示,DaCe-DT持续优于最先进方法,在最优数据集上平均提升11.73%,在次优数据集上提升13.34%,展示了其从不完美数据中稳定学习和提升整体多任务表现的有效性。
FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
重点:从特权状态到带有受控模态切换和表示对齐的RGB-D
- Authors: Filip Grigorov, Kourosh Darvish, Nandita Vijaykumar
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.11119
- Pdf link: https://arxiv.org/pdf/2610.11119
- Abstract
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy. Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap. We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents. Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them. Together, these mechanisms limit RGB-D rollouts when the actor's action distributions from RGB-D and privileged-state latents disagree. As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time. Across five manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 relative to the strongest RGB-D-at-test baseline on each task. When accounting for each method's complete training pipeline, budget-normalized training-success AUC increases from 0.47 to 0.65. On Pick-and-Place, test success rises from 0.47 to 0.86, while AUC increases from 0.12 to 0.61, a 5.0x improvement in learning efficiency over the fixed interaction budget.
- 中文摘要
基于视觉的机器人操作强化学习样本效率低,因为RGB-D观测是高维且噪声大的。仿真中可用的特权状态信息可以加速训练,但在测试时缺乏该信息会导致训练-测试模态差距。我们提出了FOCUS框架,这是一种单阶段PPO框架,它在训练批评者进行特权状态训练的同时,自动调节参与者是从RGB-D还是特权状态潜伏收集滚出数据。调控由两种模态诱导的动作分布之间的KL发散驱动,而表征对齐则鼓励在它们之间保持一致的动作选择。这些机制共同作用,当参与者从RGB-D和特权状态潜伏的动作分布不一致时,限制了RGB-D的展开。随着它们的对齐,RGB-D暴露增加,策略训练转向测试时使用的RGB-D输入。在五个操作任务中,FOCUS将平均测试成功率从0.71提升至0.93,相较于每个任务中最强的RGB-D测试基线。考虑每种方法的完整训练流程后,预算归一化训练成功AUC从0.47提升至0.65。在Pick-and-Place测试中,测试成功率从0.47升至0.86,AUC从0.12提升至0.61,学习效率比固定交互预算提升5.0倍。
When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
当接口发声:数据感知生成式UI驱动主动交互
- Authors: Xiaolong Li, Xiaohan Xu, Jinyang Li, Xinnuo Xu, Ge Qu, Nan Huo, Jack Williams, Reynold Cheng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11123
- Pdf link: https://arxiv.org/pdf/2610.11123
- Abstract
Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
- 中文摘要
如今大多数人与代理的交互仍以文本为基础。自然语言可能带来认知过载、模糊性、信息混乱和复杂任务输入缓慢;短暂生成界面可以呈现结构化信息并引导用户完成任务。我们提出GenUI-Harness,这是一种多代理框架,将用于信息检索和任务执行的工具代理与识别歧义并生成结构化界面前端代码的图形化程序员代理配对。用强化学习训练程序员具有挑战性:交互式UI生成的可验证奖励需要高昂的执行成本,而作为法官的LLM奖励则容易被奖励黑客攻击。我们用Dynamic UX(一个轻量级的动态交互和奖励收集包单一沙盒)解决了第一个挑战,第二个则是Reward Auditor,一种元奖励机制,监控奖励分布并将诊断模式提炼成共享的评分标准和评分规范。我们引入了UI-TAU Bench,这是一个通过生成UI代码实现主动人与客服互动的基准测试,基于10个由公共数据源构建的真实领域数据库,基于Tau-Bench工具使用设置,设有Lite(300项任务)和Full(1000项任务)分段。GenUI-Har Pass@3 ness在Lite上比smolagents平均提升4.48个百分点的 3。使用 GenUI-Harness 训练将 4B 骨干链从9.33%提升到58.00%Pass@3,优于大型前沿模型如 Claude Opus 5(46.67%)。GenUI-Harness在模糊和非歧义查询上也保持了稳健性。在一项比较通信渠道的评审调查中,生成的UI将平均对话轮数从3.4降至1.2。这些结果表明,数据感知型生成接口能够支持有效任务完成,并减少评估数据库支持工作流程中的对话轮次。
Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning
LLMs是否能从情境中的奖励中学习?:重新思考情境内奖励在情境强化学习中的作用
- Authors: Minchan Kwon, Seunghee Koh, Sunghyun Baek, Minsung Bae, Junmo Kim
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11152
- Pdf link: https://arxiv.org/pdf/2610.11152
- Abstract
LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters. This process is often described as in-context reinforcement learning (ICRL). Whether in-context learning (ICL) can actually play the role of RL, however, has not been tested. We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal. Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
- 中文摘要
LLM代理越来越多地通过在推理时积累上下文经验而非更新参数来提升。这一过程通常被称为上下文强化学习(ICRL)。然而,上下文学习(ICL)是否真的能发挥强化学习的作用尚未被测试。我们以最简单的形式——直接ICRL研究这个问题,模型直接基于原始轨迹-奖励对,并询问奖励是否作为学习信号。通过对六个模型四个基准测试的受控实验,我们发现奖励被读取,但其效果很小:翻转、随机化或移除奖励,改进曲线几乎不变,即使在元提示中明确指示模型探索、利用或推理奖励的情况下,这一现象依然成立。轨迹推动改进,但不是通过其语义内容:洗牌或损坏的轨迹与真实轨迹同样有效。这些模式与ICL中已知的模式高度相似,表明直接ICRL更适合被理解为ICL的一个特例,而非推理时间RL。这种重构对代理记忆设计具有启示意义:输入分布和演示等ICL因素可能比奖励塑造和探索等强化学习元素更为重要。
PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
PIVOT:基于困惑度的KD到RL垂直域少数精馏过渡调度
- Authors: Heng Li, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11167
- Pdf link: https://arxiv.org/pdf/2610.11167
- Abstract
Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.
- 中文摘要
垂直域少数样本分类对于小型语言模型来说依然具有挑战性,因为有限的监督使得获得领域特定的决策知识变得困难。策略提纯(OPD)可以通过监督学生生成的推广来提升教师引导的适应性,而基于GRPO的强化学习则可以进一步优化下游预测。然而,现有的知识驱动到强化学习流程通常依赖全局固定的过渡计划,忽视了不同样本在奖励驱动细化前可能需要不同程度的教师引导学习。我们提出了PIVOT(困惑感知情过渡优化),这是一个动态过渡框架,根据教师评估的序列困惑度在OPD和GRPO之间路由样本。PIVOT将低困惑度样本移至GRPO进行奖励驱动的精炼,同时将高困惑度样本保留在OPD下,以便持续领域知识获取。在Banking77和HWU64上的实验显示,在相同热身后优化步骤下,PIVOT持续优于持续OPD和全局同步OPD$\rightarrow$GRPO基线,实现更强的下游表现和更稳定的训练动态。
Higher-Order Action Supervision Makes A Strong Policy Class
高阶行动监督构成强有力的政策类别
- Authors: Peng Cheng, Yunxian Hou, Zhi Zhou, Qian Zhang, Chang Huang, Xianyuan Zhan
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11175
- Pdf link: https://arxiv.org/pdf/2610.11175
- Abstract
Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.
- 中文摘要
现代数据驱动决策方法,如模仿学习(IL)和强化学习(RL),在解决许多复杂任务方面取得了巨大成功。然而,这些方法在机器人和自动驾驶等实际应用中常常存在严重的控制不稳定性和鲁棒性问题,给其实际应用带来了显著挑战。我们认为,这种不稳定性问题主要源于它们仅监督和优化零阶动作(即动作标签)的局限性,未能考虑高阶动作动态和时间一致性。本文表明,同时监督零阶和一阶动作可以显著提升策略的性能和控制鲁棒性。为此,我们引入了一种新颖且优雅的损耗方案,并配有形式理论保证,能够为任何现成策略模型(如确定性、随机策略或流策略)配备高阶动作监督能力,无需任何结构修改。此外,我们提出的方法可作为轻量化即插即用模块,无缝集成多种现有离线强化学习框架。对OGBench和D4RL的广泛评估表明,我们的方法在多种连续控制环境中显著提升性能和鲁棒性。值得注意的是,我们的方法还能增强策略在低数据环境中的分布外(OOD)泛化能力,使其成为解决许多现实控制问题的理想工具。
RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
RAG压力:探讨检索增强生成中证据依赖的极限
- Authors: Shunyuan Zhou, Hao Chen, Tianyu Wang, Goose Lin, Zaiyuan Wang, Haiying Zhao
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11183
- Pdf link: https://arxiv.org/pdf/2610.11183
- Abstract
Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Mean MR follows End $>$ Beginning $>$ Middle under both policies, although individual models do not uniformly follow this ordering. A separate paired audit of 500 questions and two checkpoints supports increased harmful override without establishing a corresponding improvement in beneficial correction. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.
- 中文摘要
遵循检索的证据并不保证事实正确性:误导性证据可能导致模型替换其先前正确给出的答案。标准准确度测量通过将答案替换与已有错误结合来掩盖这种行为。我们引入了RAG-Stress,一种受控诊断协议,用于检验检索增强生成中证据依赖的极限。该协议固定问题和参考答案,编辑一个断言以支持指定的错误答案,并在证据文本中交叉两个来源优先级策略,覆盖答案跨度三个位置。我们测量每个模型中未检索正确回答问题子集的误导率(MR),同时对完整评估集进行清晰准确率。我们评估了15个系统,涵盖API模型、开放模型和通过TriviaQA-RC、HotpotQA和SearchQA进行强化学习训练的搜索代理,并额外进行了英文和中文的MedQA评估。优先排序文档的指令持续产生更高的MR,而允许依赖先验知识的指令则更准确。在模型和位置上平均,三个QA数据集的差距在10.9%到13.5个百分点之间。平均MR在两种策略下都遵循“结束$>$ 起始$>$ 中间”,尽管单个模型并不统一遵循此排序。一项单独的500个问题和两个检查点的配对审计支持了增加有害覆盖,但未确立相应的有益纠正改进。这些发现区分了证据遵循性与事实可靠性,并促使评估检索证据是否保留、替代或纠正模型的答案。
Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
为什么策略提炼有时失败:学习信号的消失
- Authors: Lei Zhao, Qichao Zhao, Bowen Zuo, Qishi Zhan
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11247
- Pdf link: https://arxiv.org/pdf/2610.11247
- Abstract
On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $>0.98$ across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at this https URL.
- 中文摘要
策略提纯(OPD)实现了语言模型间有效的能力转移,但其失败机制尚未完全明瞭。在代码生成和数学推理中,大规模教师的OPD表现出早期损失平台期,200次更新后平均最终损失减少25.1%,而自我强化学习教师则减少了96.2%,这也是通过对初始学生进行进一步强化学习(RL)训练获得的。为理解这一差异,我们将OPD分析为一个理想化的连续时间动力学系统,处于小学习速率极限。我们的训练日志诊断将这些平台期与基于梯度的学习信号代理的早期下降联系起来,同时显著损失依然存在;这些测量未能确定为何底层梯度会减弱。我们还进一步证明了在正规条件下,与初始学生足够接近的教师具有局部恢复保证,为自强化学习教师在实验中的成功提供了条件解释。在有损失和无损失平台的运行中,我们观察到学生在OPD前后表征之间的相对参数变化较小(0.025-0.098%),且高度相似(跨层线性CKA>约0.98美元)。这些观察表明有限的表征适应可能促成学习信号崩溃,这一假说仍有待验证。代码可在此 https 网址获取。
How to post-train on a surrogate: Envelope sampling mitigates reward hacking
如何对代理进行后期训练:信封采样有助于缓解奖励黑客行为
- Authors: Sanjit Dandapanthula, Shuvom Sadhuka, Samir Khan, Michael Oberst, Aaditya Ramdas, Alexandra Chouldechova
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Methodology (stat.ME)
- Arxiv link: https://arxiv.org/abs/2610.11281
- Pdf link: https://arxiv.org/pdf/2610.11281
- Abstract
Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.
- 中文摘要
大型语言模型(LLMs)通常会对LLM评判和其他廉价替代者进行后训练,因为真正的奖励,如人类偏好,在大规模查询中成本过高。这种做法常导致奖励黑客行为,即对校准错误的替代者进行强化学习,导致不良副作用。在本研究中,我们研究了一个环境:少量$n$的模型输出被注释为基于事实的标签(例如专家评审),用于重新校准LLM评判者,然后再进行优化。以往的评判重新校准方法成本高昂或启发式,已知当代理在罕见输出集合上校准错误时,策略抽样会失败。在本研究中,我们提出了包络采样,这是一种理论基础的评委重新校准方法,旨在最小化训练后模型遗憾的上限,假设人类奖励和重新校准后的奖励在评委周围形成一个$L^2$的球体。我们提出了通过拒绝或对修改后奖励微调从包络线中采样的实用算法,临床笔记生成和受控谄媚任务的实验表明,重新校准包络样本能减轻奖励黑客,而重新校准基模型样本则无效。
SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
SynCo:通过多智能体强化学习实现自演化大型语言模型的数据综合协同训练
- Authors: Wei Yang, Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason, Xuezhe Ma, Yue Zhao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11345
- Pdf link: https://arxiv.org/pdf/2610.11345
- Abstract
Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.
- 中文摘要
自我进化的LLM代理承诺通过持续互动和学习实现自主提升,减少对人工管理的依赖。实现这一承诺不仅需要更新代理,还需随着能力变化演进其训练体验。然而,大多数现有流水线依赖静态数据集或单独更新的综合模型,导致原本有用的任务变得琐碎,而过于困难的任务则缺乏信息。代理能力与训练经验之间日益加剧的不匹配限制了持续的自我提升。为解决这一问题,我们提出了SynCo,这是一个基于多智能体强化学习的智能体数据综合共训练框架,用于自我演化LLMs。SynCo联合优化了两个独立参数化的智能体:一个从理性者能力进化状态构建训练任务的合成器,以及从所得经验中学习的推理器。每个综合任务会诱导多次推理者展开,其结果为双方智能体提供互补奖励。正确性反馈提升推理者,而任务质量、答案可靠性和基于结果的可教学性则指导合成器。这些更新反馈回后续综合轮次,使任务解决策略与训练分布同步演进。涵盖八个数学推理基准的广泛实验表明,SynCo在众多现有合成数据方法和受控基线中表现显著优于,实现了最强的整体性能,同时大部分收益来自先前未解决的问题。
RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
RL-ARC:通过推理引导不确定性校准大型推理模型
- Authors: Gukhyeon Lee, SangKeun Lee
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.11352
- Pdf link: https://arxiv.org/pdf/2610.11352
- Abstract
Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-aware training methods for LMs, which incorporate objectives for uncertainty estimation into training, improve calibration but still exhibit overconfidence under distribution shift, while sacrificing reasoning performance. To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence. Specifically, RL-ARC leverages reasoning confidence as an auxiliary signal for calibrating answer confidence, applying it as reasoning-guided regularization for correct cases and as an overconfidence penalty for incorrect cases. Comprehensive results across ID and OOD settings show that, beyond improving calibration, RL-ARC enables reasoning models to adaptively estimate confidence based on the given question without substantially sacrificing reasoning performance, thereby highlighting the importance of reasoning confidence for training reliable reasoning models.
- 中文摘要
语言模型(LM)通常通过可验证奖励强化学习(RLVR)训练以增强推理能力。然而,由于RLVR未明确考虑训练中的校准,可能导致严重的校准退化,包括过度自信。近期针对LM的校准感知训练方法,将不确定性估计目标纳入训练中,虽然校准效果有所提升,但在分布偏移下仍表现出过度自信,同时牺牲推理表现。为此,我们提出了RL-ARC这一校准感知训练框架,结合推理信心和答案信心。具体来说,RL-ARC将推理信心作为辅助信号校准答案置信,作为推理引导的正则化(正确案例)和错误案例的过度自信惩罚。跨教学设计和外勤教学环境的综合结果表明,除了提升校准能力外,RL-ARC还使推理模型能够基于给定问题自适应估计置信度,同时不显著牺牲推理性能,从而凸显推理置信度在训练可靠推理模型中的重要性。
From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning
从提示到词汇:不断演变的功能性词汇助力LLM持续学习
- Authors: Fengyuan Liu, Yue Wang, Hangxi Guo, Fengyuan Liu, Chenxu Wu, Yanguang Liu, Mengnan Du
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11373
- Pdf link: https://arxiv.org/pdf/2610.11373
- Abstract
Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids costly parameter updates while achieving competitive or even superior performance to reinforcement learning methods such as GRPO on individual knowledge-intensive and reasoning tasks. This raises a natural question: \textit{Can prompt optimization, as an efficient adaptation approach, be directly applied to continual learning?} Our analysis shows that, under sequential task adaptation, it suffers from catastrophic forgetting, while optimized prompts accumulate rules that overfit to local task distributions. To address these limitations, we propose \emph{Evolving Functional REpertoires} (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones. On a three-task continual-learning stream, EFRE achieves a final average performance 7.50 percentage points higher than GRPO. Moreover, after adaptation to the Bio task, its performance on FinQA decreases by only 1.56 percentage points, compared with 25.10 percentage points for the base prompt optimization method. We further instantiate EFRE in a minimal agent system and observe consistent improvements across different backbone models. Overall, these results demonstrate EFRE's strong performance in continual learning for large language models and highlight its substantial potential for continual learning in advanced agent systems.
- 中文摘要
对于大型语言模型来说,持续学习依然具有挑战性,因为它们必须让模型能够在不削弱现有能力的前提下获得新技能和知识。现有方法通常通过精心设计模型参数的更新方式来应对这一挑战。相比之下,提示优化避免了昂贵的参数更新,同时在单个知识密集型和推理任务上,能够与强化学习方法如GRPO竞争甚至更优。这自然引出了一个问题:\textit{提示优化作为一种高效的适应方法,能否直接应用于持续学习?}我们的分析显示,在顺序任务适应下,提示存在灾难性遗忘现象,而优化提示则积累了与局部任务分布过拟合的规则。为解决这些限制,我们提出了 \emph{演化功能库}(EFRE),它用一系列随着新任务进化而演变的功能库取代单一提示:兼容的更新优化现有功能,冲突的更新触发新功能的出现。在三任务持续学习流中,EFRE 最终平均表现比 GRPO 高出 7.50 个百分点。此外,适应生物任务后,EFRE 在 FinQA 上的表现仅下降 1.56 个百分点,而基础提示优化方法下降了 25.10 个百分点。我们还进一步在最小代理系统中实例化 EFRE,观察到不同骨干模型间的持续改进。总体而言,这些结果展示了 EFRE 在大型语言模型持续学习中的强劲表现,凸显了其在高级代理系统中持续学习的巨大潜力。
Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
环境反馈建模的重要性:重新思考能动性回顾自我蒸馏中的反馈处理
- Authors: Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, Mengnan Du
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11384
- Pdf link: https://arxiv.org/pdf/2610.11384
- Abstract
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $\tau$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
- 中文摘要
强化学习常用于在互动环境中训练语言代理,但当奖励不可得时,无法直接应用。近期方法将环境反馈作为事后自提纯的特权上下文,但我们的分析表明,仅仅让教师条件反射反馈不足以实现反馈,促使我们重新思考环境反馈在能动自我蒸馏中的应用方式。鉴于环境反馈包含丰富的监督,用于建模环境如何响应代理行为,我们引入了\textit{agentic SElf-distillation with Environmental Feedback Modeling}(SELF),这是一个结合环境反馈建模和事后自提纯的框架。SELF学习预测环境反应,同时将反馈条件自教师的指导提炼进策略。我们的分析揭示了一个相互强化的机制:环境反馈建模增强了事后诸葛亮的监督和政策学习,而自我蒸馏则增强了模型对环境反馈的建模能力。通过Qwen3-8B,SELF在$\tau$-bench成功率上分别优于SDPO和GRPO6.4个百分点和4.1个百分点,在AppWorld任务目标完成率上分别高出10.71个百分点和3.57个百分点。这些结果表明,SELF在代理自我蒸馏中更有效地利用环境反馈,提升了代理的能力。
DRL-Based AoI Minimization for RSMA in Finite-Blocklength MU-MISO
基于DRL的有限区块长度MU-MISO中RSMA的AOI最小化
- Authors: Enes Kaya, Mehmet Alp Demircioğlu, Elif Tugce Ceran, Melda Yuksel
- Subjects: Subjects:
Information Theory (cs.IT)
- Arxiv link: https://arxiv.org/abs/2610.11396
- Pdf link: https://arxiv.org/pdf/2610.11396
- Abstract
This letter investigates age-of-information (AoI) minimization in multi-user wireless networks operating in the finite-blocklength (FBL) regime, which is critical for low-latency transmission of short state-update packets. While rate-splitting multiple access (RSMA) provides a powerful and flexible framework for interference management in multi-user FBL systems, the joint optimization of its parameters, such as precoding vectors, power allocation, and rate-splitting ratios, to guarantee information freshness results in analytically intractable complexity. To address this challenge, we propose an actor--critic deep reinforcement learning (DRL) framework to learn dynamic resource-allocation policies in multi-user multiple-input single-output (MU-MISO) broadcast channels. Simulation results show that the proposed RSMA-RL framework achieves consistently lower AoI than the state-of-the-art benchmarks, with substantial gains observed at low signal-to-noise ratio (SNR) and short blocklengths, while matching benchmark performance at high SNR with significantly lower online complexity via a single neural-network forward pass at execution.
- 中文摘要
本信探讨了在有限区块长度(FBL)模式下运行的多用户无线网络中的信息年龄(AoI)最小化,这对于低延迟传输短状态更新数据包至关重要。虽然速率分割多址(RSMA)为多用户FBL系统中的干扰管理提供了强大且灵活的框架,但其参数(如预编码向量、功耗分配和速率分配比)的联合优化,以确保信息的新鲜性,导致分析上难以处理的复杂性。为应对这一挑战,我们提出了一个actor-critic深度强化学习(DRL)框架,用于学习多用户多输入单输出(MU-MISO)广播通道中的动态资源分配策略。模拟结果显示,所提出的RSMA-RL框架在持续低于最先进基准测试的AoI上,在低信噪比(SNR)和短区块长度下显著提升,同时通过单次神经网络前向执行实现,在高SNR下以显著较低的在线复杂度与基准性能相当。
SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
SAIL:通过科学感知循环实现科学代理智能
- Authors: SAIL Model Team: Boyuan Sun, Bryan Dai, Che Liu, Chi Liu, Derek Li, Hongming Piao, Mengzhuo Chen, Xidong Wang, Yan Shu, Yinda Chen, Ziyang Zeng
- Subjects: Subjects:
Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
- Arxiv link: https://arxiv.org/abs/2610.11451
- Pdf link: https://arxiv.org/pdf/2610.11451
- Abstract
We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.
- 中文摘要
我们引入了SAIL,这是一个开放模型,拥有35B总参数和3B活跃参数,用于文献研究、科学编码和多步研究工作流程。SAIL通过科学意识改进循环开发:基于前沿AI模型的代理分析任务失败,构建针对能力缺口的训练任务。诊断涵盖文献任务中的检索与证据选择、编码中的科学假设与推理,以及长期调查中的规划与修订。代理利用纸质集合和科学代码库构建问题、交互轨迹和可执行任务,配合所需环境和工具。我们在多个开发周期重复这一循环,并通过监督微调、专业培训、多教师政策提炼和代理强化学习来训练SAIL。SAIL在科学研究任务中以远少于领先开放权重模型的参数实现竞争性能。
Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
Fed-GRPO:奖励信号驱动联邦集团相对政策优化
- Authors: Pengxin Guo, Shuang Zeng, Zonggen Li, Weiying Zheng, Mengting Liu, Liangqiong Qu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11502
- Pdf link: https://arxiv.org/pdf/2610.11502
- Abstract
Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by $32\times$ and supporting up to $621\times$ compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at this https URL.
- 中文摘要
大型语言模型(LLM)在通过强化学习(RL)微调后,尤其是通过群体相对策略优化(GRPO)展现出强大的推理能力。然而,现有GRPO方法假设训练数据是集中访问,但由于隐私或监管限制,这种方式在实际中可能无法实现。为此,我们提出了Fed-GRPO,一种联邦GRPO训练框架,通过实现协作推理训练而不共享原始数据,利用GRPO训练中自然产生的奖励统计数据作为零成本信号,指导聚合、局部训练和沟通。Fed-GRPO包含三种奖励信号驱动机制:(i) \emph{信号加权聚合},按奖励标准差加权客户,优先考虑学习信号更强的客户;(ii) \emph{全局奖励校准},根据本地-全局奖励差距重新加权每个提示目标,引导每个客户端针对其相对弱点;以及 (iii) \emph{自适应稀疏通信},根据每个客户端更新的信息量分配带宽。数学推理任务的大量实验表明,Fed-GRPO 在所有联邦方法中表现最佳,明显优于 FedAvg,且接近集中式训练性能,同时无损地减少 32 倍倍的通信,并在有限带宽预算下支持高达 621 倍的压缩,且仅有优雅的精度下降。我们的代码可在此 https URL 访问。
DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors
DAMP:通过去噪信念学习和对抗运动先验实现类人运动
- Authors: Puying Shen, Wenhao Cui, Huaxing Huang, Bangyu Qin, Shengtao Li, Ziyang Dong, Guoteng Zhang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.11505
- Pdf link: https://arxiv.org/pdf/2610.11505
- Abstract
Humanoid robots possess the structural capability to traverse complex terrains. However, achieving stable t raversal without relying on perceived information remains challenging, particularly in complex environments. This paper introduces DAMP, a reinforcement learning framework aimed at achieving robust and naturalistic humanoid locomotion over challenging terrains, with the assumption that no perceived information is available. The framework leverages recurrent neural networks to capture temporal dependencies and implicitly infer privileged and other task-relevant latent information. By aligning the learned representations with the task objective, the method enables robust and goal-consistent policy learning. This end-to-end framework achieves transfer learning from simulation to real-world environments, demonstrating the proposed method's robustness and generalization capabilities. The video of the real-world demonstration can be found at the following link: this https URL.
- 中文摘要
类人机器人具备穿越复杂地形的结构能力。然而,在不依赖感知信息的情况下实现稳定的t形态转变仍然具有挑战性,尤其是在复杂环境中。本文介绍了DAMP,这是一种强化学习框架,旨在实现在挑战地形上稳健自然的人形移动,假设没有感知信息可用。该框架利用循环神经网络捕捉时间依赖关系,并隐式推断特权及其他任务相关潜在信息。通过将学习到的表征与任务目标对齐,该方法实现了稳健且目标一致的策略学习。该端到端框架实现了从模拟到现实环境的迁移学习,展示了所提方法的稳健性和泛化能力。真实演示视频可在以下链接观看:此 https URL。
Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
残余优势:学生-亲属教师强化学习指导,并提供可验证的奖励
- Authors: Xiaobing Chen, Zhiqi Pang
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11519
- Pdf link: https://arxiv.org/pdf/2610.11519
- Abstract
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
- 中文摘要
带可验证奖励的强化学习(RLVR)和策略提纯(OPD)已成为训练后推理模型的两种主要范式。RLVR为每个响应赋予单一的结果标签,使其中的步骤不计分。OPD在学生访问前缀处提供代币级指导,但其点级信号并未直接反映词汇中教师与学生的分歧模式。密集且无界的对数比监督可以放大教师的影响力,但当学生的解答路径偏离教师时,强求解器未必是合适的指导。我们提出残差优势(\RA{}),将教师-学生概率残差视为有界一步奖励,在学生策略下减去相应状态值形成标准优势,并将结果置于每个响应中,然后再加入验证者优势。指导项在每个回答中均值为零,因此验证者优势仍为回答的平均标签,教师仅在其中的步骤间重新分配学分。\CoRA{} 进一步更新教师 LoRA,并对同一评分学生批次的验证者优势进行更新,并在下一迭代的剩余中使用更新教师,调整指导以适应学生的尝试。对于 Qwen3-1.7B 基础和 Qwen3-4B-基础学生和一位 Qwen3-8B 教师,\RA{} 与 GRPO 或 REINFORCE++ 结合,在三个数学基准测试的 24 次比较中,改进了底层序列优势算法,宏观Avg@8提升了 1.7--3.6 分,Pass@8 提高了 3.9-6.3 分。这两种组合都超过了仅教师的OPD,\CoRA{}又增加了1.0-1.5 Avg@8点。
When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning
何时干预?联邦强化学习中的状态感知稀疏操作
- Authors: Shutong Zheng, Sijia Chen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.11523
- Pdf link: https://arxiv.org/pdf/2610.11523
- Abstract
Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-level intervention timing remains largely implicit. In sequential decision making, however, where an intervention is applied can alter subsequent trajectories and learning signals. Through controlled experiments, we find that changing the selected trajectory states materially alters attack efficacy even when the malicious-update construction is fixed. We therefore identify when as a distinct attack dimension and introduce the Viability-constrained Behavioral Steering Attack (V-BSA), which uses local policy uncertainty to select sparse intervention states and applies envelope-constrained behavioral steering. Across discrete-action benchmarks, V-BSA achieves substantial degradation against robust aggregators and ensemble defenses with only a fraction of the interventions used by dense poisoning, while revealing task- and aggregation-dependent boundaries. Overall, our results highlight intervention timing as a distinct dimension of sequential robustness in FRL. The code is available at this https URL
- 中文摘要
联邦强化学习(FRL)使分布式智能体能够协同训练决策策略,但其去中心化训练过程也使全球策略学习暴露于复杂操控之下。现有的毒化攻击主要关注如何构建恶意更新,而轨迹级干预时机则大多隐含。然而,在顺序决策中,干预的应用可以改变后续轨迹和学习信号。通过受控实验,我们发现即使恶意更新结构被固定,改变所选轨迹状态也会实质性改变攻击效能。因此,我们将“当”识别为一个独立的攻击维度,并引入可行性约束的行为引导攻击(V-BSA),利用局部策略不确定性选择稀疏干预状态,并应用包络约束的行为引导。在离散行动基准测试中,V-BSA仅以密集中毒干预的一小部分,在对稳健聚合和集合防御方面实现了显著降解,同时揭示了任务和聚合相关的边界。总体而言,我们的结果强调干预时机作为FRL中顺序鲁棒性的一个独特维度。代码可在此 https URL 获取
Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating
学习协调进化搜索:动态DE-CMA-ES优化与结构模型更新中的进阶感知深度强化学习
- Authors: Lechen Li (1 and 2), Rongye Shi (3), Wanhuan Zhou (1) ((1) State Key Laboratory of Internet of Things for Smart City, University of Macau, Macau 519000, China, (2) College of Water Conservancy and Hydropower Engineering, Hohai University, Nanjing 210098, China, (3) School of Artificial Intelligence, Beihang University, Beijing 100191, China)
- Subjects: Subjects:
Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11546
- Pdf link: https://arxiv.org/pdf/2610.11546
- Abstract
Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.
- 中文摘要
解决高维结构模型更新问题需要一个能够导航复杂且非凸且参数相关的算法。现有的混合进化算法通常依赖静态架构或固定切换规则,导致搜索阶段不相连。为此,本研究提出了一种深度强化学习控制的动态DE-CMAES编排(DRL-DCO)算法,其中基于深度确定性策略梯度(DDPG)的actor-critic代理作为单一统一系统持续治理进化过程,而非机械性算法的串接。在进度感知状态表示和多样性知情奖励的引导下,智能体能够在差分向量探索(DE)和协方差引导的CMA-ES利用之间流畅地重新分配计算资源,同时共同调节种群规模、精英保护和重启机制以逃离局部最优。这使得DRL-DCO能够自主地在探索主导、利用主导和混合策略模式间跨代切换。训练阶段之后,受训练的参与者可以处于无监督的推理模式,内化策略自主地通过前向推断从观察到的搜索状态调控DE和CMA-ES,无需批评评估或权重更新,从而实现更快部署同时保持完整效能。经过高维单目标优化基准测试和IASC-ASCE结构健康监测基准测试验证,DRL-DCO相比最先进的自适应和混合进化算法,以及单操作员DRL控制基线,实现了更优越的收敛精度和鲁棒性。
Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
即时战略游戏中带有强盗策略选择的受限指令条件强化学习
- Authors: Nick Leenders, Roy Lindelauf, Joost van Oijen, Boris Cule
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11663
- Pdf link: https://arxiv.org/pdf/2610.11663
- Abstract
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
- 中文摘要
深度强化学习代理在实时战略游戏中表现优异,但对训练分布之外的对手可能脆弱。将战略指令选择与学习到的单位控制分离,允许在重复使用相同执行策略的同时,为不同对手选择不同策略。这需要执行者能够遵循不同命令并具备可测量的评估标准。我们引入了一种受限命令条件的近端策略优化(PPO)策略,即执行者,适用于MicroRTS,即实时战略环境。离散指令在多个环境步骤中指定经济、军队组成、军事态势和工人政策的战略目标和行为要求;执行者决定用于实现这些目标的单位级行动。汤普森采样的强盗作为战略家,从基于游戏内观察而非对手身份构建的对手战略估计中选择指令元组。与采用相同架构、预算、课程和自玩联赛训练的平坦PPO基线进行受控比较时,战略执行者系统在训练地图上对四个最强对手中的三个获胜率显著更高,包括两个最强的对手(0.55比0.97和0.01比0.34),与其他对手无显著差异。
Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
自回归检索器:提升从项目反馈中获取查询的理解,实现通用多模态检索
- Authors: Jianfei Zhao, Yifan Wang, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan, Yang Luo, Boyuan Pan, Xu Kai, Yao Hu
- Subjects: Subjects:
Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.11666
- Pdf link: https://arxiv.org/pdf/2610.11666
- Abstract
Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.
- 中文摘要
通用多模态检索通常对查询编码一次,并通过嵌入相似性对独立索引的项目进行排名。该设计支持高效搜索,但即使检索到的条目有助于澄清信息需求,查询表示方式仍保持不变。我们介绍了自回归检索器(ARR),这是一种多模态检索模型,既学习选择信息性条目,又利用其内容细化后续检索。ARR在检索条目和更新查询嵌入之间交替进行,然后用最终嵌入对集合进行排名。监督微调通过逐步对比监督教编码器使用反馈。强化学习将反馈项视为动作,并利用相关项的最终倒数排名优化其选择。查询端适配器支持针对固定条目索引的优化。ARR在域内和零样本基准测试中均表现出强劲的检索性能,平均优于比较基线。进一步分析显示,反馈在推断时提升检索能力,且反馈训练还能改善初始查询嵌入,在观察到任何项目之前。
Can Jev be Your Q or Policy in Reinforcement Learning?
Jev可以作为你在强化学习中的Q或政策吗?
- Authors: Yi Ma, Tianpei Yang, Yaodong Yang, Weixun Wang, Hongyao Tang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11692
- Pdf link: https://arxiv.org/pdf/2610.11692
- Abstract
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements that the objects of an RL system place on the answers they consume, and establish that Jev can fulfill all of them except the cardinal use of a value function. The remaining objects form positions that admit several roles each. We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance. Across nine MiniGrid tasks and three Atari games, training with Jev outperforms a standard RL learner, including where the learner makes no progress alone, while the model itself remains untrained. To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.
- 中文摘要
基础模型为强化学习(RL)提供了先验,弥补了其长期存在的样本效率和迁移弱点,但其逐个代币生成使查询顺序化且成本高昂。Jev是最近发布的决策模型,不生成任何数据,仅在一次前向传递中返回校准后的输入答案。现有研究研究将基础模型作为训练模型或提示生成器,而Jev不属于这两类,迄今仅作为单一领域中的黑箱。因此,该模型在强化学习环境中自主决策的能力,以及如何提升强化学习作为训练组成部分,仍未被探讨。为此,本文首先考察强化学习系统对象对其获取答案的要求,并确认Jev除了价值函数的基本使用外,可以满足所有要求。其余对象形成各位置,每个位置可承担多个角色。随后,我们构建了算法,将Jev作为参考策略、探索裁判和重放评分器三个位置,以提升样本效率、探索和学习表现。在九个MiniGrid任务和三款Atari游戏中,Jev训练优于标准强化学习者,包括在学习者单独无进展、模型本身未训练的情况下。据我们所知,我们首次展示了Jev在强化学习过程中的应用,并建立了冻结决策模型作为强化学习的可用组成部分,进一步探讨Jev及其他高级决策模型如何提升强化学习。
A 3D Characterization Framework for Intelligent Sequential Decision Making
一个用于智能顺序决策的三维特征框架
- Authors: Sadig Gojayev, Carolina Fortuna
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.11696
- Pdf link: https://arxiv.org/pdf/2610.11696
- Abstract
Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.
- 中文摘要
谜题被广泛用于评估人工智能(AI)系统在顺序决策中的推理能力,但来自不同范式的方法很少在统一条件下进行比较。为弥补这一空白,我们引入了一个三维特征框架,使人工智能方法分析者能够通过1)将其投射到马尔可夫决策过程(MDP)顺序决策形式主义,2)通过人类对设计的先行排序赋予自主性程度,3)技能和计算成本。利用该框架,我们分析了基于代表性图、强化学习和基于大型语言模型(LLM)的方法在设计选择和性能特性上的差异,分别由Neurosolver、前向-后向强化学习(FBRL)和自动思考搜索(AutoToS)实现,其中包括双代理的思维搜索扩展(DA-ToS)。分析依赖于河内塔谜题,该谜题提供了受控基准,具有明确定义的规则和可扩展的复杂度,从而实现在不断增加的问题规模间的一致比较。三维表征显示,基于LLM的方法由于其动作空间设计受限较弱,将复杂性从架构转向推理时间验证,导致内存和运行时间成本远高于Neurosolver和FBRL。
What is the goal of unsupervised machine learning?
无监督机器学习的目标是什么?
- Authors: Aapo Hyvärinen
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.11697
- Pdf link: https://arxiv.org/pdf/2610.11697
- Abstract
Unsupervised learning is one of the main branches of machine learning. Here I argue that unlike the other branches of machine learning (supervised and reinforcement learning), unsupervised learning is a rather heterogenous field that can serve several different goals. It seems futile to try to define one single goal for unsupervised learning. I identify four different goals for unsupervised learning: 1) Estimating the distribution, 2) Generating new data points, 3) Extracting features for downstream tasks, and 4) Understanding the data.
- 中文摘要
无监督学习是机器学习的主要分支之一。在这里我认为,与机器学习的其他分支(监督学习和强化学习)不同,无监督学习是一个相当异质的领域,可以服务于多个不同的目标。试图为无监督学习定义一个单一目标似乎是徒劳的。我为无监督学习确定了四个不同的目标:1)估计分布,2)生成新数据点,3)为后续任务提取特征,4)理解数据。
Distributed Constrained Resource Management in 6G Networks: A Scalable Hybrid Model-Learning Framework
6G网络中的分布式受限资源管理:一个可扩展的混合模型学习框架
- Authors: Thang X. Vu, Cuong Le, Symeon Chatzinotas, Bjorn Ottersten
- Subjects: Subjects:
Information Theory (cs.IT)
- Arxiv link: https://arxiv.org/abs/2610.11739
- Pdf link: https://arxiv.org/pdf/2610.11739
- Abstract
Radio Resource Management (RRM) is a fundamental challenge in 6G wireless networks, particularly under dynamic user demands, inter-cell interference, and heterogeneous QoS constraints. Centralized optimization solutions are often infeasible in practice due to the lack of system-wide state information, dynamic conditions, and signaling delays, making distributed learning-based approaches attractive. However, conventional multi-agent reinforcement learning (MARL) struggles with scalability and constraint satisfaction in such highly dynamic environments. We introduce a scalable two-phase hybrid learning framework to address the RRM challenges where orthogonal-frequency division multiplexing (OFDM) domain knowledge is explicitly incorporated into the MARL pipeline. In our proposed two-phase learning framework, the first phase allocates a minimum resource to satisfy users' QoS requirements based on channel statistics, avoiding the inefficiencies of training policies under hard QoS constraints. Subsequently, a multiagent system is employed in the second phase to optimally allocate the remaining resources for system throughput maximization. By exploiting the structural interference information, we propose a scalable MARL algorithm which decomposes the original learning problem into smaller subproblems that can be handled independently, thereby avoiding exponential growth of the action space without compromising performance. Extensive simulations in realistic scenarios with 50MHz bandwidth and different numerologies show that our method significantly outperforms existing optimization and learning baselines, offering up to 40% throughput improvement, and 100% constraint satisfaction with minimal resource usage.
- 中文摘要
无线资源管理(RRM)是6G无线网络中的一个根本性挑战,尤其是在动态用户需求、小区间干扰和异构服务质量约束下。由于缺乏系统范围的状态信息、动态条件和信令延迟,集中式优化方案在实际中往往不可行,这使得基于分布式学习的方法更具吸引力。然而,传统的多智能体强化学习(MARL)在如此高度动态的环境中存在扩展性和约束满足的困难。我们引入了一个可扩展的两阶段混合学习框架,以应对将正交频率分割复用(OFDM)领域知识明确纳入MARL流水线的RRM挑战。在我们提出的两阶段学习框架中,第一阶段根据信道统计分配最低资源以满足用户的服务质量需求,避免了在严格的QoS约束下训练策略的低效。随后,第二阶段采用多智能体系统,以优化剩余资源以实现系统吞吐量最大化。通过利用结构干扰信息,我们提出了一种可扩展的MARL算法,将原始学习问题分解为可独立处理的更小子问题,从而避免动作空间的指数增长而不牺牲性能。在50MHz带宽和不同数字学的现实场景中进行大量模拟表明,我们的方法显著优于现有的优化和学习基线,在最小资源使用下实现最高40%的吞吐量提升和100%的约束满足。
WAND: Learning Robust Navigation under Complex Wind Disturbances and Dense Obstacles for Quadrotors
WAND:在复杂风扰和密集障碍条件下学习四旋翼机的稳健导航
- Authors: Zhonghan Tang, Chenhui Li, Shuai Liang, Zhongrui You, Jianan Li, Bin Zhao, Zhigang Wang, Xuelong Li
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.11809
- Pdf link: https://arxiv.org/pdf/2610.11809
- Abstract
Robust navigation in cluttered environments remains a fundamental challenge for quadrotors, particularly when strong wind disturbances arise, which perturb vehicle dynamics, limit control authority, and substantially increase collision risk. Existing learning-based navigation policies typically rely on obstacle perception and proprioceptive observations, requiring the policy to infer time-varying disturbance effects implicitly and thereby limiting robustness under partial observability. This paper proposes WAND (Wind-Aware Navigation with Disturbance Estimation), a reinforcement learning framework for navigation under time-varying wind disturbances in dense obstacle fields. Specifically, WAND estimates wind-induced disturbance acceleration from historical proprioceptive states using a Temporal Convolutional Network (TCN). This estimation is integrated into the policy via a zero-initialized residual module, \emph{WindAdapter}, while simultaneously providing feedforward compensation for low-level control. The dual use of the estimate couples disturbance-conditioned navigation with feedforward disturbance rejection. Across 12 wind-disturbed simulation settings, WAND improved the observed success rate by 8.3 percentage points on average relative to feedforward compensation alone. Controlled opposite-crosswind experiments further showed wind-direction-dependent trajectory adaptation. In indoor fan-induced flight tests, WAND succeeded in 18 of 20 trials, demonstrating the feasibility of real-time onboard navigation.
- 中文摘要
在杂乱环境中实现稳健导航仍是四旋翼机面临的根本挑战,尤其是在强风扰动时,这些扰动扰动了车辆动力学,限制了控制权,并显著增加了碰撞风险。现有基于学习的导航策略通常依赖于障碍物感知和本体感觉观察,要求策略隐式推断时间变化的扰动效应,从而限制部分可观测性下的鲁棒性。本文提出了WAND(风感知导航与扰动估计),这是一种用于在密集障碍场中时间变化风扰动下导航的强化学习框架。具体来说,WAND利用时间卷积网络(TCN)估算历史本体感觉状态中风引起的扰动加速度。该估计通过零初始化残差模块\emph{WindAdapter}集成到策略中,同时为低水平控制提供前馈补偿。该估计的双重用途将扰动条件导航与前馈干扰抑制结合起来。在12种风扰模拟设置中,WAND相较单前馈补偿平均提高了8.3个百分点的观测成功率。受控的逆侧风实验进一步显示了风向相关的轨迹适应性。在室内风扇诱导飞行测试中,WAND在20项试验中成功了18项,证明了实时机载导航的可行性。
ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration
ConventionPlay:能力有限的培训,促进强有力的临时协作
- Authors: Abhishek Sriraman, Eleni Vasilaki, Robert Loftin
- Subjects: Subjects:
Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2610.11842
- Pdf link: https://arxiv.org/pdf/2610.11842
- Abstract
Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.
- 中文摘要
临时协作通常要求代理识别并遵守合作任务中的某种共享约定。现有关于临时协作强化学习(RL)的研究侧重于训练能够适应其伙伴建立的约定的智能体。这些方法忽视了这样一种可能性:虽然有些伙伴可能只遵循单一固定约定,而其他则可能能够适应多个约定。这里我们介绍了ConventionPlay,这是一种基于强化学习的方法,教导主体通过训练一组在不同惯例中表现出不同适应性程度的伙伴来发现伙伴的最优约定。其中一些伙伴遵循单一固定约定,而另一些则能够适应该任务中可能约定的子集。支持有限约定子集的伙伴存在,迫使受训练于该群体的代理积极探究伙伴的能力,并引导其伙伴朝着他们能够遵循的最有效联合策略前进。我们的实验结果表明,通过ConventionPlay训练的代理在测试群体中表现优于现有的临时协作方法,且这些合作伙伴与多个公约兼容。
GRPODropout: Less is More for Online Reinforcement Learning Rollouts
GRPODropout:在线强化学习推广的“少即是多”
- Authors: Hexuan Deng, Zihao Yan, Xuebo Liu, Shuo Nie, Yue Wang, Chen Wang, Zhaohua Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Min Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11854
- Pdf link: https://arxiv.org/pdf/2610.11854
- Abstract
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at this https URL.
- 中文摘要
强化学习(RL)方法如GRPO显著提升了大型语言模型推理能力,但常常存在策略熵崩溃的问题:采样多样性的丧失削弱了探索并限制了进一步改进。现有方法通过算法层面干预(如奖励修改和熵/KL正则化)或代币级重权重来解决这个问题。我们探讨了一个补充的视角:熵崩溃也可以通过改变哪些生成的推广对策略更新的贡献来缓解。在同一抽样预算下,并非所有推广都对更新有积极贡献,选择性排除部分可以改善学习。为此,我们提出了GRPODropout:在标准更新前,我们采用一种简单策略,选择性地剔除少量高概率正向优势的推广,并将保留的优势重新集中。为推动这一设计,我们开发了一套推广层级理论分析,指导方法设计和阈值选择。该方法仅改变推广使用情况,计算开销几乎不增加。实验显示,在更新时使用更少的推广样本,准确率高于原始GRPO和更高的演员熵,说明“少即是多”。这项工作提供了关于强化学习推广使用情况的洞见:移除部分推广可以提升性能。代码可在此 https URL 获取。
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
MiMo-v2.6:面向自我提升的扩展强化学习
- Authors: Xiaomi LLM-Core Team: Zongming Qiao, Ziyue Hua, Zirui Ou, Zihao Yue, Zihan Jiang, Zhuo Huang, Zhiyang Chen, Zhixian Zheng, Zhipeng Xu, Zhengrui Ma, Yuyang Hu, Yuhang Dong, Yuechen Zhang, Yudong Wang, Yuanxin Liu, Yixin Yang, Yishuo Cai, Yikai Zhao, Yihan Yan, Yifan Zhang, Yifan Song, Xiyu Wei, Xing Zhang, Xin Zhang, Xiaoqian Liu, Xiaodong Ji, Xiangwei Deng, Xueyu Guo, Wenhan Ma, Weimin Xiong, Weikun Wang, Weiji Zhuang, Shuo Liu, Shuhuai Ren, Shuhao Gu, Shimao Chen, Shijie Cao, Shihua Yu, Shicheng Li, Shengjie Zhou, Shaolei Zhang, Rang Li, Qiying Wang, Qingkai Fang, Qianli Chen, Minzheng Wang, Liwen Wang, Linli Yao, Linghao Zhang, Liangyu Cheng, Liang Zhao, Lei Li, Jinhao Dong, Jinyu Xiang, Jianyu Wei, Jiangshan Duo, Huaqiu Liu, Huanjie Fan, Hongyi Guan, Hongshen Xu, Hao Tian, Hanyu Li, Hailin Zhang, Gang Wang, Fuli Luo, Feng Wei, Dong Zhang, Dawei Zhu, Chiheng Lou, Chenhong He, Chenhao He, Chenghua Liu, Bowen Ye, Bowen Shen, Boshen Xu, Bo Yang, Bingquan Xia, Bangjun Xiao, Baixuan Xu, Zhouxiang Mao, Zhiyang Zhang, Zhixiang Xu, Zhenru Lin, Zhengju Tang, Zhaojun Huang, Yuzhe Weng, Yuxing Xiang, Yuxiao Li, Yuheng Yang, Yuhang Wang, Yuchen Liu, Yuanyuan Tian, Yuanliang Dong, Yu Cheng, Yongzhe He, Yongshun Liang, Yong Wang, Yiyan Wang, Yitian Gong
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.11959
- Pdf link: https://arxiv.org/pdf/2610.11959
- Abstract
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
- 中文摘要
强化学习(RL)是推动大型基础模型自我提升的核心训练范式。本报告介绍了MiMo-V2.6系列,这是一个全模态家族,通过扩展强化学习计算推动模型智能的前沿。在强化学习之前,我们对广泛的多模态语料库进行中期训练,以提供充足的探索空间,并在预训练混合-SWA架构上构建坚实基础设施以支持后续扩展。我们沿三维扩展RL计算:(1)更大批量和更高吞吐量,采用异步训练,每步消耗1568个样本和2.7-3.7亿令牌,上下文长度最高达1M;(2)更多样化和复杂的环境,涵盖代码、通用、可视化和网络领域,采用多种代理机束;以及(3)通过分组代理评分,增加评分计算,使长视野任务获得更准确的奖励信号,并引导模型朝向更短、更高效的令牌解决方案。为保持大规模训练稳定,我们冻结了MoE路由器,建立了多层防御机制,防止奖励黑客攻击。我们进一步构建混合任务智能强化学习的基础设施,包括统一轨迹表示、高并发多框架展开、解耦控制平面和数据平面以及训练-推理一致性。我们将训练动态、强化学习环境和强化学习框架开源,以促进复现和对大规模强化学习和模型自我提升的研究。
CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
能力:通过行为潜在编码实现能力感知策略适应
- Authors: Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Deyuan Qu, Zhiyuan Gao, Yanxiang Zhan, Jeroen Schafer, Andrew Melnik, Qing Yang, Michael Beetz
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.11971
- Pdf link: https://arxiv.org/pdf/2610.11971
- Abstract
Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning. CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction. The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers. Across 28 LIBERO tasks, CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator, and additional evaluations characterize transfer to unseen fault families and demonstrate recovery on a physical Franka Panda. this https URL
- 中文摘要
视觉-语言-动作(VLA)策略假设其训练对象,当关节故障改变指令动作的物理执行方式时,可能失败。现有的故障恢复方法通常需要任务特定再训练、故障标签、显式诊断或特权实现信息。我们介绍了CAPABLE,这是一个统一的能力感知适应框架,适用于冻结VLA,将自监督能力推断与残差强化学习整合。CAPABLE通过命令响应历史和运动学推断能力、每个关节实际实现的指令运动量及其对终端执行器行为的贡献,在线分析,使用跨关节的时间编码器、雅可比接地、交叉关节注意和自我监督物理预测。所得表示条件是一个残差策略,向VLA臂动作添加有界修正,且无故障标签或故障关节标识符。在28个LIBERO任务中,CAPABLE将排除在故障训练之外的执行器的成功率从24.8%提升至59.3%,比参数匹配的全局历史基线高出17.4分,同时保持了健康性能。在六个节点中进行的“缺失一个执行器”实验表明,这种转移并非针对单一执行器,进一步评估还描述了向未见故障族的转移,并展示了在物理Franka Panda上的恢复。此链接
When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation
代理者何时应思考?通过交叉回合估计的自适应推理
- Authors: Yiruo Cheng, Shen Huang, Xiaoshuai Song, Jiejun Tan, Guanting Dong, Pengjun Xie, Ji-Rong Wen, Zhicheng Dou
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.12061
- Pdf link: https://arxiv.org/pdf/2610.12061
- Abstract
Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.
- 中文摘要
基于大型语言模型(LLM)的智能体已在复杂任务中展现出强大能力。它们通常在交互轨迹中的每个动作前进行推理。然而,推理并非每个环节都必须,因为之前产生的推理可以继续支持后续动作。因此,一个关键挑战是判断现有推理是否足够,何时需要新的推理步骤,而不依赖昂贵的基于生成的验证。我们发现,去除额外推理后后续引用动作可能性的下降,紧密追踪这些行为在早期推理下是否仍可恢复,提供了有效且轻量级的信号来估算交叉转向动作支持。基于这一观察,我们提出了通过交叉转向估计实现推理适应(RACE)的方法,这是一种用于自适应智能体推理的训练方法。RACE引入了一种似然引导渐进推理覆盖检测(LoGiC)程序,逐步识别那些移除后对当前及后续参考动作影响有限的推理回合。由此产生的移除信号被纳入监督微调和代理强化学习中,使策略能够学习何时推理、何时直接行动。对四个代表性代理基准的广泛实验表明,RACE在保持或提升任务性能的同时,显著降低了推理成本。
Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design
通过逆最优设计实现未知非线性系统的预定义时间积分强化学习
- Authors: Tien Dat Vu
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.12103
- Pdf link: https://arxiv.org/pdf/2610.12103
- Abstract
This paper develops a new predefined-time integral reinforcement learning framework for optimal control of unknown nonlinear systems. The unknown drift is first approximated by a radial basis function (RBF) neural network, together with a data-driven online identification law for updating the corresponding neural weights. After the identifier converges to a sufficiently small neighborhood of the true dynamics, the learned model is incorporated into the integral reinforcement learning (IRL) problem. Unlike conventional reinforcement-learning-based optimal control, the desired convergence time is introduced directly into the control objective: a Lyapunov function and its prescribed decay behavior are specified by the designer, and inverse-optimal control is then used to construct a compatible running cost whose optimal policy inherits the predefined-time stabilization property. The value function is approximated by a second RBF neural network, and a new critic update law is developed to impose predefined-time convergence on the critic weights. Finite informative learning data are stored in a replay buffer and reused during the critic update, thereby avoiding the need for persistent excitation throughout the closed-loop operation. Theoretical analysis proves that the critic-weight error enters a prescribed residual set within the allocated learning horizon, while the closed-loop state reaches a small neighborhood of the origin within the overall designer-specified deadline. Numerical simulations on an unknown nonlinear system verify accurate drift reconstruction, predefined-time critic learning, and closed-loop convergence.
- 中文摘要
本文开发了一个新的预定义时间积分强化学习框架,用于对未知非线性系统的最优控制。未知漂移首先通过径向基函数(RBF)神经网络近似,并结合数据驱动的在线识别定律以更新相应的神经权重。当标识符收敛到真实动力学的足够小邻域后,所学模型被纳入积分强化学习(IRL)问题中。与传统的基于强化学习的最优控制不同,期望的收敛时间直接引入控制目标:设计者指定了李雅普诺夫函数及其规定的衰减行为,然后使用逆最优控制构建兼容的运行成本,其最优策略继承预定义的时间稳定性质。该值函数由第二个RBF神经网络近似,并开发了新的批判更新定律以对批判权重施加预定义时间收敛。有限的信息学习数据存储在回放缓冲区中并在批判更新期间重复使用,从而避免了闭环操作过程中持续激励的需求。理论分析证明,批判权重误差进入分配学习视野内的规定残差集,而闭环状态则在设计者指定的整体截止时间内到达原点的一个小邻域。在未知非线性系统上的数值模拟验证了准确的漂移重建、预定义时间批判学习和闭环收敛。
Credal Machine Learning for Risk-Averse Decision Making
Credal 机器学习用于风险规避决策
- Authors: Timo Löhr, Paul Hofman, Maximilian Muschalik, Eyke Hüllermeier
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.12115
- Pdf link: https://arxiv.org/pdf/2610.12115
- Abstract
In many machine learning applications, it is necessary to guard against worst-case scenarios and predictions that could result in substantial losses. In principle, this can be achieved by training risk-averse predictive models that minimize loss functions such as conditional value-at-risk (CVaR), rather than relying on models that perform well on average. In practice, however, the effectiveness of this approach to risk aversion is undermined by the learner's uncertainty regarding the true loss distribution and, consequently, the true CVaR. To achieve reliable risk-aversion, we propose a method in which this (epistemic) uncertainty is represented in terms of credal sets, i.e., sets of probability distributions. More specifically, we develop an efficient yet reliable learner that produces predictions in the form of credal sets and combine it with a novel decision rule that maps each credal set to a single predictive distribution for CVaR minimization. Across classification, under distribution shift, and in reinforcement learning, our approach reliably avoids catastrophic decisions, while sacrificing little in expected performance.
- 中文摘要
在许多机器学习应用中,有必要防范可能导致重大损失的最坏情景和预测。原则上,这可以通过训练风险规避型预测模型实现,这些模型最小化风险值(如风险条件值值,CVaR),而非依赖平均表现良好的模型。然而,在实际操作中,学习者对真实损失分布及真实CVaR的不确定性会削弱这种风险规避方法的有效性。为实现可靠的风险规避,我们提出一种方法,将这种(认知)不确定性用信度集表示,即概率分布集合。更具体地说,我们开发了一个高效且可靠的学习器,以信度集形式产生预测,并结合一种新颖的决策规则,将每个信件集映射到单一预测分布以实现CVaR最小化。在分类、分布偏移下以及强化学习中,我们的方法可靠地避免了灾难性决策,同时几乎不牺牲预期性能。
Q-Shaped Options for Hierarchical Reinforcement Learning
Q 形的层级强化学习选项
- Authors: Clarisse Wibault, Antoine Gorceix, Antonio Léon Villares, Alexey Zakharov, Evangelos Chatzaroulas, Michael Matthews, Eduardo Pignatelli, Jakob Foerster
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.12135
- Pdf link: https://arxiv.org/pdf/2610.12135
- Abstract
Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, using an action abstraction to represent temporally extended behaviour as options reduces the effective decision horizon. Second, enabling different state abstractions at each level of the decision process permits greater data aggregation for learning. However, realising these two benefits of a hierarchical policy depends on learning an appropriate action abstraction. Current HRL algorithms fail in one of two ways. Some discard distinctions between options needed for optimal control, undermining hierarchy altogether. Others retain unnecessary distinctions, preserving horizon reduction, but forfeiting coarser state abstraction. In this work, we characterise three desiderata for an action abstraction. We introduce Q-Shaped Options (QSO) to address all three. QSO builds on an architecture with distinct state-value functions, Q functions and policies at each level of the hierarchy. It learns the action abstraction between consecutive levels as a shared encoder shaped by their respective Q functions. The low-level Q function uses the option as a goal, encouraging the abstraction to retain distinctions necessary for optimal control. The high-level Q function uses it as an action, encouraging unnecessary distinctions to be discarded. Across offline goal-conditioned locomotion and manipulation environments, QSO learns semantically meaningful option spaces and outperforms baselines, achieving non-zero performance in tasks where all other evaluated algorithms fail.
- 中文摘要
学习应对长视野、目标条件任务需要智能体在较长时间尺度内进行推理,并在广泛的状态范围内行动。原则上,层级强化学习(HRL)通过动作(时间)与状态(空间)抽象的交互来解决这两个挑战。首先,使用动作抽象将时间延伸行为表示为选项,从而缩短了有效决策视野。其次,在决策过程的每个层面启用不同的状态抽象,可以实现更广泛的学习数据聚合。然而,实现分层策略的这两个好处依赖于学习合适的动作抽象。当前的HRL算法在两种方面失败。有些算法放弃了最佳控制所需的选项区分,完全破坏了层级结构。另一些算法保留不必要的区分,保留视野缩减,但放弃了更粗糙的状态抽象。本研究中,我们描述了动作抽象的三个期望。我们引入了Q形选项(QSO)来解决这三项需求。QSO基于具有不同状态值函数、Q函数和策略的层级结构。它学习连续层级之间的动作抽象,作为由各自Q函数塑造的共享编码器。低级Q函数以选项为目标,鼓励抽象保留对最佳控制所需的区分。高级Q函数将其作为动作,鼓励丢弃不必要的区分。在离线目标条件的移动和操作环境中,QSO学习语义有意义的选项空间,并超越基线,在所有评估算法失败的任务中实现非零性能。
Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility
利用人机参与的演示在强化学习中实现数字孪生驱动的机器人灵活性
- Authors: Yuzhu Sun, Mien Van, Nguyen Minh Nhat, Stephen McIlvanna, Sean McLoone
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.12140
- Pdf link: https://arxiv.org/pdf/2610.12140
- Abstract
Growing automation makes collaborative robots work in more variable environments, increasing the need for adaptation. We propose a human-in-the-loop online training framework combining a digital twin (DT), reinforcement learning (RL), and human demonstrations. Unlike DTs used mainly to generate synthetic data before task execution, our DT is synchronized with the physical system in real time through camera feeds, allowing the virtual robot to update its observations and policy from real-world feedback. A dual actor framework integrates imitation learning (IL) without adding a direct imitation loss to the RL actor, so demonstrations can guide adaptation instead of manual reprogramming. The proposed framework is demonstrated on the Ufactory Xarm5 collaborative robot, where the robot's end-effector aims to reach the target position while avoiding obstacles. The experiments show that the framework can resume training after a change in the physical workspace and that, with a fixed set of non-optimal demonstrations, the dual actor framework achieves a much higher final success rate than two methods that add an imitation loss to the actor. The same pattern holds with real human demonstrations collected in virtual reality (VR): with demonstrations that never reach the goal, the dual actor framework reached 83-100% mean deterministic evaluation success, against 0-17% for the two imitation-loss methods.
- 中文摘要
自动化的增长使协作机器人能够在更多样化的环境中工作,增加了适应性的需求。我们提出了一种结合数字孪生(DT)、强化学习(RL)和人类演示的在线人机环路训练框架。与主要用于生成合成数据的DT不同,我们的DT通过摄像头实时同步与物理系统同步,使虚拟机器人能够根据真实反馈更新观察和政策。双角色框架集成了模拟学习(IL),但不直接为RL演员添加模拟损失,因此演示可以引导适应,而非手动重编程。该框架在Ufactory Xarm5协作机器人上演示,机器人的端执行器旨在避开障碍物,达到目标位置。实验表明,框架在物理工作空间发生变化后可以恢复训练,并且在固定的非最优演示集下,双角色框架的最终成功率远高于给演员增加模拟损失的两种方法。同样的模式也适用于虚拟现实(VR)中收集的真实人类演示:对于未达到目标的演示,双角色框架的平均确定性评估成功率达到了83%-100%,而两种模仿损失方法均为0-17%。
Sim-to-Real RL for ASVs using SysID
使用 SysID 的模拟到真实 ASV 强化学习
- Authors: Cody Sheltraw, Tsimafei Lazouski, Maani Ghaffari, Alan Papalia
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.12202
- Pdf link: https://arxiv.org/pdf/2610.12202
- Abstract
Autonomous Surface Vehicles (ASVs) operating in dynamic marine environments require robust control policies for tasks such as path following and station keeping, making reinforcement learning (RL) a promising alternative to classical controllers. However, existing ASV simulators rarely support parallel environments for RL training. Such existing simulators require accurate hydrodynamic modeling from computational fluid dynamics solvers or towing tank tests for setting hydrodynamic parameters to address the sim-to-real gap. To address these challenges, we present an ASV simulator and accompanying pipeline that enables training policies starting from unknown vehicle dynamics. Our framework uses only a CAD model and brief set of open-water field trajectories for approximating and refining both hydrodynamic and thruster parameters. Real-world deployments on a BlueBoat ASV demonstrate successful zero-shot sim-to-real transfer in path following and station-keeping tasks without prior hydrodynamic and propeller information.
- 中文摘要
在动态海洋环境中运行的自主水面载具(ASV)需要稳健的控制策略来执行路径跟踪和站位保持等任务,使强化学习(RL)成为传统控制器的有前景替代方案。然而,现有ASV模拟器很少支持强化学习的并行环境训练。现有模拟器需要通过计算流体动力学求解器或拖曳罐测试进行精确的流体动力学建模,以设置水动力参数以弥补模拟与实际差距。为应对这些挑战,我们提出了ASV模拟器及其配套流程,支持从未知车辆动力学出发的训练策略。我们的框架仅使用CAD模型和简短的开阔水域场地轨迹,用于近似和优化水动力学和推进器参数。BlueBoat ASV的实际部署展示了在没有先前水动力学和螺旋桨信息的情况下,成功实现零发射模拟到实的路径跟踪和站位保持任务。
VibeEdit: Image Editing with Canvas Instructions
VibeEdit:使用Canvas操作说明进行图像编辑
- Authors: Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.12229
- Pdf link: https://arxiv.org/pdf/2610.12229
- Abstract
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
- 中文摘要
在文本引导图像编辑中,描述期望的变更通常很简单,但识别目标对象或区域可能很繁琐,尤其是当多个对象看起来相似时。我们引入了一个新的图像编辑界面,允许用户直接在图像上放置空间标记和可选的简短注释。这些注释共同构成一个画布指令,指定编辑位置和更改内容。我们的编辑器VibeEdit遵循这些指令,在不需单独文本提示的情况下执行对象添加、移除、替换、属性修改和移动。我们构建了155万对源-目标编辑对,包含对象掩码和结构化编辑描述,并在训练时从中渲染画布指令。我们将Qwen-Image-Edit采用层解耦条件,分别编码源图像和画布指令进行图像编辑。我们通过区域加权监督微调训练模型,随后采用评分标准引导强化学习,以提升编辑完成度、局部编辑质量及未编辑区域的保存。我们基于独立构建的、由人类策划的419个案例基准评估VibeEdit,重点在相似对象中进行目标选择。VibeEdit的VLM评分为79.9,外区域PSNR为32.8 dB,而FireRed是我们评估中评分最高的文本指导基线,分别为67.4和24.0 dB。
Walking on Roofs: Exploring the Potential of Walking Robots for Construction Work on Roofs
屋顶行走:探索步行机器人在屋顶施工中的潜力
- Authors: Bjoern-Felix Dettmar, Arne Roennau
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.12272
- Pdf link: https://arxiv.org/pdf/2610.12272
- Abstract
This paper investigates the feasibility of deploying quadruped walking robots for the automation of work in roof environments. While quadrupeds have demonstrated versatility across various domains, their large-scale deployment remains limited, partly due to lack of application-specific designs. Roof environments represent a novel and unexplored use case, combining high safety risks for human workers with repetitive, strenuous tasks that could benefit from robotic assistance. A dedicated test rig of a roof's surface was designed to evaluate the baseline performance of a commercial quadruped, the \emph{Unitree Go2}, in this new environment. Experiments revealed that standard ball feet are inherently inadequate for locomotion on sloped roofs: slippage increased quadratically with incline angle, get-up and lie-down sequences were only possible on small inclines, and critical failures already occurred regularly on moderate inclines of 25°. This work provides the first systematic assessment of quadruped locomotion in roof environments, highlighting both the potential and the current limitations of this application and establishes a baseline for future research. Effective solutions will require a combination of task-oriented foot designs, advanced contact mechanics and environment-specific control strategies, such as reinforcement learning for roof-adapted gaits.
- 中文摘要
本文探讨了在屋顶环境中部署四足行走机器人自动化工作的可行性。尽管四足机器人在多个领域展现出多功能性,但其大规模部署仍有限,部分原因是缺乏针对特定应用的设计。屋顶环境是一种新颖且尚未被探索的应用场景,结合了对人类工人的高安全风险与重复性、艰巨的任务,这些工作本可受益于机器人辅助。专门设计了屋顶表面测试设备,以评估商业四足动物\emph{Unitree Go2}在该新环境中的基线性能。实验显示,标准球脚在斜屋顶上本质上不够适宜:滑动随倾斜角呈平方加剧,起伏和躺下序列仅在小坡度上可行,且在25°中等坡度上已频繁发生关键故障。这项工作首次系统评估了屋顶环境中的四足行走,突出了该应用的潜力与当前局限性,并为未来研究奠定了基础。有效的解决方案需要结合任务导向的脚部设计、先进的接触力学和环境特定控制策略,如对屋顶适应步态的强化学习。
HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
HarnessSQL:现实数据库环境中的SQL代理本地开发培训
- Authors: Haolin Yang, Jipeng Zhang, Jian Xie, Shuaishuai Gong, Sirui Han, Yike Guo
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.12274
- Pdf link: https://arxiv.org/pdf/2610.12274
- Abstract
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
- 中文摘要
文本转SQL模型通常训练为将问题直接映射为静态查询,而现实中的数据库代理则通过与实时数据库进行有状态、多回合交互——检查模式、执行探针查询、诊断错误和修正假设。这造成了关键的训练-部署不匹配,因为中介这种交互的执行束只在推理时引入。为弥合这一差距,我们提出了HarnessSQL,一种基于约束的原生后训练框架,在监督微调和强化学习过程中保持完整的交互结构。HarnessSQL构建孤立的可执行数据库环境,配合隐藏的执行预言机,直接在目标SQL框架内部署教师,并仅保留经过验证的全序列SFT轨迹,随后执行奖励强化学习。在Spider 2.0-SQLite中,HarnessSQL大幅提升了紧凑模型的执行准确性,Qwen3-8B从15.5%提升至45.2%,Qwen3-14B从22.2%提升至54.8%,同时有效转移至BIRD-Interactive和LiveSQLBench等非发行交互基准。我们的发现表明,直接在其执行框架内训练数据库代理对于掌握复杂且长期的数据库工作流至关重要。
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
人工智能代理能否通过学习登顶?在长期游戏代理竞赛中评估启发式学习
- Authors: Kaisen Yang, Qingle Liu, Kejin Wang, Yicheng Zhao, Jieming Li, Shenghan Zheng, Ruize Yang, Bojun Yang, Heng Gong, Xiang Gao, Lanyue Zhang, Kaiyu Zhong, Zhuo Liu, Shaoxuan Li, Chengxi Li, Yong Yan, Weixuan Zhang, Tianwei Luo, Situ Wang, Youjie Zheng, Sihan Zhao, Shengyuan Wang, Huan-ang Gao, Jiazheng Xu, Xiaohui Xie, Wentao Han, Hongning Wang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.12341
- Pdf link: https://arxiv.org/pdf/2610.12341
- Abstract
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
- 中文摘要
对抗性博弈推动了从启发式搜索向强化学习的进步,但从有限样本中学习和调整策略仍具挑战性。AI代理通过将游戏体验转化为可执行策略的修订,提供了另一种选择。基于启发式学习(HL),我们正式提出了对抗性启发式学习(AHL),这是一种利用AI代理作为学习引擎,优化游戏策略和支持软件,同时保持模型权重固定的范式。我们介绍AAArena,这是一个基准测试,包含12款真实对抗博弈和1920个存档的人类程序,其评估协议基于真实比赛竞赛。代理解释规则、选择对手、分析回放并修订游戏代理,以在固定比赛和评估预算内获得最高排名。我们评估了 \val{completedmodels} 模型和 harness 配置:Opus5.5 与 Claude Code 共获得 6 枚金牌,而没有任何已评估配置能超过其余 6 个人梯子。在规则规范更复杂的游戏中,性能通常较弱。进一步实验显示,对手选择和密集反馈支持策略改进,代理能从自身比赛的策略内重放和其他玩家的非策略重赛中学习。这些结果凸显了 HL 在对抗性博弈中的潜力,并识别出博弈理解、战略实施和长期策略制定中的持续挑战。
AgentGarten: Code Worlds for Evolving Agents
AgentGarten:进化特工的代码世界
- Authors: Jiawei Chi, Shangchen Miao, Zhiyuan Shi, Kailu Wu, Hanyang Wang, Weiliang Chen, Qiyu Dai, Jinshan Ren, Jun Gao, Mingsheng Long, Yueqi Duan, Jiangran Lyu, Jialong Wu, Fangfu Liu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.12374
- Pdf link: https://arxiv.org/pdf/2610.12374
- Abstract
Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.
- 中文摘要
交互式虚拟世界允许代理通过探索和互动来学习。代理能学习的内容受限于他们所实践的环境,必须忠实于一致的状态、规则和动态,同时观察必须符合现实世界的视觉分布。在多元世界中实现这两者仍是一个瓶颈。我们介绍了AgentGarten,一个将模拟器和游戏引擎与共享神经渲染器结合起来的框架,构建实时交互环境。其模拟后端维护持久世界状态并执行程序定义的交互规则,而渲染器则从结构化条件生成视觉观察,通过通用接口导出。为了构建神经渲染器,我们将预训练的视频模型适配到几何条件,通过提出的对抗强迫进行提炼,并优化实时交互的推理。对抗强制使历史预填充可微分,通过精确回放,后续预测的损失更新渲染器对先前观察的编码,并添加真实数据的对抗监督以提升视觉质量。在AgentGarten中,代理通过视觉观察感知世界,实时互动,并通过将每一轮经验提炼成后续代理继承和完善的操作手册来提升。我们的实证研究显示,学习效率显著提升,代理仅用4轮学习即可学习,而传统强化学习对应者需数百万轮。随着新世界可以编写代码并通过同一界面渲染,环境可以在数量和难度上随代理同步调整,迈向通过交互体验不断进化的代理。
A Unified Bellman Operator for Safety-Critical Reinforcement Learning
一个统一的贝尔曼操作员,用于安全关键强化学习
- Authors: Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach, Jaime Fernández Fisac
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.12420
- Pdf link: https://arxiv.org/pdf/2610.12420
- Abstract
Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.
- 中文摘要
在安全关键领域进行强化学习,需要在严格遵守安全约束的同时最大化任务性能。现有的安全强化学习范式通常迫使双方权衡:要么需要先验知识以提供严格的安全保障(例如安全滤波器),要么支持联合学习但平均仅满足安全约束。本研究提出了一种新颖的贝尔曼算子,将性能与安全目标统一为联合价值函数。我们证明了联合贝尔曼算符的时间差分学习在两个时间尺度的随机近似框架下收敛。在快速时间尺度上,估计学习联合策略的安全值,而在慢时间尺度估计联合值。通过将极限动力学表述为职业平均的微分包含关系,并证明其渐近收敛到一组极限且安全约束约束的任务值函数,从而确保收敛。理论上,一旦收敛,最终的最优策略在始终保持安全的情况下最大化任务回报。对连续控制任务进行神经近似的实证评估显示,测试时几乎没有安全违规,收敛稳定。
FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems
FAITH:可行性感知安全过滤强化学习,适用于高维系统
- Authors: Songyuan Zhang, Baljeet Singh, Sarthak Ranjeet Kaingade, Chuchu Fan, Bryan Trinh
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.12432
- Pdf link: https://arxiv.org/pdf/2610.12432
- Abstract
Safe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filters separate them at action execution, but classical designs require an analytic safety function and dynamics model, and standard minimal-intervention filters are myopic to long-horizon task return because they minimize only instantaneous action deviation. Hard projections are also undefined when no safe action exists. We present FAITH, a feasibility-aware, model-free framework that approximates the optimal state-action safety value and amortizes minimal-intervention filtering with a feedforward network. The task policy optimizes the task return through the filtered dynamics, which recovers the feasible constrained problem without a competing safety term in the task-policy update. When no action satisfies the learned safety condition, the same filter approaches the action with minimum predicted peak harm. On a double integrator example and a Safety Gym environment, FAITH achieves the highest return among methods with no feasible-start violations and matches the lowest harm from infeasible starts. On a 29-DoF humanoid, it reaches a 99.95% safety rate while retaining 97% of the unfiltered return in Walking-Avoid, and obtains the highest measured safety rate in Push-Avoid by learning to sacrifice balancing and fall away from the protected region. The same policies are also demonstrated on a real-world Unitree G1 humanoid.
- 中文摘要
安全强化学习通常将安全和任务绩效置于同一策略目标中,两者可以引入竞争的更新。安全过滤器在动作执行时将它们分开,但经典设计要求分析安全函数和动态模型,标准最小干预过滤器因仅最小化瞬时动作偏差,导致任务返回视角短至长视角。当不存在安全动作时,硬预测也未定义。我们提出FAITH,一个可行性意识、无模型的框架,近似最优状态动作安全值,并通过前馈网络摊销最小干预过滤。任务策略通过过滤动力学优化任务返回,恢复可行受限问题,且任务策略更新中无竞争安全项。当无行动满足学习的安全条件时,同一过滤器以最小预测峰值伤害接近该动作。在双积分器示例和安全健身房环境中,FAITH在无可行起始违规的方法中获得最高回报,且匹配不可行起始伤害最低。在29 DoF类人形生物上,行走避让中保持97%未滤波回报的99.95%安全率,通过学会牺牲平衡并远离保护区域,获得推避最高安全率。同样的策略也在现实世界的Unitree G1人形生物中得到验证。
Generative Neural Retargeting for Human-to-Robot Dexterous Manipulation
生成神经重定向用于人对机器人灵巧操作
- Authors: Dechen Gao, Yue Yang, Ben Abbatematteo, Nathan Godwin, Pengcheng Wang, Roger Boldu, Steven Man, Zhiyang Dou, Chuan Qin, Sho Nakagome, Eric Whitmire
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.12440
- Pdf link: https://arxiv.org/pdf/2610.12440
- Abstract
Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only $8.5\%$ of the samples required by MPC, achieving a success rate of $56.20\%$ compared to $27.20\%$ for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning $223$k demonstrations and $3.3$k object geometries.
- 中文摘要
人类演示是学习灵巧操作的可扩展数据源,但身体差距阻止了人类运动直接在机器人上执行。逆运动学(IK)将人类运动高效地重新定位到机器人,但忽略动力学,常常产生不可行的运动。强化学习(RL)和基于抽样的模型预测控制(MPC)常用于产生动态可行的运动,但两者样本效率低且对超参数敏感。强化学习存在昂贵且不稳定的训练和繁琐的奖励工程;MPC避免策略优化,却单独重新定位每个轨迹,解决一个轨迹并不会让下一个更容易。采样成本随着数据集规模和任务难度迅速增长。我们假设动态可行轨迹集中在演示共享的低维流形附近,因此重定向可简化为从该流形中采样,条件为人体运动,而非每次演示都需重新解一个优化问题。我们提出了 \textbf{生成神经重定向}(GNR),利用流量匹配模型采样可行轨迹。GNR 仅以 MPC 所需样本的 8.5 美元,成功率超过 56.20 美元,而 MPC 仅为 27.20% 美元。GNR可用于大规模、长视距和毫米级精度人体演示的可扩展且高效的再定向:通过在实模拟数据引擎中应用GNR,我们生成了一个灵活的操作数据集,带有密集的接触力标签,涵盖22.3万美元的演示和3.3万美元的物体几何体。
A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
平衡数据饮食:解决机器人控制大规模强化学习探索瓶颈
- Authors: Octi Zhang, Mateo Guaman Castro, Patrick Yin, Ignacio Dagnino, Abhishek Gupta, Rosario Scalise, Byron Boots
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.12465
- Pdf link: https://arxiv.org/pdf/2610.12465
- Abstract
General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations. Recent work has shown that diverse simulator resets, combined with massively parallel simulation, can alleviate much of this engineering burden on several manipulation problems. However, we find that naively scaling this paradigm to more precise or dynamic problems remains non-trivial. While simulator resets can help with exploration, uniformly sampling over this distribution wastes a growing fraction of learning experience on task configurations the policy has already mastered or cannot yet attempt. This makes it challenging to see the expected benefits of scaling parallel environments for RL, since much of the learning signal in a batch is wasted during learning. To mitigate this, we introduce Success Guided Sampling (SGS), a simple adaptive sampler that concentrates RL training on task configurations around the frontier of the policy's capabilities. Doing so allows large-scale simulated RL to make the most out of the experience in a batch, enabling much more effective scaling to large-scale parallel simulation. Across experiments using up to $2^{20}$ (over one million) parallel environments, SGS enables RL to solve challenging multi-terrain quadruped locomotion and contact-rich assembly tasks that prior methods fail to solve. Finally, we distill the learned manipulation policies into RGB-based policies and demonstrate zero-shot transfer to several challenging assembly tasks on real hardware. Project website: this https URL.
- 中文摘要
通用机器人必须执行从敏捷移动到灵巧操作的广泛任务。虽然模拟到现实强化学习(RL)已被证明是实现这一目标的有用工具,但当前的强化学习流程依赖于工程化、每任务结构先验,如形状奖励和演示。最新研究表明,多样化的模拟器重置结合大规模并行模拟,可以减轻多个操作问题的工程负担。然而,我们发现将这一范式天真地扩展到更精确或动态的问题仍然不简单。虽然模拟器重置有助于探索,但对该分布进行均匀采样会浪费越来越多的学习经验在策略已掌握或无法尝试的任务配置上。这使得在学习过程中难以看到并行环境扩展的预期效益,因为批量学习中的大量学习信号被浪费。为缓解这一问题,我们引入了成功引导采样(SGS),这是一种简单的自适应采样器,将强化学习集中在策略能力前沿的任务配置上。这样做使大规模模拟强化学习能够在批量中最大化利用体验,从而实现更高效的大规模并行模拟扩展。在使用高达2^{20}$(超过一百万美元)并行环境的实验中,SGS使强化学习能够解决以往方法无法解决的多地形四足行走和接触丰富组装任务。最后,我们将所学的操作策略提炼为基于RGB的策略,并演示了零样本转移至多个真实硬件上具有挑战性的组装任务。项目网站:此 https URL。
Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
Dex-One2Many:通过一次人类演示学习灵巧操作
- Authors: Jusuk Lee, Sungha Kim, Yeonsoo Park, Jonguk Cheon, Yoonkyo Jung, Yongjun You, H. Jin Kim, Jia-Bin Huang, Furong Huang, Youngseok Jang, Seungjae Lee
- Subjects: Subjects:
Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.12470
- Pdf link: https://arxiv.org/pdf/2610.12470
- Abstract
While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.
- 中文摘要
虽然从单一人类视频中学习灵巧操作提供了替代昂贵机器人演示的有前景替代方案,但许多近期方法主要模仿演示的动作。这种严格的动作匹配通常限制了推广到视频中未显示的初始物体姿态、目标姿势和抓取。另一种,通过强化学习(RL)发现策略可以实现广泛的推广,但如果没有先验指导,在复杂多阶段任务中高维探索时会遇到困难。为应对这些结合的泛化与探索挑战,我们提出了Dex-One2Many,一个从真实到模拟再到现实的框架,能够从单一人类视频中学习可推广的灵活操作策略。我们的关键见解是将视频抽象为序列场景图,引导RL,实现高效探索,同时保持广泛的泛化性。这些图表作为生成约束,用于采样多样化的重置状态,并为每个阶段提供密集的奖励。由于这些图表限制的是关系而非精确姿态,这些复位状态覆盖了视频之外的物体姿态和抓取,同时从它们初始化每个阶段并获得密集的奖励,使探索时间短且有指导性。Dex-One2Many 完全在模拟中训练,将零射击转移到真实多指手中。在五个工具使用和操作任务中,Dex-One2Many 在可见配置中比基线多出 6.5%,其稳健的泛化则在未见场景中将差距扩大到 71%。
Keyword: diffusion policy
Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
诊断与恢复长视界技能层的观测空间转移
- Authors: Pranav Wagh, Yu Fang, Yue Yang, Mingyu Ding
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10810
- Pdf link: https://arxiv.org/pdf/2610.10810
- Abstract
Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned $\pi_{0.5}$ policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
- 中文摘要
长视野机器人操作通常通过串联独立训练的技能构建。虽然每个技能单独可可靠,但当技能链化时性能急剧下降:每个下游技能必须从前一个技能留下的状态开始,而非从训练分布开始。我们研究了这种失败模式——观察-空间转移(OSS),并探讨这些技能接缝失效的原因。利用特权模拟器重置,我们发现主导的转移来自场景状态的位移(例如,抽屉打开或先前技能留下的次要物体),而非机器人的关节配置或下游技能操作的对象。为验证这一诊断,我们构建了一个完全学习的检测-恢复-恢复系统:任务进度监控检测停滞,学习策略恢复错位场景组件,缝隙稳健微调使技能恢复。它恢复了所有测试方案失败的接缝,我们将其视为诊断证据,而非通用方法。在BOSS-44基准测试中,系统将全链成功率从7.6%提升至26.5%,比基础策略提升3.5倍,提升特权恢复预言机的51%,而K中最佳重采样、扩散策略和世界模型基线则无法从评估的接缝状态恢复。在运行微调$\pi_{0.5}$策略的真实Franka臂上,同一显示器受限于外部摄像头可观测性,但闭环仍能恢复一些终端故障,促使手腕和握把感知。这些结果表明,某些长期合成失败通过恢复场景再恢复策略,比从非支持状态重试更为有效。
SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills
SkillWeave:将异质演示融入长远视野操控技能
- Authors: Ryosei Tamura, Xiaoxiang Dong, Uksang Yoo, Yuemin Mao, Romina Mir, Jonathan Francis, Jeffrey Ichnowski
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.12046
- Pdf link: https://arxiv.org/pdf/2610.12046
- Abstract
Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills. To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution. Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. Videos and code are available at this http URL .
- 中文摘要
灵巧操作既需要大规模任务进展,也需要精确的丰富接触互动,因此很难收集有效支持这两种模式的演示。我们提出了SkillWeave,一种异构的远程远程操作演示框架,用于粗略伸手和传输,以及精准、接触丰富技能的动觉教学。为解决演示者在动觉数据收集过程中出现的视觉不匹配,我们提出了一种对象-掩码条件的扩散策略,使用离线对象分割进行训练监督,并在部署时使用轻量级学习遮罩预测器,避免在线分割和图像修复。为减少独立训练子任务策略之间的分布转移,我们引入了后继感知终端引导,从前任策略中抽样的动作中选择,引导系统走向后继策略已展示的初始状态分布支持的状态。在三个现实世界的长期视野任务中,SkillWeave 实现了27%的平均端到端成功率。掩膜条件化的动觉策略将灵巧子任务的成功率提升至平均65%,而继承者意识切换的平均组成效率达到87%。这些结果表明,匹配示范模态与交互机制、明确解决动觉视觉不匹配以及引导政策交接到继承支持状态,显著提升了长期的灵巧操作。视频和代码可在此 http 网址获取。