生成时间: 2026-08-07 17:01:08 (UTC+8); Arxiv 发布时间: 2026-08-07 20:00 EDT (2026-08-08 08:00 UTC+8)
今天共有 32 篇相关文章
Keyword: reinforcement learning
Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
当工具能完成作业时教授入门人工智能:课程重设计与学生权利法案
- Authors: Yusuf Pisan
- Subjects: Subjects:
Computers and Society (cs.CY)
- Arxiv link: https://arxiv.org/abs/2608.05175
- Pdf link: https://arxiv.org/pdf/2608.05175
- Abstract
Large language models can complete most of the assignments in an introductory artificial intelligence course. This paper is an experience report on redesigning one such course, CSS~382 at the University of Washington Bothell, in response. Rather than freeze the curriculum, the redesign retained the course's classical core (search, adversarial search, Markov decision processes, reinforcement learning) and added a strand in which students build a large language model from scratch, so that a tool they are required to use is also one they are required to understand. Assessment was rebuilt around tasks that resist unattributed automation: in-class exercises, reflective writing, and a defended team project, with examinations removed entirely. The policy on AI was inverted, from unmentioned in 2023 to required in 2026. The center of the paper is a participatory ethics sequence in which a cohort of students deliberated on and endorsed a "Student Bill of AI Rights" governing their instructor's own use of AI, including a requirement that the instructor personally complete any AI-generated assignment before issuing it. The provisions were scaffolded by an AI-generated prompt and ratified by the students, and that provenance is part of what the account examines. The design, the student-authored artifacts, and the tensions that followed are reported, including student objections to AI-generated course materials, with explicit attention to the limits of what a single-cohort design narrative can claim.
- 中文摘要
大型语言模型可以完成入门人工智能课程中的大部分作业。本文是一篇关于重新设计华盛顿大学博塞尔分校CSS~382课程的经验报告。这次重新设计没有冻结课程内容,而是保留了课程的经典核心内容(搜索、对抗性搜索、马尔可夫决策过程、强化学习),并增加了一条方向,让学生从零开始构建大型语言模型,使他们必须使用的工具也必须理解。评估系统围绕那些难以被无归属自动化的任务重组:课堂练习、反思写作和一个有防御性的团队项目,考试完全取消。人工智能政策被颠倒,从2023年未提及变为2026年强制执行。论文的核心是一个参与式伦理序列,一组学生讨论并支持了一项“学生人工智能权利法案”,该法案规范其教师自身使用人工智能,其中包括教师在发布作业前必须亲自完成。这些条款由AI生成的提示搭建,并由学生们批准,这一来源是本讲述所探讨的内容之一。书中报道了设计、学生创作的产物以及随之而来的紧张关系,包括学生对AI生成课程材料的反对,并明确关注单一群体设计叙事所能主张的局限性。
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
Search2Skill:通过基于评分标准的强化学习超越知识边界的技能提炼
- Authors: Muyang Ye, Tian Lan, Feihu Jiang, Yongshi Ye, Wuyunsiqin, Bin Zhu, Qianghuai Jia, Zhao Xu, Weihua Luo, Ye Wang, Jinyang Zhang, Longyue Wang, Lingfeng Bao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.05245
- Pdf link: https://arxiv.org/pdf/2608.05245
- Abstract
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model's parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent's capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.
- 中文摘要
可重复使用的技能,涵盖解决现实职业任务所需的程序知识,为基于LLM的智能体提供了一条通往专家领域自我进化的路径。现有的自我演化技能方法通过模型的参数化知识或轨迹内部构建技能,因此受限于模型已有的知识。然而,专业技能背后的领域惯例和标准程序往往超出这一界限,单靠代理人很难获得。为解决这一问题,我们提出了一个新框架Search2Skill,自动识别代理能力缺口,搜索外部资源以应对,并将检索到的证据提炼成结构化、可重复使用的技能。具体来说,Search2Skill通过基于评分标准的强化学习方案进行优化,共同提升了何时搜索、如何搜索以及如何生成技能。基于三个基准测试的八个专家级领域的实验显示,Search2Skill在流式和保留式评估协议下,始终优于搜索增强和基于轨迹的技能学习基线。进一步分析显示,这些提升来自技能抽象,而非原始检索的证据,且获得的技能能跨模型尺度转移。
An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals
新兴的零售投资组合管理应用:个性化、税务意识强化学习,结合自然语言目标
- Authors: Ramin Pishehvar
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2608.05255
- Pdf link: https://arxiv.org/pdf/2608.05255
- Abstract
Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums and technology stacks unavailable to individual investors. We present a fully built, integration-tested application that closes this gap: a FastAPI backend and web dashboard that let a user describe an investment goal in plain language (e.g. "I want steady growth but need to sell some shares next month for a down payment"), routes that goal to one of six investment mandates, and produces a live, broker-integrated portfolio recommendation from athree-phase reinforcement learning system -- a self-supervised cross-asset encoder, a Mixture-of-Experts (MoE) allocation policy with a learned intent router, and a lightweight LoRA adapter that personalizes recommendations from an individual's revealed brokerage behavior without retraining the shared model. The system is functionally complete and integration-tested end-to-end against a live brokerage API (Alpaca, paper-trading mode), including multi-user authentication, a trust first preview-before-apply confirmation flow, daily email digests, and an auditable action-integrity chain, but has not yet been opened to real end-users; we report this honestly as an emerging, pre-deployment application with a concrete path to full deployment, alongside 14-day walk-forward backtests (bootstrapped confidence intervals included) as preliminary, pre-deployment validation rather than production performance. We also report several practical engineering lessons -- silently-inactive integration paths, hanging third-party API calls, and the value of end-to-end empirical verification over trusting checkpoint metadata -- that we believe generalize to other applied RL systems built on external, live data sources.
- 中文摘要
散户投资者缺乏机构客户理所当然享有的个性化、税务意识的投资组合管理——现有机器人顾问采用静态、基于规则的配置,机构级系统则要求账户最低要求和个人投资者无法获得的技术栈。我们提供了一个完整构建、经过集成测试的应用,弥合了这一差距:一个FastAPI后端和网页仪表盘,允许用户用通俗易懂的语言描述投资目标(例如“我想要稳定增长,但下个月需要卖出一些股票作为首付”),将该目标路由到六个投资任务之一,并生成一个实时的, 来自三阶段强化学习系统的经纪人集成投资组合推荐——一个自我监督的跨资产编码器、一个带有学习意图路由器的专家混合(MoE)配置策略,以及一个轻量级LoRA适配器,能够根据个人揭示的经纪行为个性化推荐,而无需重新训练共享模型。该系统功能完整,并通过端到端集成测试,基于实时经纪API(Alpaca,纸币交易模式),包括多用户认证、信任优先预览后应用确认流程、每日邮件摘要和可审计的操作完整性链,但尚未向真实终端用户开放;我们诚实地报告,这是一款新兴的部署前应用,拥有一条明确的全面部署路径,同时还将14天的前行回测(包括自举置信区间)作为初步部署前验证,而非生产性能。我们还报告了几个实际工程经验——无声非激活的集成路径、第三方API调用的延迟,以及端到端经验验证相较于信任检查点元数据的价值——我们相信这些经验可以推广到基于外部实时数据源的其他应用强化学习系统。
Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks
TSN网络中队列级XR流量调度的多代理变换器
- Authors: Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.05340
- Pdf link: https://arxiv.org/pdf/2608.05340
- Abstract
Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (XR). However, the widespread adoption of XR introduces significant challenges due to co-located services in MEC environments, leading to contention for shared network resources. Moreover, XR traffic types have distinct characteristics and criticality in terms of timing requirements, further increasing the complexity and dynamics of such environments. Although reinforcement learning has shown promise for TSN scheduling optimization in dynamic network scenarios, existing approaches rely on centralized or high-level multi-agent designs and are typically tailored to periodic and predictable industrial traffic, limiting their applicability to XR workloads. As a result, these approaches suffer from (i) limited ability to capture inter-queue dependencies due to coarse-grained control, and (ii) poor adaptability to highly dynamic and heterogeneous XR traffic. To address these gaps, we propose a multi-agent reinforcement learning approach for queue-level XR traffic scheduling. We adopt the multi-agent transformer (MAT) to model inter-queue dependencies via attention over agents' observations and actions, enabling implicit coordination across heterogeneous co-located XR applications. Our simulation results show that the proposed method outperforms baselines, achieving up to 71.42% latency reduction and up to 83.2% reduction in failure rate, while consistently achieving high reliability across all queues.
- 中文摘要
时敏网络(TSN)和移动边缘计算(MEC)在实现超可靠低延迟通信方面具有强大潜力,适用于如扩展现实(XR)等时间敏感应用。然而,XR的广泛采用带来了重大挑战,因为MEC环境中服务共址,导致共享网络资源的竞争。此外,XR流量类型在时序要求方面具有独特的特性和关键性,进一步增加了此类环境的复杂性和动态性。尽管强化学习在动态网络场景下的TSN调度优化中展现出潜力,但现有方法依赖集中式或高级多代理设计,且通常针对周期性和可预测的工业流量,限制了其适用于XR工作负载。因此,这些方法存在:(i)由于粗粒度控制,捕捉队列间依赖的能力有限,以及(ii)对高度动态和异构的XR流量适应能力较差。为弥补这些空白,我们提出了一种多代理强化学习方法,用于队列级XR流量调度。我们采用多代理变换器(MAT)来通过关注代理的观察和动作来建模队列间依赖关系,实现跨异构共址XR应用的隐性协调。我们的模拟结果显示,所提方法优于基线,延迟降低高达71.42%,故障率降低高达83.2%,同时在所有队列中始终保持高可靠性。
Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application
多智能体强化学习用于时间敏感应用中的在线流量调度
- Authors: Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.05346
- Pdf link: https://arxiv.org/pdf/2608.05346
- Abstract
Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). However, existing TSN scheduling solutions predominantly rely on static optimization techniques or centralized learning models that are based on fixed traffic patterns, limiting their effectiveness in dynamic environments. In practice, MEC environments often host multiple co-located XR traffic flows whose characteristics evolve over time, creating complex inter-queue dependencies that current schedulers fail to capture. Addressing these challenges requires adaptive, decentralized scheduling mechanisms capable of coordinating multiple TSN queues under varying traffic conditions. To this end, this paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent. The Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues. The simulation results demonstrate that the proposed approach reduces average frame waiting times by up to 26.8% and worst-case delays by approximately 16.8%, highlighting its effectiveness in dynamic XR-driven MEC scenarios.
- 中文摘要
时间敏感网络(TSN)正日益集成到移动边缘计算(MEC)中,以支持延迟要求严格的应用,如扩展现实(XR)。然而,现有的TSN调度解决方案主要依赖基于固定流量模式的静态优化技术或集中学习模型,限制了其在动态环境中的有效性。实际上,MEC环境常常托管多个共址的XR流量流,这些流量的特性随时间演变,形成复杂的队列间依赖关系,当前调度器未能捕捉这些差异。应对这些挑战需要适应性、去中心化调度机制,能够在不同流量条件下协调多个TSN队列。为此,本文提出了一种多智能体强化学习(MARL)TSN调度框架,每个TSN队列都被建模为自主代理。异构代理近端策略优化(HAPPO)算法被用来显式建模代理间依赖关系,并联合优化跨队列的服务交付。模拟结果表明,该方法平均帧等待时间可减少多达26.8%,最坏情况下延迟约16.8%,凸显其在动态XR驱动MEC场景下的有效性。
Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability
部分可观测性下稳健的UUV导航统一规划-学习框架
- Authors: Md Ether Deowan, Eleni Kelasidi
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.05365
- Pdf link: https://arxiv.org/pdf/2608.05365
- Abstract
This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboard sonar and depth image observations, adapts a clearance-constrained global planner (GP) to provide long-horizon structure, and integrates a reinforcement learning (RL) policy to handle short-range tracking and reactive avoidance. To further support decision-making under partial observability, the system learns a compact latent state representation from onboard sensor data, encoding environmental structure, obstacle dynamics, and uncertainty. Behavior tree (BT) distillation with staged supervision is introduced to improve safety and training stability, while an uncertainty-calibrated distillation mechanism reweights teacher guidance using online latent-model uncertainty, emphasizing uncertain regimes during learning, with time-to-collision (TTC) and clearance cues remaining explicit in planning and local policy features. To demonstrate the efficacy of the framework, a reproducible multi-seed evaluation protocol is established in high-fidelity GPU-accelerated simulation using NVIDIA Isaac Sim, and performance is benchmarked against BT-only and standard RL baselines. The results obtained demonstrate improved robustness and safety under dynamic conditions, thus providing a general pipeline with a unified hybrid planning learning architecture and a reproducible methodology for robust UUV autonomy under partial observability.
- 中文摘要
本文提出了一个仅观察的自主性框架,用于动态水下环境中无人水下载具(UUV)导航,集成了持续占用映射、全球许可感知规划和风险感知的局部控制。拟议的管线仅根据舰载声纳和深度图像观测构建占用地图,采用受许可限制的全局规划器(GP)以提供长视野结构,并整合强化学习(RL)策略以处理短程跟踪和被动规避。为了进一步支持部分可观测性下的决策,系统从机载传感器数据中学习紧凑的潜在状态表示,编码环境结构、障碍动力学和不确定性。引入了带有分阶段监督的行为树(BT)蒸馏,以提高安全性和培训稳定性,同时采用不确定性校准的提炼机制,利用在线潜在模型不确定性重新加权教师指导,强调学习过程中的不确定性,同时在规划和本地政策特征中明确指出碰撞时间(TTC)和清除线索。为验证该框架的有效性,采用NVIDIA Isaac Sim在高精度GPU加速仿真中建立了可重复的多种子评估协议,并以纯光线和标准强化学习基准测试了性能。结果显示了在动态条件下更强的鲁棒性和安全性,从而提供了一套统一的混合规划学习架构和可重复的方法论,支持部分可观测性下稳健的UUV自主性。
IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
IFlowNets:扩展生成采样器以学习不完整信息博弈中的策略
- Authors: Conor M. Artman, Nicholas Di, Scott Perkins
- Subjects: Subjects:
Machine Learning (cs.LG); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2608.05422
- Pdf link: https://arxiv.org/pdf/2608.05422
- Abstract
While many algorithms blend reinforcement learning (RL) with counterfactual regret (CFR) methods to leverage tradeoffs in computational speed and performance, there are fewer investigations into generative sampling frameworks in game theoretic applications in incomplete information games. We extend a generative flow network framework, Adversarial Flow Networks (AFlowNets), to incomplete information games, called Information Flow Networks (IFNs). We prove that previously established constraints for generative flow networks in complete information games are inadmissible for obtaining valid densities (corresponding to player strategies) and a valid training objective. We show that our proposed generalization, IFlowNets, alleviates this issue and strictly generalizes AFlowNets. In preliminary results for three standard game environments, IFlowNets perform comparably to or better than Outcome Sampling Monte Carlo Counterfactual Regret (OSMCCFR) and standard RL-based methods in performance and speed.
- 中文摘要
虽然许多算法结合强化学习(RL)和反事实遗憾(CFR)方法来利用计算速度和性能的权衡,但关于生成抽样框架在博弈论中不完全信息博弈应用的研究较少。我们扩展了一个生成流网络框架——对抗流网络(AFlowNets),用于不完整的信息博弈,称为信息流网络(IFNs)。我们证明了,在完全信息博弈中,生成流网络的先前约束对于获得有效密度(对应玩家策略)和有效训练目标是不可接受的。我们证明我们提出的推广方法IFlowNets缓解了这一问题,并严格推广了AFlowNets。在三种标准游戏环境的初步结果中,IFlowNets在性能和速度上与蒙特卡洛反事实遗憾采样(OSMCCFR)和基于强化学习的标准方法相当或更好。
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
搜索辅助联合代理-环境强化学习,实现带旋转的多智能体终生路径寻觅
- Authors: He Jiang, Jingtian Yan, Yulun Zhang, Yimin Tang, Tanishq Duhan, Rishi Veerapaneni, Guillaume Sartoretti, Jiaoyang Li
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2608.05588
- Pdf link: https://arxiv.org/pdf/2608.05588
- Abstract
Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones. While many learning-based planners have been proposed for LMAPF, most rely on oversimplified kinematic assumptions that may overlook motion constraints critical to real-world performance. In this work, we study a more realistic LMAPF model derived from many real-world automated warehouse systems, termed LMAPF-R2, which incorporates robust safety constraints and in-place rotation constraints. These constraints substantially increase coordination difficulty, particularly in highly constrained spaces. To address these challenges, we propose Search-Aided Joint Reinforcement Learning (SJRL). We first augment neural policies with Causal PIBT, a single-step search-based planner that resolves agents' collisions and propagates their intentions. We then introduce a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search. Experiments demonstrate that SJRL achieves significant improvements over the strong search-based planner, Causal-PIBT, across multiple high-density maps. We further validate SJRL in a challenging mixed-reality warehouse environment with 8 physical robots and 248 virtual robots.
- 中文摘要
终身多智能体路径寻寻(LMAPF)要求反复规划无碰撞路径,供智能体在达到当前目标后不断获得新目标。虽然许多基于学习的规划器被提出用于LMAPF,但大多数依赖于过于简化的运动学假设,可能忽视了对现实性能至关重要的运动约束。在本研究中,我们研究了一个更真实的LMAPF模型,该模型源自许多现实世界的自动化仓库系统,称为LMAPF-R2,它包含了稳健的安全约束和原位旋转约束。这些约束显著增加了协调难度,尤其是在高度受限的空间中。为应对这些挑战,我们提出了搜索辅助联合强化学习(SJRL)。我们首先用因果PIBT来增强神经策略,这是一种单步搜索规划器,可以解决智能体的碰撞并传播他们的意图。随后,我们引入了统一的强化学习表述,联合优化代理和环境策略,环境策略通过学习图边成本,通过向后迪克斯特拉搜索提供全局移动指导。实验表明,SJRL在多个高密度地图上相比基于搜索的强力规划器Causal-PIBT实现了显著提升。我们还进一步验证了SJRL在具有挑战性的混合现实仓库环境中,拥有8台实体机器人和248台虚拟机器人。
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
LC-GRPO:基于流的GRPO与朗之文修正的串联推断间隙桥接
- Authors: Yingqing Guo, Hui Yuan, Zijian He, Mengdi Wang, Zheng Ding
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.05600
- Pdf link: https://arxiv.org/pdf/2608.05600
- Abstract
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training. Although the ODE and SDE share the same marginal distributions in continuous time, their finite-step discretizations can differ substantially. In particular, SDE rollouts often become blurry as the exploration noise increases, creating a mismatch between the samples used for reinforcement learning and those generated by the test-time ODE sampler. We introduce LC-GRPO, a flow-based GRPO framework with Langevin correction. Each rollout transition first takes an inference-aligned ODE Euler step and then applies a stochastic Langevin correction targeting the marginal distribution at the resulting timestep. The required score is recovered directly from the flow velocity, requiring no additional score model, while the resulting transition remains an isotropic Gaussian with a tractable likelihood for policy optimization. We theoretically show that, under suitable conditions, one Langevin correction step reduces the Wasserstein error of an imperfect ODE Euler step. At a matched randomness level, we further show that the proposed transition can be more accurate than the standard Euler--Maruyama discretization of the reverse SDE. Experiments on SD3.5-Medium, FLUX.1-Dev, and HunyuanVideo demonstrate that LC-GRPO consistently improves reward optimization across text-to-image and text-to-video tasks, preserves generation quality, and substantially narrows the gap between stochastic training rollouts and deterministic test-time ODE inference.
- 中文摘要
基于流的生成模型通常通过求解确定性常微分方程(ODE)来抽样,而在线强化学习则需要随机展开以进行策略探索和优化。因此,现有的GRPO流模型方法在训练时用随机微分方程(SDE)替代推断时间常微分方程。尽管常微分方程和单微分方程在连续时间内共享相同的边际分布,但它们的有限步离散化可能有很大差异。特别是,随着探索噪声增加,SDE的展开常常变得模糊,导致用于强化学习的样本与测试时常微分方程采样器生成的样本不匹配。我们介绍了基于流的 GRPO 框架 LC-GRPO,并带有朗之文修正。每个滚动转变先进行一个推理对齐的常微分方程欧拉步骤,然后在该时间步对边际分布施加随机朗之文修正。所需的分数直接从流速中恢复,无需额外评分模型,而所得的跃迁仍是一个各向同性高斯分布,具有可解的策略优化可能性。我们理论上证明,在适当条件下,一个朗之文修正步可以减少不完美常微分方程欧拉步的瓦瑟斯坦误差。在匹配随机性水平下,我们进一步证明所提议的转变比标准欧拉-丸山离散化的逆SDE更为准确。在SD3.5-Medium、FLUX.1-Dev和HunyuanVideo上的实验表明,LC-GRPO在文本到图像和文本到视频任务中持续提升奖励优化,保持生成质量,并显著缩小随机训练推广与确定性测试时间常微分方程推断之间的差距。
ChronoVision: Temporal Reasoning via Latent State Reconstruction
时间视觉:通过潜在状态重建进行时间推理
- Authors: Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.05631
- Pdf link: https://arxiv.org/pdf/2608.05631
- Abstract
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.
- 中文摘要
多模态大型语言模型擅长被动感知,但在复杂的视觉认知任务中表现不佳,这些任务需要多步时间推理。这种退化主要源于基于语言的推理固有的模糊性,常常无法准确表达连续的视觉变换。为此,我们提出了ChronoVision,一种多模态框架,旨在将视觉逻辑与潜在影像对齐。在监督微调过程中,重建视觉头预测最终转换状态的潜在表征,而ROI注意力定位模块则通过语义范围查询聚焦模型的关键视觉证据。在后期培训中,我们采用隐式过程基础机制的强化学习,由复合奖励函数引导,评估结果正确性、潜在过程对齐和无监督视觉聚焦。此外,我们引入了Vbvr-VQA,这是一种新颖的数据集,通过将视频推理重新表述为严格的图像排序任务来评估时间追踪。实验显示,ChronoVision在Vbvr-VQA上实现了最先进的性能,域内准确率为74.8%,域外准确率为71.6%,同时在IntPhys2这一极具挑战性的跨领域基准测试中也达到了55.0%的强劲准确率。
When Agentic AI Meets Integrated Sensing and Communication
当智能人工智能遇上综合感知与通信
- Authors: Kai Li, Conggai Li, Sarah Ali Siddiqui, Syed Sohail Ahmed, Xin Yuan, Shenghong Li, Wei Ni
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.05792
- Pdf link: https://arxiv.org/pdf/2608.05792
- Abstract
Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent surfaces (RIS), edge intelligence, multi-agent coordination, and resilient networking has developed largely in isolation. This survey unifies the literature within a six-stage closed-loop framework comprising observation, contextualization, reasoning and prediction, planning and orchestration, execution and collaboration, and feedback and resilience. It also introduces five levels of agentic maturity, ranging from physical-layer primitives to fully closed-loop agentic ISAC. We use this framework to review advances in multimodal intelligence, large language models, reinforcement learning, federated learning, RIS-assisted control, Unmanned Aerial Vehicle (UAV) and vehicular networks, and AI-native network management, and analyze privacy, security, resilience, and sustainability as cross-cutting requirements of the full perception-reasoning-action loop. An audit of representative studies against nine agentic-specific evaluation criteria shows that no system reports more than one or two of them, exposing a gap between claimed and demonstrated agentic maturity. We identify open challenges in physical-to-semantic grounding, predictive world models, real-time agent-PHY interaction, safe tool use, heterogeneous multi-agent collaboration, benchmarking, and resource-efficient autonomy.
- 中文摘要
智能人工智能(AI)正在将集成感知与通信(ISAC)从功能导向的物理层技术转变为目标驱动的闭环智能系统,这一范式我们称之为AISAC。基于学习的感知、资源分配、可重构智能表面(RIS)、边缘智能、多智能体协调和韧性网络等领域的现有研究大多是孤立发展的。本调查将文献整合在一个六阶段闭环框架内,包括观察、情境化、推理与预测、规划与编排、执行与协作,以及反馈与韧性。它还引入了五个智能体成熟度层级,从物理层原语到完全闭环智能体ISAC。我们利用该框架回顾多模态智能、大型语言模型、强化学习、联邦学习、RIS辅助控制、无人机(UAV)和车载网络,以及AI原生网络管理的进展,并分析隐私、安全、韧性和可持续性作为感知-推理-行动循环的跨领域需求。对代表性研究的九项针对性评估标准的审计显示,没有任何系统报告超过一到两项,暴露出主张与实际主体成熟度之间的差距。我们识别出物理语义基础、预测世界模型、实时代理-PHY交互、安全工具使用、异构多代理协作、基准测试以及资源高效自主等方面的未决挑战。
On-Policy Delta Distillation for Multilingual Math Reasoning
多语言数学推理的策略上Delta提纯
- Authors: Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
- Subjects: Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.05802
- Pdf link: https://arxiv.org/pdf/2608.05802
- Abstract
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.
- 中文摘要
策略提炼(OPD)正作为LLM培训后强化学习的有前景替代方案而兴起,但其在多语言环境中的有效性仍未被充分探索。我们用英语、韩语和日语进行数学推理,研究OPD及其高级变体On-Policy Delta Distillation(OPD$^2$)。OPD$^2$ 通过利用后培训教师与其基础模型之间的概率差距作为学习信号,从而改进 OPD。Qwen3的实验显示,OPD$^2$持续优于原版OPD,韩语和日语表现尤为显著,且总体缩小了英韩性能差距。我们还发现,纯英语的OPD也能提升韩语和日语的表现,但往往会使回答偏向英语,凸显了多语言数据对保持目标语言回答的重要性。
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
M$^3$R-Bench:基于证据的多模态隐喻理解的统一基准
- Authors: Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun, Yuming Yang, Jiang Zhong, Ruirui Chen, Jingman Shi, Hao Wu, Nayu Liu, Xinyi Jiang, Kaiwen Wei
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.05817
- Pdf link: https://arxiv.org/pdf/2608.05817
- Abstract
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual this http URL address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at this https URL.
- 中文摘要
隐喻通过跨领域映射实现对抽象概念的理解,同时传达情感态度。在多模态场景中,视觉和文本信息共同构建目标-源映射,这需要概念理解和跨模态推理。然而,现有基准主要通过孤立子任务评估隐喻理解,缺乏基于证据的解释,这使得评估模型是否建立基于视觉和文本的映射变得困难。http URL解决了这些局限性,我们引入了M$^3$R-Bench,一个统一且基于证据的基准测试,包含1000个图像-文本实例,并配有人工验证的注释。在概念隐喻理论和非字面语言理解理论的指导下,M$^3$R-Bench 提供了隐喻出现、目标-来源映射、情感及“证据识别——建立映射——情感推断'的阶段解释。对M$^3$R-Bench的评估显示,现有模型常常忽视视觉证据,依赖表面文本线索,并产生不准确的目标--来源映射,暴露出跨模态证据--映射不匹配。为解决这一不匹配,我们提出了M$^3$R-推理器,结合基于课程的推理监督与任务感知强化学习,使模型推理与隐喻解释相匹配。实验显示,仅有8B参数骨干的M$^3$R-Reasoner在四个统一任务指标上优于更大型的专有MLLMs,视觉证据和情感正当性得分分别提升28.45分和30.11分,平均评分评分超过Claude-Sonnet-4.68.00分。数据集和代码可在该 https URL 访问。
Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
通过层级推理和话语级目标奖励提升LLM中的社会智力
- Authors: Xiaofeng Wang, Kakam Chong, Shuai Xiao, DeXin Kong, Qingyuan Tian, Chen Ju, Xu Yan, Shuai Zhao, Fei Huang, Rui Wang, Shuguang Han, jufeng chen
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.05832
- Pdf link: https://arxiv.org/pdf/2608.05832
- Abstract
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
- 中文摘要
大型语言模型(LLMs)擅长结构化任务,但在动态社交互动中存在困难,而成功需要长期目标协调和快速适应。当前方法常常对每句话都采用统一的目标奖励,忽视了每个对话回合目标的具体性,也未能考虑潜在策略的合理性。受计划行为理论启发,我们提出了思考-战略-响应(TSR)框架,将社会对话分解为两个层级阶段:高层战略规划和低层语言执行。为了优化TSR,我们引入了带有方差门控奖励的线性层级强化学习(LHRL-VGR),这是一种新颖算法,能够根据目标达成分数的方差动态路由奖励——平衡目标完成与策略遵循。SOTOPIA基准测试的实验显示,我们的方法能使Qwen2.5-7B智能体在目标完成率上超越GPT-40基线7.32%,展现出多智能体社会谈判任务中的最先进性能。
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
AppDeltaWorld:移动图形界面代理的过渡接地Delta代码世界模型
- Authors: Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.05891
- Pdf link: https://arxiv.org/pdf/2608.05891
- Abstract
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive apps and privacy-critical operations. At the same time, existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic. To address these limitations, we propose AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description. AppDeltaWorld retrieves app-specific Level-1 HTML references under an action-transition constraint, generates Level-2 executable HTML conditioned on the current screen, action, predicted next-screen text, and retrieved structure, and inserts generated visual assets into image slots before browser rendering. As a world model, AppDeltaWorld achieves the highest fidelity on CMGUIBench-500 under Code2World evaluation, with clear gains in structural layout and UI element reconstruction over image-only and code-only baselines. As a training environment, AppDeltaWorld supports filtered closed-loop SFT data construction that, when combined with public supervision, enables AppDeltaAgent to achieve state-of-the-art performance on AndroidLens and consistent gains on MobileGym and MobileWorld. Moreover, world-model-based test-time reinforcement learning enables policy adaptation and shows further improvements without additional interaction with real apps.
- 中文摘要
移动图形界面代理可以通过像素感知和触摸操作操作应用,使其成为收集和改进长期移动交互策略的有前景界面。然而,对于敏感应用和隐私关键操作,真实轨迹难以获得。与此同时,现有的模拟环境扩展成本高昂,图形界面世界模型仍存在生成不稳定、模态覆盖有限和动作转换逻辑不一致的问题。为解决这些局限性,我们提出了AppDeltaWorld,一种基于过渡的Delta代码世界模型,预测下一个图形界面为可达的代码更新,而非无约束的图像或文本描述。AppDeltaWorld 在动作-过渡约束下检索应用特定的一级 HTML 引用,生成基于当前屏幕条件的二级可执行 HTML,条件为当前画面、动作、预测的下一屏文本和检索的结构,并在浏览器渲染前将生成的视觉资产插入图像槽。作为世界模型,AppDeltaWorld在Code2World评估下实现了CMGUIBench-500最高的保真度,结构布局和UI元素重建明显优于仅图像和代码基线。作为培训环境,AppDeltaWorld 支持过滤闭环 SFT 数据构建,结合公共监督,使 AppDeltaAgent 能够在 AndroidLens 上实现最先进的性能,并在 MobileGym 和 MobileWorld 上持续提升。此外,基于世界模型的测试时间强化学习使策略适应成为可能,并在无需额外与真实应用交互的情况下展现出进一步改进。
VLMs for Videogame Data Annotation
用于电子游戏数据注释的VLM
- Authors: Katrin Schmid, Iuri Frosio
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.05949
- Pdf link: https://arxiv.org/pdf/2608.05949
- Abstract
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.
- 中文摘要
视觉语言模型(VLM)和人工智能(AI)代理彻底改变了工程师在现实应用中处理复杂问题的方式。但它们在电子游戏中的采用受限于合成场景的极端变异性以及对现实物理的不适应性。本文探讨了VLM在视频游戏帧序列中用奖励信号进行注释的应用,这一任务具有多种潜在应用,包括条件训练和离线强化学习等。我们展示了VLM常常难以回答赛车电子游戏的基本问题(尽管我们在其他游戏类型中观察到类似行为),并讨论了VLM输出混合和提示优化等对策。我们还展示了输入序列长度、分辨率和问题批处理如何影响注释质量及其令牌消耗。
Training a Conditioned Video Game Agent on a VLM Annotated Dataset
在VLM注释数据集上训练条件化视频游戏代理
- Authors: Katrin Schmid, Iuri Frosio
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.05954
- Pdf link: https://arxiv.org/pdf/2608.05954
- Abstract
Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards generally requires a difficult trial-and-error approach. Lastly, rewards are often sparse and understanding how they eventually affect the learned policy is a non-trivial exercise. To ease these issues we propose annotating a video game dataset with Vision Language Models (VLMs) instructed to extract human defined rewards. We show that offline RL can then be used to train a conditioned agent that responds accordingly to the desired returns and we discuss the difficulties and limitations that emerged in our early experiments.
- 中文摘要
强化学习(RL)是一种强大但远非易用的策略学习技术。在电子游戏的具体案例中,需要访问游戏引擎才能获得训练奖励(例如从环境中收集奖励)。此外,正确识别和加权奖励通常需要复杂的试错方法。最后,奖励往往稀少,理解它们最终如何影响所学政策并非易事。为缓解这些问题,我们提议用视觉语言模型(VLM)标注视频游戏数据集,以提取人类定义的奖励。我们展示了离线强化学习可用于训练条件化智能体,使其能根据期望回报做出相应反应,并讨论了早期实验中出现的困难和局限性。
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
AgentOPSD:用于智能强化学习的递归自蒸馏
- Authors: Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.05987
- Pdf link: https://arxiv.org/pdf/2608.05987
- Abstract
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
- 中文摘要
带有可验证奖励的强化学习(RL)构建了轨迹层面的优势估计,但它常常未能认可决定长期、多回合代理任务结果的关键决策。近期工作引入了特权自提纯用于信用分配,提供了更密集的监督,但目前尚不清楚此类局部信号应如何表示连续信用。我们提出了AgentOPSD,这是一种无批评、递归的代理强化学习中轮级学分分配方法。AgentOPSD将令牌级教师-学生对数概率差距汇总为回合级证据,并在对数赔率空间中递归更新贝叶斯信念状态。这产生了一种原则性的加权方案,将稀疏的结果监督转化为回合级的信用信号,并通过连续状态之间的边际信念修正识别关键转折。该方法与标准策略优化完全兼容,无需额外的批判者或额外的推广。我们在ALFWorld、WebShop和Search-QA上使用Qwen2.5模型在两个尺度(3B和7B)评估AgentOPSD。AgentOPSD优于GRPO和强自蒸馏基线,在ALFWorld上Qwen2.5-7B测试成功率为89.1%。消融研究将这些收益归因于回合级聚合和历史依赖的递归信念更新。
Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
基于观察的自我预测强化学习用于视觉连续控制
- Authors: Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.05989
- Pdf link: https://arxiv.org/pdf/2608.05989
- Abstract
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation space (observation prediction). However, state-of-the-art methods from both categories still struggle on challenging visual control tasks when training data is limited. We posit that relying on either predictive objective alone may be insufficient. In contrast, observation prediction grounds learned representations in observation-level dynamics, but does not directly regularize the temporal predictability of latent representations over extended horizons. In this paper, we propose Observation-Grounded Self-Predictive Representations (OG-SPR), a model-free visual RL algorithm for continuous control that learns representations that are both temporally predictive in latent space and grounded in observation-level dynamics. OG-SPR incorporates two core auxiliary objectives: multi-step latent self-prediction and next-observation prediction. We empirically show that directly imposing latent self-prediction on the shared representation may over-constrain it and does not necessarily improve performance. To address this issue, OG-SPR introduces two lightweight adapters for latent self-prediction, allowing the shared representation to benefit from temporally predictive signals without being forced to directly satisfy the self-prediction objective. Experiments on 28 visual control tasks from the DeepMind Control Suite show that OG-SPR improves aggregate performance over state-of-the-art self-predictive and observation-predictive RL methods, with particularly pronounced gains in challenging domains such as dog and humanoid.
- 中文摘要
从像素中高效采样策略学习是强化学习(RL)中长期存在的挑战。最新的基于动态的表示学习方法通过辅助预测(在潜空间(自预测)或观察空间(观察预测)中学习动态感知表征,显著提高了无模型视觉强化学习的样本效率。然而,这两类最先进的方法在训练数据有限时,在具有挑战性的视觉控制任务上仍然存在困难。我们认为,仅依赖任一预测目标可能都不够。相比之下,观测预测将学习的表征建立在观测层级动力学中,但并未直接规范潜在表征在延伸视野上的时间可预测性。本文提出观察基础自预测表征(OG-SPR),一种无模型的视觉强化学习算法,用于连续控制,学习既能在潜空间中时间预测又基于观察层动态的表征。OG-SPR包含两个核心辅助目标:多步潜在自我预测和下一次观测预测。我们通过实证表明,直接对共享表示施加潜在自我预测可能会过度限制其,且不一定能提升性能。为解决这一问题,OG-SPR引入了两个轻量级的潜在自我预测适配器,使共享表示能够受益于时间预测信号,而无需直接满足自我预测目标。DeepMind 控制套件中对28项视觉控制任务的实验显示,OG-SPR在整体表现上优于最先进的自我预测和观察预测强化学习方法,尤其在犬类和类人生物等具有挑战性的领域上取得了显著提升。
Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation
超越平面政策:具身智能体机器人操作的层级后期培训
- Authors: He Kong, Zengjue Chen, Qi Wang, Qianli Xing, Runliang Niu, Peidong Liu, Jiawei Li, Shiqi Wang, Yi Chang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.05999
- Pdf link: https://arxiv.org/pdf/2608.05999
- Abstract
Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-training methods predominantly optimize VLA models as flat policies, making it difficult to explicitly model task progression and perform robust long-horizon manipulation. Although hierarchical approaches introduce task decomposition, they mainly rely on supervised learning from offline demonstrations and cannot effectively improve execution through online interaction. To address this limitation, we propose Hierarchical Robotic Control (HiRoC), a hierarchical post-training framework that decouples high-level task planning from low-level action execution. The planner decomposes complex tasks into executable subgoals to provide explicit semantic guidance, while the executor continuously improves subgoal-conditioned action generation through reinforcement learning. To enable effective collaboration between the two modules, we further align the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution. Extensive experiments across diverse robotic manipulation benchmarks demonstrate that HiRoC consistently outperforms strong baselines. Comprehensive analyses further validate the effectiveness of hierarchical post-training and the contribution of each key component.
- 中文摘要
视觉-语言-行动(VLA)模型通过利用预训练的视觉语言模型,在机器人操作方面展现出了卓越的能力。然而,现有的训练后方法主要将VLA模型优化为平面策略,这使得显式建模任务进程和执行稳健的长视野操作变得困难。虽然分层方法引入了任务分解,但它们主要依赖于线下演示的监督学习,无法通过在线互动有效提升执行。为解决这一局限性,我们提出了分层机器人控制(HiRoC),这是一种分层的训练后框架,将高层次任务规划与低层次行动执行分离。规划器将复杂任务分解为可执行的子目标,以提供明确的语义指导,而执行者则通过强化学习持续改进子目标条件下的动作生成。为了实现两个模块之间的有效协作,我们在强化学习前进一步将执行者与规划者生成的子目标对齐,从而减少规划与执行之间的分布错位。在多种机器人操作基准测试中进行的大量实验表明,HiRoC始终优于强基线。综合分析进一步验证了层级式培训后的有效性及各关键组成部分的贡献。
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
OneEmo:一种用于情感感知、理解与互动的统一多模态推理模型
- Authors: Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang
- Subjects: Subjects:
Human-Computer Interaction (cs.HC)
- Arxiv link: https://arxiv.org/abs/2608.06013
- Pdf link: https://arxiv.org/pdf/2608.06013
- Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at this https URL.
- 中文摘要
多模态大型语言模型(MLLM)在情商方面展现出了卓越的能力。然而,主流研究主要关注任务特定专精,常忽视任务间协同效应,潜在推理潜力未被充分挖掘。为了弥合这一差距,我们引入了OneEmo,一种能够掌握情绪感知、理解和互动的统一情感通才。为此,我们首先构建了EmoWorld-130K,这是一个综合数据集,通过人机参与工作流程将专业情感知识提炼为显式推理轨迹。对该语料库的监督微调揭示了多任务学习带来的显著互利。其次,为了充分释放潜在的推理潜力,我们提出了Emo-Chord,一种通过统一多任务奖励分配稳定优化的新型强化学习策略。大量实验表明,OneEmo在大多数基准测试中,在类似规模的基线下都能实现最先进的性能。值得注意的是,尽管参数远少于商业模型,OneEmo依然带来了极具竞争力的成果。本文为更可靠、更易理解的情感计算铺平了道路。代码可在该 https URL 访问。
ProDVI: Programmatic Dynamics Priors for Value Network Initialization
ProDVI:用于值网络初始化的程序化动态先验
- Authors: Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.06015
- Pdf link: https://arxiv.org/pdf/2608.06015
- Abstract
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. Existing approaches obtain informative initializations through pre-collected datasets, high-fidelity simulators, or meta-learning over related tasks, but these prerequisites may be difficult to access or even unavailable. In this paper, we propose Programmatic Dynamics Priors for Value Network Initialization (ProDVI), a framework that leverages the commonsense and domain knowledge encoded in large language models to initialize RL agents without relying on these resources. Specifically, ProDVI prompts a code-generating language model to produce executable Python functions that encode coarse hypotheses about environment dynamics. These functions are then used to generate synthetic transitions. Based on these transitions, we construct an auxiliary dynamics prediction objective to pretrain the state-action encoder of the value network in an actor-critic framework. The learned representation provides dynamics-aware inductive biases before online RL begins. Importantly, the generated programs are used only for representation pretraining and are not required to faithfully simulate the target environment. While the generated programs may be inaccurate, their induced initialization can be corrected through online learning from real transitions and rewards. Experiments on OpenAI Gym and DeepMind Control Suite tasks show that ProDVI can effectively improve the sample efficiency of model-free RL algorithms.
- 中文摘要
深度强化学习(RL)以样本效率低著称。一个促成因素是强化学习代理通常是从零开始初始化的,迫使他们通过在线交互获取与任务相关的知识。现有方法通过预先收集的数据集、高精度模拟器或相关任务的元学习获得有用的初始化,但这些前提条件可能难以访问甚至无法获得。本文提出了用于价值网络初始化的程序化动态先验(ProDVI),该框架利用大型语言模型中编码的常识和领域知识,在不依赖这些资源的情况下初始化强化学习代理。具体来说,ProDVI 提示一个代码生成语言模型生成可执行的 Python 函数,编码关于环境动力学的粗略假设。这些函数随后用于生成合成跃迁。基于这些转变,我们构建了一个辅助动力学预测目标,用于在actor-critic框架下预训练值网络的状态-动作编码器。学习到的表征在在线强化学习开始前提供了动态感知的归纳偏差。重要的是,生成的程序仅用于表示预训练,并不需要忠实模拟目标环境。虽然生成的程序可能不准确,但其诱导初始化可以通过在线学习从真实过渡和奖励中纠正。OpenAI Gym和DeepMind Control Suite任务的实验表明,ProDVI能够有效提升无模型强化学习算法的样本效率。
Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference
混合自适应线程调优以缓解高性能强化学习推理中仿真执行瓶颈
- Authors: Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu
- Subjects: Subjects:
Machine Learning (cs.LG); Multiagent Systems (cs.MA); Performance (cs.PF)
- Arxiv link: https://arxiv.org/abs/2608.06025
- Pdf link: https://arxiv.org/pdf/2608.06025
- Abstract
In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling overhead, and reduced throughput. Through empirical analysis, we identify the ratio of task execution time to scheduling time as the key factor determining the optimal thread count. Building on this insight, we propose AutoThread, a hybrid adaptive thread-tuning method for mitigating simulation bottlenecks in RL inference. AutoThread employs a Physics-Informed Neural Operator (PINO) as a thread-count predictor and incorporates a finite-source M/M/1 queueing model to constrain and guide prediction, enabling fast and accurate estimation under dynamic workloads. It further performs load-aware online fine-tuning to compensate for prediction errors and refine resource allocation. Experiments show that AutoThread improves average speedup by 18.4\% over static strategies, achieves average throughput of 1.7x and 1.8x that of XGBoost and Reinforcer, respectively, and reduces execution time by up to 83.8\% compared with state-of-the-art methods. Our code and dataset are publicly available at this https URL.
- 中文摘要
在仿真环路决策系统中,强化学习(RL)推理通常受限于模拟器端的执行开销,因为工作负载高度动态且对运行时线程配置敏感。现有的多线程策略在执行前或执行过程中难以匹配线程资源,导致资源争用、调度开销和吞吐量下降。通过实证分析,我们确定任务执行时间与调度时间的比值是决定最优线程数的关键因素。基于这一见解,我们提出了AutoThread,一种混合自适应线程调优方法,用于缓解强化学习推断中的仿真瓶颈。AutoThread 采用物理知情神经算子(PINO)作为线程计数预测器,并结合有限源 M/M/1 队列模型来约束和指导预测,实现动态工作负载下的快速且准确的估计。它还进一步进行负载感知的在线微调,以补偿预测误差并优化资源分配。实验显示,AutoThread 相比静态策略平均速度提升了 18.4%,平均吞吐量分别是 XGBoost 和 Reinforcer 的 1.7 倍和 1.8 倍,执行时间也比最先进方法减少了高达 83.8%。我们的代码和数据集在此 https URL 公开。
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
从失败中学习:通过硬否定实现统一多模态检索的以检索为中心的CoT
- Authors: Zelong Sun, Jun Wang, Kaicheng Yang, Tiancheng Gu, Ziyong Feng, Zhiwu Lu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.06060
- Pdf link: https://arxiv.org/pdf/2608.06060
- Abstract
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.
- 中文摘要
统一多模态检索旨在识别满足通过异构输入表达复杂用户意图的候选对象。尽管基于大型视觉语言模型(LVLM)的检索器高效且可扩展,但直接编码原始多模态输入常常遗漏细粒度的判别线索,导致语义相似候选者之间产生混淆。最新方法通过生成思维链(Chain-of-Thought,简称CoT)理由来丰富查询表示,从而缓解了这一限制。然而,这种推理通常仅基于查询本身:它解释了查询描述的内容,但不解释检索者误解了什么。我们认为,有效的检索推理应以检索反馈为条件。基于这一见解,我们介绍了UniME-R1,一个嵌入-顾问框架,学习对最初检索的候选人进行推理,并生成以检索为中心的思维链(RC-CoT)。顾问会逐个分析候选人,以识别嵌入者混淆的判别线索。如果目标出现在初始的顶K集合中,UniME-R1会直接重新排序候选人;否则,它生成RC-CoT以精细取回取方向,并用双模嵌入器进行全语料库重取。为训练该框架,我们挖掘硬负面以模拟真实的检索失败,联合优化直接检索和RC-CoT增强检索,并通过监督学习和以检索为导向的强化学习使顾问与检索结果对齐。MMEB-V2及多样多模态反演基准测试的广泛实验表明,UniME-R1在强基线条件下持续提升反演性能。
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
潜在上下文有帮助吗?北极航运中逆向强化学习的受控评估
- Authors: Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw, Jaswanth Kumar, Md Mahbub Alam, Gabriel Spadon
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2608.06105
- Pdf link: https://arxiv.org/pdf/2608.06105
- Abstract
Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. Inverse reinforcement learning (IRL) provides a framework for recovering such rewards from vessel trajectories, while recent meta-IRL methods introduce latent context variables to capture behavioral heterogeneity. However, it remains unclear whether these latent representations recover genuinely hidden preferences or simply re-encode information already available in the observed state. We conduct a controlled evaluation on 3,186 AIS-derived voyages from 202 vessels across nine Arctic shipping seasons, comparing a linear shared reward, a nonlinear shared reward, and a latent-context model built on the same nonlinear architecture. The nonlinear reward improves held-out likelihood by 50.9% over the linear baseline, whereas adding vessel-specific latent context reduces performance by 16.5%. Behavioral analysis, context probes, and a pre-registered feature-hiding ablation show that apparent vessel-level variation is largely explained by observable route and environmental conditions rather than hidden vessel-specific factors. Moreover, predictive accuracy, route fidelity, and reward transfer yield different model rankings, demonstrating that no single metric is sufficient to evaluate learned rewards. These findings motivate testing whether the observed route, environmental, and vessel features already explain behavioral variation before adding per-vessel latent context. This supports more trustworthy AI deployment in safety-critical domains.
- 中文摘要
人工智能(AI)辅助导航可以帮助北极航运适应快速变化的海冰状况,但可靠的部署需要可解释且能适应环境变化的奖励模型。逆向强化学习(IRL)为从血管轨迹中恢复此类奖励提供了框架,而最新的元IRL方法则引入潜在上下文变量以捕捉行为异质性。然而,目前尚不清楚这些潜在表征是否恢复了真正隐藏的偏好,还是仅仅重新编码了观察到状态中已有的信息。我们对202艘船只、跨九个北极航运季节的3186次AIS衍生航程进行了受控评估,比较线性共享奖励、非线性共享奖励和基于同一非线性架构的潜在上下文模型。非线性奖励比线性基线提升了50.9%的保留可能性,而增加血管特异性潜伏情境则降低16.5%的性能。行为分析、情境探针和预注册的隐藏特征消融显示,表面血管级变化主要由可观察的路径和环境条件解释,而非隐藏的血管特异因素。此外,预测准确性、路线忠实度和奖励转移会给出不同的模型排名,表明没有单一指标足以评估已学到的奖励。这些发现促使在加入每条血管潜在上下文之前,测试观察到的路径、环境和血管特征是否已经解释了行为变异。这支持了在安全关键领域更值得信赖的人工智能部署。
Contextual Information Policy Optimization for Search Agents
搜索代理的上下文信息策略优化
- Authors: Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.06128
- Pdf link: https://arxiv.org/pdf/2608.06128
- Abstract
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer cor rectness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reason ing: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirma tion bias and inefficient this http URL, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning ac tions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to pre this http URL,CIPOdiscourages evidence-detached guesses and promotes reasoning trajecto ries in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive exper iments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven rea soning and achieves excellent performance on most tasks.
- 中文摘要
搜索代理通过在多步推理中获取和使用外部证据,将大型语言模型扩展到静态参数记忆之外。对于涉及复杂或不断演变信息的知识密集型任务,其可靠性不仅依赖于检索相关EV识别,还依赖于利用这些识别来指导后续推理。然而,现有方法主要奖励最终答案的正确性或中间进展,而未直接评估检索后行为是否基于检索证据。这种错位促使先验驱动的推理ing:代理基于内部知识形成结论,并主要通过检索来确认,导致确认偏差和效率低下。http URL,我们提出了情境信息策略优化(CIPO),一种以证据为导向的强化学习框架,明确将策略优化与外部证据的使用对齐。CIPO为受检索信息影响的推理行为赋予密集的回合级信用,同时将该证据使用信号与全局结果奖励结合,CIPO不鼓励无证据的猜测,促进推理轨迹,使检索到的事实能够指导或修正后续推理。重要的是,CIPO既不需要人工过程注释,也不需要额外的奖励模型。对七个域内外基准测试的广泛经验表明,CIPO降低了先验驱动的rea生成现象,并在大多数任务中表现出色。
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
iARCS:可控3D场景生成的迭代代理强化学习
- Authors: Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.06161
- Pdf link: https://arxiv.org/pdf/2608.06161
- Abstract
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
- 中文摘要
合成三维场景生成越来越多地被用作计算机视觉和具身人工智能的数据源,但现有生成器往往在优化感知真实性时,未能可靠满足关键任务的功能约束。这种不匹配限制了合成数据在下游训练中的价值,因为在这些训练中,可访问性、可遍历性和空间规则合规性往往至关重要。我们介绍iARCS,一种迭代的代理强化学习框架,将预训练场景生成器适配到自然语言任务需求。iARCS采用两阶段策略:普遍奖励预训练以提升物理合理性和布局质量,随后通过LLM生成的奖励程序进行针对任务的微调,这些奖励程序通过训练反馈迭代优化。实验显示,在可行性、可达性和以净空为中心的任务上,约束保真度提升,任务特定约束优化有效,并实现了竞争场景多样性。我们还进一步证明,iARCS生成的数据改进了基础生成器,支持其作为实用合成数据生成工具的价值,而不仅仅是可控的场景编辑方法。
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
EnvACE:通过世界演练内化环境动态以实现智能强化学习
- Authors: Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.06197
- Pdf link: https://arxiv.org/pdf/2608.06197
- Abstract
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at this https URL.
- 中文摘要
用于长期工具的大型语言模型代理训练通常依赖于与真实或合成的可执行环境交互,这些环境的构建和验证成本高昂,或依赖难以建立基础的外部模拟器。我们介绍EnvACE,一种能动强化学习方法,用世界演练替代训练期间的外部环境互动。该政策在行动和排练之间交替进行:首先生成工具调用,然后扮演环境的角色,产生该行动引发的反应,并根据排练后的反应来决定后续决策。这两个角色通过任务成功奖励实现了端到端的联合优化。通过世界演练,该政策将行为与环境反应之间的关系内化在其参数中,形成一个直接支持决策的代理世界模型。在BFCL-v4、tau^2-Bench、VitaBench和FinMCP-Bench上,EnvACE实现了强大且可迁移的性能,在整体评估中优于环境缩放基线。受控研究进一步表明,世界演练在不同模型尺度上持续提升政策学习。在测试阶段,内化世界模型允许在承诺执行前进行私人排练,在适度的排练预算下实现进一步提升,无需额外外部干预。我们的发现确立了世界彩排作为超越外部环境限制扩展LLM代理培训的新途径。我们的代码在此 https URL 公开。
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
DASH:政策内推理模型自我提炼的发散自适应监督视野
- Authors: ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.06243
- Pdf link: https://arxiv.org/pdf/2608.06243
- Abstract
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: this https URL
- 中文摘要
带有可验证奖励的强化学习(RLVR)通过自动验证的结果信号提升大型语言模型的推理能力,但这些信号通常较为稀疏且处于序列层面。策略上自蒸馏(OPSD)通过查询特权教师在学生访问前缀并提供密集的代币级分发监督,来缓解这种稀疏性。尽管这种密集的监督缓解了信号稀疏性,但我们发现标准OPSD仍然未能充分发挥推广的时间结构。它赋予每个局部散度相同的系数,无论其位置或散度序列如何。在政策自回归生成中,相同的发散幅度可能遵循不同的差异历史,反映教师与学生之间不匹配的不同演变。由于仅靠局部标量无法区分这些时间上下文,标准OPSD无法将其标记级权重调整为实现的差异序列。为解决这一限制,我们提出了发散自适应监督视野(DASH)。DASH将每个局部蒸馏信号与序列级均值之间的间隙映射到自适应传播门,然后利用这些门控制向后多步聚合。通过这样做,DASH根据生成过程中局部发散的变化调整代币级监管权重。在三个模型尺度上的三个数学推理基准测试实验显示,DASH在所有三个尺度的基准测试中均优于我们匹配的原版OPSD重运行。DASH重用OPSD已计算的教师和学生分布,因此增益无需额外的教师或学生前向传递。代码:这个 https URL
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
RRC:通过基于排名的奖励构建解锁LLM强化学习中的生成奖励模型
- Authors: Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.06310
- Pdf link: https://arxiv.org/pdf/2608.06310
- Abstract
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at this https URL.
- 中文摘要
近期的奖励建模进展显示,从判别性奖励模型向生成性奖励模型的范式转变。然而,尽管生成奖励模型在反应排名方面表现出色,但它们在强化学习(RL)中的潜力尚未被充分发挥。我们的分析显示,这一局限源于生成奖励建模的比较性与现有强化学习算法采用的标量评分范式之间的不匹配。为弥合这一差距,我们提出了基于排名的奖励构建(RRC)方法,使生成奖励模型能够通过从相对偏好排名中推导奖励,提供更有效的强化学习信号。RRC引入了两种互补策略:自竞争排名,利用抽样回答之间的比较;锚点引导排名,使基于排名的奖励构建能够基于规模,仅使用少量参考回答。开放式聊天和推理基准测试的实验表明,RRC通过生成奖励模型显著提升了强化学习训练,取得了相较现有奖励构建方法的持续提升。我们的代码可以在这个 https URL 找到。
Keyword: diffusion policy
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
SkillMemo:专家指导的技能记忆框架,用于构成具身操作
- Authors: Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.05970
- Pdf link: https://arxiv.org/pdf/2608.05970
- Abstract
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model's current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming $\pi_{0.5}$, while exhibiting strong compositional generalization to unseen task configurations.
- 中文摘要
具身视觉运动模型,包括扩散策略(DP)和视觉-语言-行动(VLA)模型,在机器人操作基准测试中展现出有前景的性能。然而,其潜力仍受制于大规模具象轨迹数据集的稀缺,导致在分发外(OOD)场景中组合推广不足,且捕获可重复使用技能结构的能力有限。为解决这一限制,我们提出了基于技能的记忆(SkillMemo)框架,隐式将长期的演示分解为潜在的原子技能,并将技能层面特征整合进动态的情节记忆库以解决作曲任务。具体来说,我们首先引入了一个基于专家混合(Mixture-of-Experts,MoE)架构的专家引导轨迹分割模块,该模块隐式将轨迹划分为由学习门槛系数表示的不同技能原语。我们还设计了一种技能级情节记忆架构,将紧凑的技能表示作为可检索的键值对存储。在推理过程中,记忆库检索最相关的技能原语,随后与模型当前的门控分布融合,提供稳健的细化前动作预测。在模拟基准测试和实际操作任务上的大量实验表明,SkillMemo 能够持续增强 DP 和 VLA 骨干,实现最先进的性能,并超过 $\pi_{0.5}$,同时对未见任务配置表现出强烈的组合推广。
VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations
VIDP:多样化演示中合规机器人操作的可变阻抗扩散政策
- Authors: Hisham Khalil, Neil Fernandes, Thomas M. Kwok, Hsiu-Chin Lin, Yue Hu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.06210
- Pdf link: https://arxiv.org/pdf/2608.06210
- Abstract
Contact-rich manipulation requires precise tracking and mechanical compliance, where variable impedance control can improve robustness in task success, whereas static compliance cannot adapt to varying contact constraints. Variable impedance skills can be learned from demonstrations, avoiding complex modeling, but compliance is a hidden variable in force-agnostic kinematic data. While existing methods infer compliance from trajectory variations, these variations may reflect geometric adaptation and not intentional compliance when subject to changing spatial layouts. Therefore, this letter introduces Variable Impedance Diffusion Policy (VIDP), an imitation learning-based variable impedance control framework leveraging a Task-Parameterized Directionality-Aware Mixture Model (TP-DAMM) to extract physically consistent trajectory distributions from diverse demonstrations. By mapping distributions to stiffness profiles, VIDP jointly predicts pose actions and task compliance without force sensors. Real-world experiments show that VIDP significantly outperforms fixed-impedance baselines in task success rate while reducing interaction forces with respect to high stiffness controllers and tracking errors with respect to low stiffness baselines.
- 中文摘要
丰富的接触操作需要精确的跟踪和机械顺应性,其中可变阻抗控制可以提高任务成功的鲁棒性,而静态顺应性则无法适应不断变化的接触约束。可变阻抗技能可以通过演示学习,避免复杂建模,但顺应性在力无关的运动学数据中是一个隐藏变量。现有方法通过轨迹变化推断顺应性,但这些变化可能反映几何适应,而非空间布局变化时的有意顺应。因此,本信介绍了可变阻抗扩散策略(VIDP),这是一种基于模拟学习的可变阻抗控制框架,利用任务参数化方向感知混合模型(TP-DAMM)从多种演示中提取物理一致的轨迹分布。通过将分布映射到刚度剖面,VIDP能够在没有力传感器的情况下共同预测姿态动作和任务顺应性。实际实验表明,VIDP在任务成功率上显著优于固定阻抗基线,同时降低高刚度控制器的相互作用力,低刚度基线的跟踪误差。