生成时间: 2026-08-18 16:39:46 (UTC+8); Arxiv 发布时间: 2026-08-18 20:00 EDT (2026-08-19 08:00 UTC+8)
今天共有 57 篇相关文章
Keyword: reinforcement learning
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
何时沟通:多智能体强化学习中原则门控的信念分布与认知分歧
- Authors: Teoman Kaman
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.14559
- Pdf link: https://arxiv.org/pdf/2608.14559
- Abstract
Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I propose a principled alternative: agents communicate only when the KL divergence between their learned belief distributions exceeds a fixed threshold. Each agent maintains a belief distribution over a latent world state computed as a softmax over its LSTM hidden state, and communicates only when belief disagreement is large enough to justify information exchange. I evaluate this approach on the Predator-Prey benchmark from IC3Net \cite{singh2019} across two environment sizes with 5 seeds each, and on MPE simple_spread \cite{lowe2017}, comparing against IC3Net, CommNet, and an independent controller. On PP 10$\times$10, IC3Net outperforms KL-belief at all thresholds. On the harder PP 20$\times$20, a threshold ablation over $\varepsilon \in {0.1, 0.3, 0.5, 1.0}$ reveals an inverted U-shape: $\varepsilon=0.5$ achieves 73.84 average steps and 42\% success rate versus IC3Net's 75.31 steps and 31\%, a gap of 1.47 steps and 11 percentage points with tighter seed variance. On MPE, the belief head improves mean reward by 12 points and reduces variance by 26$\times$ even when gating is inactive, suggesting two orthogonal contributions: principled gating when beliefs can converge, and improved latent representations that benefit coordination regardless.
- 中文摘要
多智能体强化学习中的有效沟通不仅需要智能体决定\textit{什么}进行交流,还需要决定何时进行?现有方法要么在每个时间步通信,要么通过REINFORCE策略梯度\cite{singh2019}学习二元门,这是一种高方差信号,会产生不稳定且无法解释的门控行为。我提出了一个原则性的替代方案:只有当其学习信念分布之间的KL发散超过固定阈值时,代理才进行交流。每个智能体在其LSTM隐藏状态上保持一个潜在世界状态的信念分布,该分布作为软最大值计算,只有当信念分歧足够大以支持信息交换时才进行通信。我在IC3Net \cite{singh2019}的Predator-Prey基准测试中评估了这种方法,涵盖两个环境大小、每个5个种子,以及MPE simple_spread \cite{lowe2017},并与IC3Net、CommNet和独立控制器进行了比较。在PP 10$\times$10中,IC3Net在所有门槛上都优于KL-belief。在较难的PP 20$\times$20上,对$\varepsilon \in {0.1, 0.3, 0.5, 1.0}$进行阈值消融,结果呈现倒U形:$\varepsilon=0.5$平均步数为73.84步,成功率为42%,而IC3Net为75.31步,成功率为31%,种子方差更小,差距为1.47步,差距为11个百分点。在MPE中,信念头即使在门控不激活时也能提升平均奖励12个点,并减少26$\倍数,这表明了两种正交贡献:信念可以收敛时的原则性门控,以及改进的潜在表征,无论如何都有利于协调。
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
SKILL:自纠正知识引导迭代大型语言模型代理,用于逻辑优化
- Authors: Rui Yang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.14579
- Pdf link: https://arxiv.org/pdf/2608.14579
- Abstract
Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. We introduce SKILL, a Self-correcting Knowledge-guided Iterative Large Language Model Agent that unifies multi-agent LLM reasoning and RL-based environment interaction for automated synthesis optimization. SKILL coordinates three specialized LLMs: GPT-4o for strategic planning, Claude Sonnet 4 for detailed reasoning, and Gemini 2.5 Pro for efficient analysis with a PPO-based RL agent that learns actionable policies through direct interaction with synthesis tools. A novel self-correcting module monitors environment feedback (PDA metrics), detects suboptimal behaviors, and invokes LLM-guided recovery strategies. Evaluations on IWLS, OpenCores, and EPFL benchmarks show SKILL achieves a 12.4 % PDA improvement over expert flows and 86.3% success rate on logic systems up to 500K gates.
- 中文摘要
逻辑综合优化面临巨大挑战,原因是搜索空间呈指数增长,奖励信号稀疏,逻辑结构多样。传统的专家设计流程缺乏适应性,而强化学习(RL)方法常常存在样本效率低和可解释性有限的问题。我们介绍SKILL,一款自我纠正知识引导的迭代大型语言模型代理,统一了多智能体的大型语言模型推理和基于强化学习的环境交互,实现了自动化综合优化。SKILL协调三款专用大型语言模型:用于战略规划的GPT-4o,用于详细推理的Claude Sonnet 4,以及基于PPO的强化学习代理的Gemini 2.5 Pro,通过与综合工具的直接交互学习可执行的策略,进行高效分析。一种新型自我纠正模块监控环境反馈(PDA指标),检测次优行为,并调用LLM引导的恢复策略。IWLS、OpenCores 和 EPFL 基准测试的评估显示,SKILL 在 PDA 上比专家流程提升了 12.4%,在高达 500K 门的逻辑系统上成功率为 86.3%。
Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach
城市无线网络中的智能基站部署:一种基于地理数据的数字孪生方法
- Authors: Zhenyu Tao, Yuxuan Li, Wei Xu, Yongming Huang, Xiaohu You
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.14599
- Pdf link: https://arxiv.org/pdf/2608.14599
- Abstract
The placement of base station (BS) is a fundamental determinant of coverage and capacity of urban wireless networks. Yet large-scale BS deployment optimization remains challenging due to its dependency on site-specific radio propagation and user spatial distributions, both of which are unfortunately difficult to obtain prior to deployment. To overcome this barrier, we propose an intelligent BS deployment framework that integrates a geographic data-informed wireless network digital twin (DT) with deep reinforcement learning (DRL), enabling sample-free macro BS deployment optimization from solely open geographic data, without on-site measurements, real user trajectories, or exhaustive ray tracing. The proposed DT incorporates a sample-free radio map prediction model with hybrid input representation to achieve kilometer-scale signal strength estimation in milliseconds, complemented by a diffusion-based generative model for trajectory synthesis to collectively characterize channel and user distributions. Leveraging the DT as a virtual training environment, we formulate BS deployment as a multi-step Markov decision process (MDP) and solve it via a spatially structured DRL algorithm. A local search process and a Wasserstein distance-based deployment buffer are further incorporated to efficiently explore the large combinatorial solution space. Experimental results in real-world urban scenarios demonstrate that the geographic data-informed DT attains accuracy comparable to 100-sample-based prediction, and the intelligent BS deployment framework achieves up to 98.9% of the idealized benchmark performance while reducing optimization overhead by over 99%.
- 中文摘要
基站(BS)的布置是城市无线网络覆盖和容量的根本决定因素。然而,大规模BS部署优化依然具有挑战性,因为它依赖于特定站点的无线电传播和用户空间分布,而这些信息在部署前都很难获得。为克服这一障碍,我们提出了一个智能BS部署框架,该框架将基于地理数据的无线网络数字孪生(DT)与深度强化学习(DRL)集成,实现仅凭开放地理数据实现无样本宏BS部署优化,无需现场测量、真实用户轨迹或详尽光线追踪。所提DT采用无采样无线电图预测模型和混合输入表示,实现毫秒级公里级信号强度估计,辅以基于扩散的生成模型用于轨迹综合,以集体表征信道和用户分布。利用DT作为虚拟训练环境,我们将BS部署构建为多步马尔可夫决策过程(MDP),并通过空间结构的DRL算法求解。进一步集成了局部搜索过程和基于距离的Wasserstein部署缓冲区,以高效探索庞大的组合解空间。在真实城市场景中的实验结果表明,基于地理数据的DT准确度可与基于100样本的预测相当,智能BS部署框架可实现理想基准性能的98.9%,同时优化开销降低超过99%。
Explaining Reinforcement Learning Decisions in Self-adaptive Systems
解释自我适应系统中的强化学习决策
- Authors: Jasmina Gajcin, Juan C. Rosero, Ivana Dusparic
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.14620
- Pdf link: https://arxiv.org/pdf/2608.14620
- Abstract
Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural networks, lack transparency and are difficult to understand. This can lead to diminished user trust, and makes for a more challenging verification of systems. To address this challenge, this paper introduces Explanations using Alternative Realities for Reinforcement Learning (EARL), a Python library to produce counterfactual explanations in RL settings. This library allows the user to produce explanations by exploring What-if scenarios to clarify agent behavior by comparing possible outcomes. Counterfactual explanations have been shown to be intuitive and user-friendly in psychology research, but have only recently been explored in RL, with existing implementations usually limited to toy examples and benchmarks. EARL supports counterfactual explanation generation in realistic RL-based self-adaptive systems. To demonstrate its applicability, we demonstrate its use in a simulation of CitiBikes, a self-adaptive bike-sharing system, and we provide evaluations showing how it performs in real applications.
- 中文摘要
强化学习(RL)已被广泛应用于自主和自*系统,但强化学习策略,尤其是依赖神经网络的深度强化学习政策,缺乏透明度且难以理解。这可能导致用户信任度下降,并使系统验证更具挑战性。为应对这一挑战,本文介绍了“利用替代现实进行强化学习解释”(EARL),这是一个用于在强化学习环境中生成反事实解释的Python库。该库允许用户通过探索假设情景来产生解释,通过比较可能的结果来澄清代理行为。反事实解释已被证明在心理学研究中直观且用户友好,但直到最近才在强化学习中被探索,现有的实现通常仅限于玩具示例和基准测试。EARL支持在现实的基于强化学习的自适应系统中生成反事实解释。为了展示其适用性,我们在CitiBikes(一种自适应共享单车系统)的模拟中演示了其应用,并提供了其在实际应用中的表现评估。
Inference-Time Mitigation of Adversarial Political Bias in Large Language Models
大型语言模型中对抗性政治偏见的推理时间缓解
- Authors: Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.14629
- Pdf link: https://arxiv.org/pdf/2608.14629
- Abstract
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Intelligence (AI). Current model alignment paradigms, such as reinforcement learning from human feedback (RLHF), make LLMs follow overarching safety instructions. However, this instruction tuning can be exploited via adversarial prompt injection and be used to generate unsafe content. In particular, political bias has not been specifically targeted by modern alignment techniques as harmful and biased content. To address this vulnerability of LLMs, we propose mitigation strategies using Chain of Thought (CoT) prompting and Direct Preference Optimization (DPO). Using a public dataset of legislative videos, we generate summaries using LLMs, inject bias via adversarial prompting and evaluate their performance on a four axis scale designed for political summarization. In this paper, we present different methods to shield LLMs against the injection of political bias. Our results demonstrate that the proposed Recursive Self-Correction approach raises model performance from a Political Neutrality Likert scale baseline of 2.14 to 4.56, averaged across all models, demonstrating effective inference-time mitigation of political bias in LLM-generated summaries.
- 中文摘要
随着大型语言模型(LLMs)成为信息检索和摘要任务的主力,确保它们始终保持非党派性且不受政治偏见影响,是迈向更安全、更值得信赖的人工智能(AI)的关键一步。当前的模型对齐范式,如来自人类反馈的强化学习(RLHF),使LLM遵循整体安全指令。然而,这种指令调优可以通过对抗提示注入被利用,并被用来生成不安全的内容。特别是,政治偏见并未被现代阵营技术特别针对为有害和有偏见的内容。为解决LLMs的这一脆弱性,我们提出了利用思维链(CoT)提示和直接偏好优化(DPO)的缓解策略。利用公开的立法视频数据集,我们利用大型语言模型生成摘要,通过对抗提示注入偏见,并在设计用于政治摘要的四轴尺度上评估其表现。本文提出了多种方法,帮助LLM免受政治偏见的渗透。我们的结果表明,所提出的递归自我修正方法将模型表现从所有模型平均的政治中立性李克特量表基线2.14提升至4.56,展示了在LLM生成摘要中有效减少政治偏见的推断时间。
Belayer: Efficient Fault Tolerance for LLM Agentic RL Training
边路者:LLM代理强化学习训练的高效容错
- Authors: Jiecheng Zhou, Qinghao Hu, Peng Sun, Xingcheng Zhang, Weiming Zhang
- Subjects: Subjects:
Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.14635
- Pdf link: https://arxiv.org/pdf/2608.14635
- Abstract
Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment containers whose actions may produce visible side effects, such as file edits, command execution, and dependency installation. A single trajectory can span many rounds of gen- eration and environment interaction, so a component failure can discard completed work or expose the model to an environment state that is inconsistent with its context. However, existing systems lack efficient and correct recovery mechanisms for this distributed execution model. This paper presents Belayer, an efficient fault-tolerant system for LLM agentic RL training. Belayer handles failures in both rollout engines and environment execution while targeting low failure-free overhead. For scoped worker-local rollout failures, Belayer equips each pre-initialized shadow worker with a selective GPU-state reuse protocol that retains independently owned weights and raw KV-arena allocations after owner and GPU health checks, reinitializes worker-local state, and rebuilds request-specific KV contents from logged token prefixes. For environment failures, Belayer introduces full checkpoint and full restore to jointly capture and restore container file-system and runtime state, and coordinates the recovered environment with the LLM context to preserve prefix consistency. An adaptive policy opportunistically overlaps full-state checkpointing with natural LLM inference bubbles when the predicted interval is long enough. Empirical results show low measured overhead during failure-free training, a worker-recovery-time reduction of up to 42 times faster compared with a full engine cold start, and 1.5 to 3.5 times faster recovery from environment failures.
- 中文摘要
大型语言模型(LLM)代理越来越多地接受长期沙盒环境中的强化学习训练。与传统强化学习不同,代理型强化学习将GPU密集型的部署引擎与具状态环境容器结合,这些容器的操作可能产生可见的副作用,如文件编辑、命令执行和依赖安装。单一轨迹可以跨越多轮生成和环境交互,因此组件故障可能导致已完成的工作丢失,或使模型暴露在与其上下文不一致的环境状态中。然而,现有系统缺乏高效且正确的分布式执行模型恢复机制。本文介绍了Belayer,一种高效的容错系统,用于LLM代理式强化学习训练。Belayer 在实现低且无故障的开销的同时,同时处理部署引擎和环境执行的故障。对于有范围的工人-本地部署失败,Belayer 为每个预初始化的影子工作者配备了选择性的 GPU 状态重用协议,该协议在所有者和 GPU 健康检查后保留独立拥有的权重和原始 KV-arena 分配,重新初始化工作者-本地状态,并从记录的令牌前缀重建请求特定的 KV 内容。对于环境故障,Belayer 引入了完整的检查点和全恢复,共同捕获和恢复容器文件系统及运行时状态,并将恢复的环境与 LLM 上下文协调,以保持前缀一致性。当预测区间足够长时,自适应策略机会性地将全状态检查点与自然的LLM推理气泡重叠。实证结果显示,在无故障训练期间测量到的开销较低,工人恢复时间比全发动机冷启动快多达42倍,环境故障恢复速度为1.5至3.5倍。
Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions
在每集分布中培训和评估伦理强化学习代理
- Authors: Prabhjyot Singh, Majid Ghasemi, Mark Crowley
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.14642
- Pdf link: https://arxiv.org/pdf/2608.14642
- Abstract
Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents. An agent can look ethical on average while concentrating its violations in a few bad episodes, and a creature in the environment harmed in one episode is not restored by good conduct in another. We compare four ways of training ethical behavior in Craftax, an open-ended survival benchmark. The four are: scalar penalties with termination, a linear multi-objective weight sweep, an adaptive Lagrangian constraint, and a non-compensatory utility optimized per episode under the Expected Scalarized Returns (ESR) criterion. All are evaluated under a single detector-based protocol that counts every violation in every episode without censoring. On the frontier of mean return against mean violation rate, the four methods are indistinguishable; per episode they separate sharply. At matched mean return, the ESR agent holds its stated budget of one violation in effectively every episode (worst-decile 1.04 +/- 0.07 violations), the Lagrangian leaks past the same budget (1.14 +/- 0.03), and the weight sweep's worst episodes double it (2.20 +/- 0.20). An observation-augmentation control attributes the separation to the training objective rather than to what the agent observes, and the per-episode guarantee costs nothing on the mean frontier. When ethical violations do not average away across episodes, we argue both training and evaluation must target the per-episode distribution rather than the mean.
- 中文摘要
强化学习(RL)智能体在单一奖励信号上训练,利用设计奖励与预期行为之间的差距。当我们试图将伦理行为灌输给强化学习的代理时,这尤其成为问题。一个代理人平均看起来很有道德,但却在少数糟糕的事件中集中其违规行为;而在一次事件中受害的环境中生物,在另一事件中并未因良好行为而恢复。我们比较了Craftax中四种伦理行为训练方式,Craftax是一个开放式生存基准。这四项指标分别是:带终止的标量惩罚、线性多目标权重扫过、自适应拉格朗日约束,以及在预期标量收益(ESR)准则下每集优化的非补偿效用。所有这些都采用基于检测器的单一协议评估,该协议在每集中统计每一次违规且不进行审查。在平均回报与平均违规率的边界上,这四种方法无法区分;每集他们分开得很明显。在匹配平均回报下,ESR特工几乎每集都保持其规定的一次违规预算(最坏十分位1.04 +/- 0.07次违规),拉格朗日量泄漏超过相同预算(1.14 +/- 0.03),权重扫描中最严重的集数会翻倍(2.20 +/- 0.20)。观察-增强控制将分离归因于训练目标,而非代理观察到的内容,且每集保证在平均前沿不会产生任何成本。当伦理违规在各集数之间没有平均化时,我们认为培训和评估都必须针对每集的分布,而非平均值。
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
通过离线强化学习发现高质量的国际象棋谜题
- Authors: Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, Chris Piech
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.14851
- Pdf link: https://arxiv.org/pdf/2608.14851
- Abstract
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to engage in different thinking patterns. In some domains, such as chess, puzzles are used to help students practice their skills in calculating the next moves and recognizing known patterns on a board. Giving students a practice set of puzzles to help them learn different modes of thinking is challenging because the teacher needs to carefully balance between different motifs and how many look-ahead steps a student needs to perform. Popular online platforms like this http URL and Lichess offer players millions of puzzles. Unlike chess tactics puzzles procured by human experts, where chess beginners can learn valuable insights, these puzzles are automatically generated and often regarded as having low pedagogical value. These platforms also rely on a heuristic to recommend puzzles to users for practice. Using the user history data over an entire year, a total of 1.5 billion puzzle-solving histories, we learn the pedagogical value of a puzzle and how to automatically choose a set of puzzles to better support chess learners using insights from offline reinforcement learning. We show that using offline policy evaluation, our trained policy has significant impact on beginners with puzzle-solving Elo range of 100--1000, particularly for the group of beginners whose learning growth was stagnant. We also performed a qualitative analysis of the puzzles discovered by our model by collecting annotation ratings from expert chess players. The success of our pipeline shows promise for a future where we can understand the pedagogical values of practice items given general user interaction data.
- 中文摘要
学习和掌握技能需要大量且刻意的练习。在许多学习环境中,制作高质量的教学材料可能需要高度的领域专业知识,且耗时甚长。教学材料通常需要训练学生参与不同的思维模式。在某些领域,如国际象棋,谜题被用来帮助学生练习计算下一步动作和识别棋盘上的已知模式。给学生一套练习谜题以帮助他们学习不同的思维模式很有挑战性,因为教师需要在不同主题和学生需要完成多少前瞻性步骤之间取得平衡。像这个http URL和Lichess这样的热门在线平台为玩家提供了数以百万计的谜题。与由人类专家获取的国际象棋战术谜题不同,这些谜题是自动生成的,通常被认为教学价值较低。这些平台还依赖启发式系统,推荐谜题给用户练习。通过全年用户历史数据,共计15亿条解谜历史,我们了解了谜题的教学价值,以及如何自动选择一组谜题,更好地支持国际象棋学习者,利用离线强化学习的洞见。我们通过离线策略评估表明,我们训练有素的策略对解谜等级范围为100-1000的初学者有显著影响,尤其是对学习增长停滞的初学者群体。我们还通过收集专家棋手的注释评分,对模型发现的谜题进行了定性分析。我们流程的成功为未来展望了希望,届时我们可以在一般用户交互数据下理解实践项目的教学价值。
Deep Reinforcement Learning for 6G AI-RAN: A Comprehensive Survey
6G AI-RAN深度强化学习:一项全面调查
- Authors: Jie Lu, Peihao Yan, Qijun Wang, Ruxin Lin, Huacheng Zeng
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Signal Processing (eess.SP)
- Arxiv link: https://arxiv.org/abs/2608.14877
- Pdf link: https://arxiv.org/pdf/2608.14877
- Abstract
The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous services, dynamic environments, and competing performance objectives. Open Radio Access Network (O-RAN) provides the open interfaces, disaggregated architecture, and multi-timescale control loops needed to support this transformation, while deep reinforcement learning (DRL) offers a natural framework for optimizing sequential decisions under uncertainty. However, existing surveys either address artificial intelligence (AI) and machine learning (ML) in O-RAN broadly or focus on isolated DRL use cases, leaving a gap in the systematic connection between DRL methodology, O-RAN architecture, and operational deployment. To the best of our knowledge, this article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN. We review the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provide an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure. We classify DRL applications across radio resource management, mobility management, interference control, traffic steering, energy efficiency, network slicing, integrated sensing and communication, security, and massive MIMO. We further examine multi-agent and federated coordination, foundation models and agentic AI, trustworthy DRL, sim-to-real transfer, continual adaptation, resource-efficient inference, and reinforcement learning operations. Finally, we review experimental platforms, benchmarks, standards, and industry activities, and identify research directions toward sample-efficient, safe, scalable, interoperable, and deployable DRL control for 6G Open AI-RAN.
- 中文摘要
向第六代(6G)网络的演进正在将无线接入网(RAN)转变为一个可编程且智能的控制平台,必须不断适应异构服务、动态环境和竞争的性能目标。开放无线接入网(O-RAN)提供了支持这一转型所需的开放接口、分解架构和多时间尺度控制循环,而深度强化学习(DRL)则为在不确定性下优化顺序决策提供了自然框架。然而,现有调查要么广泛关注O-RAN中的人工智能(AI)和机器学习(ML),要么只关注孤立的DRL用例,导致DRL方法论、O-RAN架构与运营部署之间的系统性联系存在空白。据我们所知,本文呈现了Open AI-RAN首次专门且全面的DRL综述。我们回顾了无模型、基于模型、离线、安全、多智能体、联邦和迁移学习的基础,并提供了一个基于O-RAN的框架,通过状态、观察、动作、奖励、约束和时间结构来构建RAN控制问题。我们将DRL应用分类涵盖无线资源管理、移动管理、干扰控制、交通引导、能源效率、网络切片、集成感测与通信、安全以及大规模MIMO。我们还进一步探讨了多智能体与联邦协调、基础模型与智能人工智能、可信的多重学习学习、模拟到现实转移、持续适应、资源高效推理以及强化学习操作。最后,我们回顾了实验平台、基准、标准和行业活动,并确定了6G开放AI-RAN样本高效、安全、可扩展、互操作且可部署的DRL控制的研究方向。
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
指令空间反事实解释帕累托条件强化学习
- Authors: Joanikij Chulev, Hendrik Baier
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
- Arxiv link: https://arxiv.org/abs/2608.14963
- Pdf link: https://arxiv.org/pdf/2608.14963
- Abstract
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN explanations as return-command interventions, using a return-only PCN variant that avoids the added ambiguity of horizon-conditioning. Second, we adapt adversarial machine learning methods to reinforcement-learning explanations. Third, we introduce a boundary-seeded directional search that improves over purely local optimization in the command-action landscape, resulting in our proposed approach CF-ZOO. The resulting explanations are actionable and intuitively expressed in the user's own preferences: "If your trade-off had shifted slightly towards X, the agent would have chosen Y."
- 中文摘要
帕累托条件网络通过对期望的返回命令进行单一策略条件,学习多目标强化学习行为。然而,从命令和状态到行动的本地映射仍然不透明。我们提出了对PCN的命令空间反事实解释:给定固定状态、原始命令和箔行动,我们在黑箱环境中搜索一个最小修改的期望返回命令,在该命令下相同训练策略会选择对照。我们的贡献有三方面。首先,我们将PCN解释表述为返回指令干预,采用仅返回的PCN变体,避免地平线条件反射带来的额外歧义。其次,我们将对抗性机器学习方法应用于强化学习的解释。第三,我们引入了边界种子方向搜索,在命令-动作景观中超越纯局部优化,形成了我们提出的方法CF-ZOO。由此产生的解释是可操作的,并且直观地表达在用户自己的偏好中:“如果你的权衡稍微偏向X,代理会选择Y。”
MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems
元理性:通过编辑元信息解决几何问题,精确交错多模态推理
- Authors: Penghao Yin, Haomin Wang, Qihong Tang, Xiaoye Qu, Hongjie Zhang, Xiao-Ping Zhang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
- Arxiv link: https://arxiv.org/abs/2608.15006
- Pdf link: https://arxiv.org/pdf/2608.15006
- Abstract
Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning. Some recent methods introduce intermediate visual states to facilitate reasoning, but they are often hindered by inaccurate geometric representations and low rendering fidelity, ultimately leading to unreliable outputs. To address these limitations, we propose MetaReason, a framework for multimodal reasoning in plane geometry that leverages structured meta-information to enable accurate auxiliary-line construction. The framework first parses geometric images into meta-information, performs controllable edits with predefined tools to synthesize high-fidelity visual states, and then conducts reasoning based on these augmented views. To support this framework, we construct TutorGeo, a comprehensive dataset containing 17k image-to-meta conversion samples, 60k text-only reasoning traces, and 60k interleaved multimodal reasoning traces. Using this dataset, we combine supervised fine-tuning and reinforcement learning to develop robust multimodal reasoning capabilities. We also introduce ExamGeo, a benchmark derived from real-world examination problems that enables systematic evaluation across varying difficulty levels. Experimental results demonstrate that MetaReason significantly outperforms existing open-source models and achieves competitive performance against proprietary models.
- 中文摘要
尽管视觉推理对于解决复杂几何任务至关重要,但现有的视觉语言模型高度依赖纯文本推理。一些最新方法引入中间视觉状态以促进推理,但常因几何表现不准确和渲染精度低而受阻,最终导致输出不可靠。为解决这些局限性,我们提出了MetaReason,这是一个利用结构化元信息实现辅助线构造的平面几何多模推理框架。该框架首先将几何图像解析为元信息,使用预定义工具进行可控编辑以合成高保真度的视觉状态,然后基于这些增强视图进行推理。为支持该框架,我们构建了 TutorGeo,这是一个包含 17k 图像到元转换样本、60k 纯文本推理痕迹和 60k 交错多模态推理痕迹的综合数据集。利用该数据集,我们将监督式微调和强化学习结合起来,开发出稳健的多模态推理能力。我们还引入了ExamGeo,这是一个基于真实考试题目的基准测试,能够在不同难度层级下进行系统性评估。实验结果显示,MetaReason 远远优于现有的开源模型,并在与专有模型竞争中取得竞争力。
LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning
基于LLM的层级协调控制与持续感知策略学习
- Authors: Changhong He, Jinda Gao, Xinkuan Liu, Le Zhang, Xizi Luo, Yu Mei
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15041
- Pdf link: https://arxiv.org/pdf/2608.15041
- Abstract
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
- 中文摘要
在复杂工程系统中协调多个交互单元具有挑战性,因为系统交互难以建模,操作信息异构,且低层次动作必须满足严格的约束条件。我们提出了基于LLM的分层框架,LLM基于异构操作上下文协调交互单元,而任务专属控制器或优化器生成可执行且约束感知的动作。我们进一步引入了延续感知GRPO,以捕捉协调决策在后续控制区间内的影响。该方法不仅仅根据决策的即时结果来评判,还评估了系统在当前政策下之后的演变。我们验证了该框架在多匝道交通控制和虚拟电站(VPP)能源管理方面,使用简化的系统模型进行训练,并使用更真实的模拟器进行评估。在这两种任务中,所提方法始终优于直接针对任务的控制与优化、端到端强化学习、基于规则和基于强化学习的层级协调,以及仅提示的LLM协调器,展示了异构上下文推理、层级执行和延续意识策略学习的价值。
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
Max-Q 人类在线机器人学习的选择性模仿
- Authors: Zihang Wang, Yishan Wang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15088
- Pdf link: https://arxiv.org/pdf/2608.15088
- Abstract
Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior. We present a training method for this setting based on two components. First, an \emph{MC Q-chunk} critic regresses chunk-level action values onto Monte Carlo returns from the replay buffer, performing sample-average (behavior) policy evaluation so that intervention trajectories are credited directly rather than diluted by current-policy TD backups. Second, \emph{max-Q selective imitation} updates the actor by imitating, at each state, the higher-$Q$ action between the current policy action and a buffer sample under a hard winner-take-all rule. This rule automatically switches between learning from interventions and on-policy self-improvement: when the autonomous policy is stronger, targets align with the policy distribution, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift. In practice we score candidates with a standard critic ensemble mean to reduce comparison noise, without softening targets or introducing score-gap thresholds. On a real USB pick-and-insertion task with 20 demonstrations, ACT QChunk-MCBC attains 99\% success within 30 minutes of HIL training, whereas HIL-SERL requires about 5 hours to converge. In simulation on Peg Insertion and Square, ACT/Flow Q-chunk variants similarly reach $\ge$96\% success within roughly half an hour of effective training, outperforming HIL-SERL, EXPO, and E2HiL on the success--time frontier.
- 中文摘要
真人机器人的在线强化学习必须迅速吸收人类干预,同时持续超越人类。我们基于两个组成部分,提出了针对该环境的训练方法。首先,\emph{MC Q-chunk} 批评者将区块级的行动值回归到回放缓冲区的蒙特卡洛返回上,进行样本平均(行为)策略评估,从而直接归功干预轨迹,而非被当前政策的TD备份稀释。其次,\emph{max-Q 选择性模仿}通过在每个状态模拟当前策略动作与缓冲区样本之间较高$Q$的动作来更新演员,采用硬性赢家通吃规则。该规则自动在从干预中学习和政策内自我提升之间切换:当自主策略更强时,目标与政策分布一致,减少导致执行时间分布偏移的政策-目标-样本差距。在实践中,我们用标准的批评集合平均值给候选人评分,以减少比较噪声,同时不软化目标或引入分数差距阈值。在真实的USB拾取插入任务中,包含20个演示,ACT QChunk-MCBC在HIL培训30分钟内成功率达99%,而HIL-SERL则需约5小时才能完成。在Peg Insertion和Square的模拟中,ACT/Flow Q块变体在有效训练约半小时内达到$$96\%,在成功时间前沿上优于HIL-SERL、EXPO和E2HiL。
StructRL: Structured Action-Space Exploration for Flow-Based VLAs
StructRL:基于流的VLA的结构化动作空间探索
- Authors: Jiarui Yang, Bin Zhu, Jingjing Chen, Na Zou, Yanwei Fu, Jianggang Zhu, Yu-Gang Jiang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.15139
- Pdf link: https://arxiv.org/pdf/2608.15139
- Abstract
Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for adapting them to new tasks. Existing RL methods typically inject stochasticity inside the denoising chain, often through isotropic or temporally independent noise. However, effective robot exploration calls for structured noise: temporally smooth and scaled differently across action groups. We show that simply switching the in-chain noise to a structured form does not suffice: noise added at an intermediate flow time can be weakened by the remaining denoising steps before execution, a phenomenon we call \emph{Structured Noise Dilution}. We propose \textbf{StructRL}, which avoids dilution by relocating policy stochasticity to the action space via three coupled choices: (i) a deterministic ODE decoder, (ii) structured noise injected directly in the action space, and (iii) last-step replay, where policy-gradient updates avoid assigning likelihoods to intermediate denoising states. This keeps structured exploration tied to the executed action while providing a tractable training signal for the flow decoder. Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL. \textbf{Project page:} this https URL
- 中文摘要
基于流程的视觉-语言-行动(VLA)模型现已被广泛用于连续机器人操作,在线强化学习(RL)正作为适应新任务的关键技术。现有的强化学习方法通常在去噪链中注入随机性,通常通过各向同性或时间无关噪声。然而,有效的机器人探索需要结构化噪声:在时间上平滑,且在不同动作组之间以不同的尺度进行。我们证明,仅仅将链内噪声转换为结构化形式是不够的:在中间流时间添加的噪声可以通过执行前剩余的去噪步骤减弱,这种现象我们称之为\emph{结构化噪声稀释}。我们提出了 \textbf{StructRL},通过三种耦合选择将策略随机性重新定位到动作空间,避免稀释:(i) 确定性常微分方程解码器,(ii) 直接注入动作空间的结构化噪声,以及 (iii) 最后一步重放,策略梯度更新避免将似然分配到中间去噪状态。这既使结构化的探索与执行动作相关联,又为流解码器提供了可处理的训练信号。通过三个基于流的VLA模型,基于多个模拟操作基准和两个真实任务,StructRL相较于以往链内基线提升了探索效率和OOD性能,展示了结构化动作空间探索在适应基于流的VLA与强化学习中的有效性。\textbf{项目页面:} 这个 https URL
PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search
PureTD:无评估时间搜索的双陆棋奖金游戏强化学习
- Authors: Alexander L. Strehl
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.15146
- Pdf link: https://arxiv.org/pdf/2608.15146
- Abstract
We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL), with minimal hand-coded logic and no expert features. In this setting, we demonstrate that pure self-play RL suffices to train models that reach near-state-of-the-art playing strength. Specifically, for cubeful money games, our search-free model evaluates faster and is substantially stronger than the open-source engines GNU Backgammon and Open Sage running a one-move (1-ply) look-ahead search.
- 中文摘要
我们在无评估时间搜索的环境下重新审视Tesauro的TD-Gammon,用于双陆棋奖金游戏。跳棋游戏和魔方动作(使用加倍魔方)都是通过自我对战强化学习(RL)从零学习的,几乎没有手工编码的逻辑,也没有专家功能。在此环境中,我们证明纯自玩强化学习足以训练模型达到接近最先进的游戏实力。具体来说,对于立方体货币游戏,我们的无搜索模型评估更快,远胜于运行单步(1层)前瞻搜索的开源引擎GNU Backgammon和Open Sage。
LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset
LAPF:基于 UAVScenes 数据集的基于 LLM 代理的路径查找器
- Authors: Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Gutiérrez Gaitán, Atefeh Hajijamali Arani, Rui Zhang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15175
- Pdf link: https://arxiv.org/pdf/2608.15175
- Abstract
Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making. Existing optimization-based, Machine Learning (ML), and Reinforcement Learning (RL) approaches often rely on predefined models or task-specific training, limiting their generalization and adaptability in uncertain scenarios. Recent Large Language Model (LLM)-assisted approaches offer promising reasoning capabilities but remain constrained by limited agentic functionality, including insufficient memory, planning, and tool interaction this http URL paper proposes an LLM-Agent-Based Path Finder (LAPF) framework for autonomous UAV navigation in town-scale outdoor environments. LAPF extends LLM-assisted navigation by integrating perception, memory, planning, and action modules into a closed-loop cognitive architecture. The proposed agent leverages prior navigation experiences, performs Chain-of-Thought (CoT) reasoning, couples each detected hazard to a bounded corrective action, and dynamically refines waypoint decisions based on environmental this http URL three independent trials per method demonstrate that LAPF achieves mean path lengths of 512.83 m and 506.37 m, compared to the straight-line optimum of 497.33 m, corresponding to path length reductions of 17.2% and 15.6% relative to CoT prompting and absolute path efficiencies of 97.1% and 98.1% in open-field and obstacle-injected scenarios, respectively. Furthermore, LAPF is the only evaluated approach that couples every detected hazard to a bounded, metric-neutral corrective action while maintaining near-goal stability, with zero clamp events in both scenarios, whereas CoT prompting increases from 9.7 to 14.0 events.
- 中文摘要
无人飞行器(UAV)越来越多地被部署用于复杂户外环境的自主导航,这些环境在动态条件和任务需求下需要智能自适应决策。现有基于优化的机器学习(ML)和强化学习(RL)方法通常依赖预定义模型或任务专属训练,限制了其在不确定场景下的泛化性和适应性。近期大型语言模型(LLM)辅助方法提供了有前景的推理能力,但仍受限于有限的代理功能,包括内存不足、规划和工具交互不足。本文 http URL 论文提出了一个用于城镇级户外环境中自主无人机导航的 LLM 代理路径查找器(LAPF)框架。LAPF通过将感知、记忆、规划和行动模块整合进闭环认知架构,扩展了LLM辅助导航。该代理利用以往的导航经验,进行思维链(CoT)推理,将每个检测到的危害与有界纠正措施结合,并基于环境动态优化航点决策。每种方法的三次独立试验表明,LAPF的平均路径长度分别为512.83米和506.37米,而直线最优距离为497.33米。 相较于CoT提示,路径长度缩短为17.2%和15.6%,在开阔地和注入障碍物场景中,绝对路径效率分别为97.1%和98.1%。此外,LAPF是唯一一种将所有检测到的危害与有界、度量中立的纠正措施结合起来的方法,同时保持接近目标的稳定性,且两种情景均无钳制事件,而CoT提示事件则从9.7增加到14.0。
Temporal Logic Guided Universal Task Representations for Reinforcement Learning
强化学习的时序逻辑引导通用任务表示
- Authors: Hao Zhang, Zhangli Zhou, Zhen Kan
- Subjects: Subjects:
Robotics (cs.RO); Formal Languages and Automata Theory (cs.FL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.15509
- Pdf link: https://arxiv.org/pdf/2608.15509
- Abstract
Task guided agents demonstrate strong performance in a wide range of complex tasks. However, most existing task representation algorithms are tailored to specific contexts and struggle to generalize across diverse scenarios. Moreover, they typically depend on gradient signals from reinforcement learning controllers to update their weights, which can degrade both representation quality and learning efficiency. To overcome these limitations, we propose LOTUS, a temporal logic inspired universal task representation framework that can be seamlessly integrated into any RL algorithm to enhance agent performance across diverse task settings. Specifically, we design a novel task representation architecture capable of modeling relationships and extracting task semantics from LTL formulas. We further introduce a more effective update mechanism that treats the LTL encoder as a policy, thereby improving representation capacity. To enhance stability and robustness, LOTUS leverages the bisimulation metric, which provides theoretical guarantees for LTL representation, including behavioral equivalence, optimality fidelity, and trajectory robustness. Experimental results show that LOTUS outperforms most existing methods in learning efficiency, generalization capability, and representation quality. Specifically, LOTUS accelerates convergence over 20% in single-task scenarios, achieves a 15%-45% higher success rate in unseen manipulation tasks, and improves generalization performance over 25% in complex multi-task environments with increased sub-goal depth or conjunctions. The corresponding code, videos, and appendix are available at: this https URL.
- 中文摘要
任务引导代理在各种复杂任务中表现出优异的性能。然而,大多数现有的任务表示算法都针对特定情境量身定制,难以在不同场景中泛化。此外,它们通常依赖强化学习控制器的梯度信号来更新权重,这会降低表示质量和学习效率。为克服这些限制,我们提出了LOTUS,一种受时间逻辑启发的通用任务表示框架,可无缝集成到任何强化学习算法中,以提升代理在不同任务环境中的性能。具体来说,我们设计了一种新型任务表示架构,能够建模关系并从LTL公式中提取任务语义。我们进一步引入了一种更有效的更新机制,将LTL编码器视为策略,从而提升表示能力。为了提升稳定性和鲁棒性,LOTUS利用双模拟指标,该指标为LTL表示提供了理论保证,包括行为等价性、最优性忠实度和轨迹鲁棒性。实验结果表明,LOTUS在学习效率、泛化能力和表示质量方面优于大多数现有方法。具体来说,LOTUS在单任务场景中收敛加速超过20%,在看不见的操作任务中成功率提高15%-45%,在复杂多任务环境中子目标深度或连接增加的泛化性能提升超过25%。相应的代码、视频和附录可在以下网站获取:https URL。
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
现在谁来领导?用于图表到代码生成的令牌级模态仲裁
- Authors: Qinghao Fu, Yarong Wang, Shunlei Ning, Yilin Wang, Shunwen Bai, Xinda Wang, Jiaotuan Wang, Yinan Nie, Wei Zhou
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15510
- Pdf link: https://arxiv.org/pdf/2608.15510
- Abstract
Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods either train visual and coding abilities separately, or fine-tune on chart-to-code data with the two abilities entangled. Neither strategy accounts for the distinct nature of the two abilities or the interference that arises when they are optimized together. We propose MoCA (Mixture of Cross-modal Arbitration), which separates the two abilities rather than blending them. MoCA is built on Cross-modal Arbitration Block (CAB), which maintains a visual branch and a code branch as two distinct pathways, and a lightweight arbiter that arbitrates their relative contributions at every layer and generated token. We train MoCA in two stages: a supervised warm-up on self-distilled reasoning trajectories that decomposes visual understanding into explicit steps, followed by reinforcement learning with rewards on both the reasoning process and the final code. Analysis shows that the arbiter learns structured rather than arbitrary allocations, with expert contributions varying systematically across tokens, layers, and instances. Across three benchmarks, MoCA delivers competitive performance against general-domain and chart-specialized models. Ablation results show that the gains cannot be attributed to a larger model size alone, but instead arise from the joint contributions of complementary visual and code branch initialization and input-conditioned arbitration through CAB.
- 中文摘要
图表到代码生成需要模型读取图表的细粒度视觉细节,并编写可执行代码以复现图表。现有的图表到代码方法要么分别训练视觉和编码能力,要么在图表到代码数据上微调,并结合这两种能力。这两种策略都没有考虑到两种能力的独特性质,也没有考虑到它们一起优化时产生的干扰。我们提出MoCA(跨模态仲裁混合),将两种能力分离而非融合。MoCA 建立在跨模态仲裁块(CAB)之上,该块维护视觉分支和代码分支作为两个不同的路径,以及一个轻量级仲裁器,仲裁它们在每一层的相对贡献和生成的代币。我们将MoCA训练分为两个阶段:一个是对自我提炼推理轨迹的监督热身,将视觉理解分解为具体步骤;随后是强化学习,并对推理过程和最终代码给予奖励。分析显示,仲裁者学习的是结构化而非任意的分配,专家贡献在代币、层和实例之间系统性地变化。在三个基准测试中,MoCA在与广域和图表专门化模型中均有竞争力。消融结果表明,这些收益不能仅归因于模型规模的增加,而是来自通过CAB实现的视觉和代码分支初始化与输入条件仲裁的互补贡献。
GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning
GLaQ:将潜在查询建立在视觉证据中,用于多模态推理
- Authors: Zesheng Yang, Lingling Zhang, Xinyu Zhang, Cheng Zhang, Pengyu Li, Heng Wang, Lin Wu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.15517
- Pdf link: https://arxiv.org/pdf/2608.15517
- Abstract
Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning steps. To address this limitation, tool-augmented thinking-with-images methods maintain visual access externally by revisiting or manipulating the image, but require predefined tools and additional inference-time processing. As an internal alternative, continuous visual latent reasoning retains intermediate computation in hidden states. However, its prevailing autoregressive construction makes each latent state depend on its predecessors, so later states may repeat information already present in the latent sequence rather than capture complementary visual details. We introduce GLaQ, a grounded latent-query framework that replaces sequential latent rollout with a fixed set of context-conditioned queries grounded in the original visual tokens. The grounded queries are reinjected for answer generation, providing direct and coordinated access to source visual evidence. We train GLaQ with localized-view supervision followed by reinforcement learning under task-level rewards. Across five benchmarks for fine-grained visual understanding and perception, GLaQ-7B gains 5.99--9.66\% over its base model and leads all compared visual latent methods, suggesting that direct query-to-image grounding can recover localized evidence from the full image without external visual operations or autoregressive latent rollouts.
- 中文摘要
思维链推理极大提升了多模态大型语言模型的问题解决能力。然而,细粒度的视觉证据在基于文本的推理步骤中仍然难以保存和重复使用。为解决这一限制,工具增强式图像思维方法通过重访或操作图像来保持外部视觉访问,但需要预设工具和额外的推理时间处理。作为内部替代方案,连续的视觉潜在推理保留了隐藏状态中的中间计算。然而,其主流的自回归构造使每个潜在状态依赖于其前置,因此后续状态可能会重复潜序列中已有的信息,而非捕捉互补的视觉细节。我们介绍GLaQ,这是一个基于原始视觉标记的固定上下文条件查询框架,取代了顺序潜在的推送。基于基础的查询被重新注入以生成答案,提供对原始视觉证据的直接协调访问。我们先用局部视角监督训练GLaQ,随后在任务级奖励下进行强化学习。在五个细粒度视觉理解和感知基准测试中,GLaQ-7B比其基础模型提升了5.99%-9.66%,领先所有比较的视觉潜在方法,表明直接查询到图像的基础化可以在无需外部视觉操作或自回归潜在展开的情况下,从完整图像中恢复局部证据。
Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback
为什么摘要变得中立:强化学习中从人类反馈中归因的政策归因
- Authors: Mikhail Krasitskii, Alexander Gelbukh, Olga Kolesnikova, Grigori Sidorov
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.15530
- Pdf link: https://arxiv.org/pdf/2608.15530
- Abstract
Reinforcement learning with human feedback (RLHF) aligns LLMs with human preferences, improving summarization fluency and safety, but causes sentiment drift: overly neutral summaries stripped of emotional nuance. We diagnose why RL acts as a sentiment neutralizer and present Policy Attribution, a framework using gradient and logit decomposition to trace drift to reward model (RM) signals and KL (Kullback-Leibler) penalty. Sentiment drift reflects a strategic bias toward "low-risk" tokens maximizing expected rewards under preference uncertainty (Stiennon et al., 2020; Gao, Schulman, and Hilton, 2023). On Reddit TL;DR and CNN/DailyMail, RLHF summaries get higher rewards but show 30-40% lower sentiment variance. Cross-lingual analysis across eight languages shows language-independent drift, with morphologically richer languages more suppressed (Krasitskii et al., 2026). We propose and validate a sentiment-aware regularization technique reducing drift by 18-22% without harming summary quality. The code and toolkit will be public.
- 中文摘要
人类反馈强化学习(RLHF)使LLM符合人类偏好,提高摘要的流畅性和安全性,但也会导致情感漂移:过于中性的总结剥离了情感细腻。我们解释了为何强化学习作为情感中和器,并提出了策略归因框架,这一框架利用梯度和对数分解来追踪漂移到奖励模型(RM)信号和KL(Kullback-Leibler)惩罚。情绪漂移反映了一种战略性倾向于“低风险”代币,在偏好不确定性下最大化预期回报(Stiennon 等,2020;高、舒尔曼和希尔顿,2023年)。在Reddit上,简而言之;DR和CNN/DailyMail、RLHF摘要的回报更高,但情绪波动率降低了30%-40%。跨八种语言的跨语言分析显示出语言无关的漂移,形态丰富的语言则被抑制得更为明显(Krasitskii 等,2026)。我们提出了并验证了一种情感感知规范化技术,可在不损害摘要质量的情况下减少18-22%的漂移。代码和工具包将是公开的。
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
当熵不足以实现:在LLM输出长度预测中重新夺回失去的语义
- Authors: Feiyang Ren, Shengtao Wen, Lingbing Guo, Yu Tian, Yuanning Cui, Xiang Chen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15592
- Pdf link: https://arxiv.org/pdf/2608.15592
- Abstract
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduces the overhead. This advantage is especially pronounced in long-context reasoning and reinforcement learning applications. Existing approaches, such as entropy-guided token pooling, use token-wise entropy as their primary signal, but they tend to ignore differences in semantic content across tokens. So, important tokens are often underweighted, and tokens carrying little information receive disproportionate emphasis. This hurts the reliability of length prediction. We introduce ESTP (Entropy-and-Semantic Token Pooling), a lightweight framework that addresses this issue by combining entropy with attention-based importance scores. These scores are derived directly from the self-attention weights computed during the LLM prefill phase, and this allows ESTP to capture both uncertainty and semantic importance with minimal additional computation. Since the framework reuses prefill activations, it adds almost no extra memory overhead and introduces only minimal latency. On the ForeLen benchmark, ESTP outperforms baseline methods, achieves better prediction accuracy and lower error rates in most scenarios. When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput and reduce the padding ratio. Our results offer a practical and effective building block for length-aware LLM serving systems.
- 中文摘要
高效的LLM服务常常被需要将序列填充到固定的最大长度所限制,这浪费了计算量并降低了吞吐量。提前预测输出长度使得采用时长感知调度成为可能,从而降低了开销。这种优势在长上下文推理和强化学习应用中尤为明显。现有方法,如熵引导令牌池,以令牌熵为主要信号,但往往忽视令牌间语义内容的差异。因此,重要的代币往往被低估,而信息较少的代币则获得不成比例的重视。这损害了长度预测的可靠性。我们引入了ESTP(熵与语义令牌池),这是一个轻量级框架,通过结合熵与基于注意力的重要性评分来解决这一问题。这些分数直接来自LLM预填充阶段计算的自注意权重,这使得ESTP能够以最小的额外计算量捕捉不确定性和语义重要性。由于该框架重用预填充激活,几乎不增加额外内存开销,且仅引入极小的延迟。在ForeLen基准测试中,ESTP优于基线方法,在大多数场景下实现了更高的预测准确性和更低的错误率。当它与长度感知调度器集成到端到端系统测试时,还能进一步提升整体吞吐量并降低填充比。我们的结果为长度感知LLM服务系统提供了实用且有效的构建模块。
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
机器人多巴胺2.0:基于历史条件和值班时间感知的过程奖励建模用于机器人操作
- Authors: Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15680
- Pdf link: https://arxiv.org/pdf/2608.15680
- Abstract
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.
- 中文摘要
视觉-语言-动作(VLA)模型提升了机器人操作能力,但仍易受累积错误、场景转换和偏离轨迹的威胁。强化学习可以优化预训练的VLA策略,但稀疏的成功信号阻碍了探索,而工程化的密集奖励则成本高昂且针对特定任务。现有的学习视觉奖励模型常依赖静态的前后对抗观察,导致时间上的模糊性和在非分布(OOD)执行下保持鲁棒性变化与任务无效失败之间的区分较弱。我们介绍了Robo-Dopamine 2.0,这是一种历史和面向OOD的过程奖励模型,具有两两预测接口。它结合了(1)历史条件的成对奖励,使用源对齐的参考面板进行合成OOD查询,在线查询则使用观察到的展开历史,同时保留查询的端点;(2)一个面向对象的签名进度空间,表示有效进度、鲁棒性、失败和恢复。带有过渡感知回放的签名跳跃课程学习粗略执行顺序,然后进行细粒度进度校准。我们还构建了OOD轨迹数据集和五个家族基准。参考面板将平均视觉顺序一致性(VOC)从0.967提升至0.986,且有效活动范围VOC从0.906提升至0.958。在相同的40万配对奖励预算下,带25%回放的签名跳训练平均VOC达到0.9872,而匹配池洗牌控制组为0.9858。在下游强化学习中,完整模型实现了86.8%的平均机器人双胞胎成功率和71/80的真实世界插入成功率。
Adaptive Mixing of Policies from Searching and Policies from Learning
从搜索和学习策略的自适应混合
- Authors: Gavin B. Rens
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.15700
- Pdf link: https://arxiv.org/pdf/2608.15700
- Abstract
Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can take exceedingly long. Objectives: Rather than perform search to the same depth every time (typically at a fixed period of steps), reduce the search depth proportionally to the quality of the policy network priors. Methods: We describe Flexer, an architecture that, for each step, mixes the policy from a neural network and the policy from Monte Carlo tree search. The mixing factor favors the MCTS policy as the policy imitation error of the network and the environment models' variance increases. Results: Flexer outperforms a version of AlphaZero (and DQN and ADP) for some experiments on three toy symbolic problems.
- 中文摘要
背景:通过搜索/规划生成的训练目标提炼在强化学习中已被证明非常有用,但搜索过程可能非常耗时。目标:与其每次(通常在固定步数下)都进行相同深度搜索,不如根据策略网络先验的质量比例降低搜索深度。方法:我们介绍了Flexer架构,该架构在每一步中混合了神经网络策略和蒙特卡洛树搜索策略。混合因子有利于MCTS策略,因为网络和环境模型的策略模仿误差增加。结果:Flexer在三个玩具符号问题的某些实验中优于AlphaZero版本(以及DQN和ADP)。
GAINS: Leveraging Inconsistent Human Intervention Signals in Reinforcement Learning
收益:利用不一致的人类干预信号在强化学习中
- Authors: Xinyi Zhang, Yinuo Zhao, Pei Ren, Lechun Jiang, Huiqian Jin, Lei Sun, Dapeng Wu, Zhengping Che, Chi Harold Liu, Jian Tang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.15707
- Pdf link: https://arxiv.org/pdf/2608.15707
- Abstract
Correcting robot manipulation policies through human intervention holds great promise for real-world deployment, yet human operators are inherently imperfect in both the actions they provide and the timing of their intervention signals. While the former has been extensively discussed in reinforcement learning (RL), the latter remains underexplored. At high control frequencies, human intervention signals are often delayed and inconsistent across time and state space. In this work, we present GAINS, a framework for leveraging inconsistent human intervention signals in RL. At the core of GAINS, we employ distributional RL with quantile Q-networks to model the return variability induced by sparse task rewards and inconsistent human interventions. Building on this distributional representation, we introduce a pessimistic exploration strategy that promotes safe and sample-efficient learning under human corrections. We evaluate GAINS on four diverse simulated manipulation tasks and two challenging real-world scenarios against state-of-the-art intervention-based methods. GAINS achieves a 22% higher task success rate than RLIF and improves recovery success by up to 43% in failure scenarios. These results highlight the importance of modeling return variability induced by human imperfection for real-world deployment of intervention-based learning.
- 中文摘要
通过人工干预纠正机器人操控政策在实际应用中具有巨大潜力,但人类操作员在行动和干预信号的时机上本质上都不完美。虽然前者在强化学习(RL)中被广泛讨论,但后者仍然未被充分探索。在高控制频率下,人工干预信号常常在时间和状态空间中延迟且不一致。在本研究中,我们提出了GAINS框架,这是一个利用强化学习中人类干预信号不一致的框架。在GAINS的核心,我们利用分布式强化学习和分位Q网络,模拟由稀疏任务奖励和不一致的人类干预引发的回报变异性。基于这种分布表征,我们引入了一种悲观的探索策略,促进在人工修正下安全且样本高效地学习。我们评估了四项多样化模拟操作任务和两种具有挑战性的现实场景,并结合最先进的干预方法。GAINS的任务成功率比RLIF高出22%,在失败情景下恢复成功率提升高达43%。这些结果凸显了对人类不完美性导致的回报变异建模对于实际应用干预式学习的重要性。
TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness
TaoLive 数字化身代理技术报告:培训代理与其安全带同步进化
- Authors: TaoLive AIGC LLM Team: Yuhan Sun, Wenhao Lin, Yongdong Luo, Yibo Hu, Meiguang Jin, Junfeng Ma, Weihang Pan, Jiaxin Zhao, Zulong Chen
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.15763
- Pdf link: https://arxiv.org/pdf/2608.15763
- Abstract
AI-powered digital-avatar streamers in live e-commerce must answer product questions, engage viewers, and execute changing business strategies in real time. This requires low latency, factual and effective replies, and rapid adaptation to updated campaign, compliance, and style requirements. We develop an evolvable Harness that decouples Skills, Hooks, system prompts, and tools from model weights, allowing runtime behavior to change without retraining. However, Harness evolution creates a moving execution environment: compact models fine-tuned on one configuration may memorize names, schemas, and prompt templates rather than follow the Harness currently provided, while stronger zero-shot models are too slow for real-time use. We address this tension with Harness-Aware Training (HAT), which makes Harness states part of the training distribution. HAT applies task-preserving Harness-State Augmentation (HSA) to Skills, tool schemas, prompt structures, and interaction constraints, and comprises three stages: HSA-based supervised fine-tuning, general on-policy distillation to recover general capabilities, and HSA-based agentic reinforcement learning in a production-informed live-room simulator. Across four evaluation sets with more than 4,500 cases, our compact 35B model scores 94.8 on real-world Live-Stream QA, versus 80.3 for the base model and 93.0 for the strongest evaluated general LLM, while scoring 94.6 on Harness-Variant QA and retaining 83.5 on IFEval. By contrast, fixed-Harness SFT reduces IFEval by 7.7 points. In a controlled complete-agent replay on one NVIDIA H20 GPU with MTP enabled, the system achieves 3.407 s P50 and 8.114 s P95 latency. These results show that HAT produces a latency-feasible compact agent that remains effective under evaluated Harness changes without sacrificing general instruction following.
- 中文摘要
AI驱动的数字虚拟主播在实时电商中必须回答产品问题,吸引观众,并实时执行不断变化的商业策略。这需要低延迟、事实准确且有效的回复,以及快速适应更新的活动、合规性和风格要求。我们开发了一种可进化的Harness,将技能、钩子、系统提示和工具与模型权重解耦,允许运行时行为发生变化而无需重新训练。然而,束带演进创造了一个动态的执行环境:在单一配置上微调的紧凑模型可能只能记忆名称、模式和提示模板,而非遵循现有的束带,而更强的零样本模型则不适合实时使用。我们通过“具觉训练”(HAT)来解决这种张力,将束带状态纳入训练分布。HAT将任务保持的束-状态增强(HSA)应用于技能、工具模式、提示结构和交互约束,包含三个阶段:基于HSA的监督微调、基于策略的一般提炼以恢复通用能力,以及基于HSA的代理强化学习,在生产导向的现场模拟器中进行。在四个评估集、超过4500个案例中,我们的紧凑型35B模型在真实世界直播质量保证中得分为94.8,基础模型为80.3分,最强的通用大型语言模型为93.0分,而在Harness-Variant QA中得分94.6分,IFEval保持83.5分。相比之下,固定线束SFT则降低了7.7个百分点的IFEval。在启用MTP的NVIDIA H20 GPU上进行受控的完整代理重放时,系统实现了3.407秒的P50和8.114秒的P95延迟。这些结果表明,HAT能够生成一个延迟可行的紧凑代理,在评估过的束带变化下依然有效,同时不牺牲整体指令的跟随性。
Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning
通过基于重中心的对抗性逆强化学习学习股票交易策略
- Authors: Arishi Orra, Himanshu Choudhary, Manoj Thakur
- Subjects: Subjects:
Machine Learning (cs.LG); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2608.15770
- Pdf link: https://arxiv.org/pdf/2608.15770
- Abstract
Designing effective trading strategies using reinforcement learning remains challenging due to delayed and noisy rewards, poor exploration, and the difficulty of enforcing explicit risk constraints. In this work, we propose BRaG, a barycenter-based adversarial inverse reinforcement learning framework for stock trading that learns trading behavior from multiple heterogeneous expert strategies. BRaG aggregates expert demonstrations using a performance-weighted Wasserstein barycenter, yielding a stable pseudo-expert representation that captures shared structure across diverse trading styles. This representation is used to pretrain a trading policy via adversarial imitation learning, which alleviates unstable exploration during reinforcement learning. The pretrained policy is subsequently refined using reinforcement learning with true market rewards. To ensure risk-aware decision-making, BRaG incorporates control barrier functions that constrain action execution and regularize policy learning to satisfy drawdown limits. We evaluate the proposed approach on four major global equity markets, including the US, UK, Indian, and Taiwanese indices. Across all the markets, the proposed approach achieves stronger performance than both classical trading rules and recent deep reinforcement learning methods, while exhibiting more stable risk characteristics.
- 中文摘要
由于奖励延迟且噪音大、探索不足以及难以执行明确风险约束,设计有效的强化学习交易策略依然充满挑战。在本研究中,我们提出了BRaG,一种基于重心的对抗性逆强化学习框架,用于股票交易,能够从多种异构专家策略中学习交易行为。BRaG利用绩效加权的Wasserstein重心汇总专家演示,生成稳定的伪专家表示,捕捉不同交易风格间的共享结构。该表示被用来通过对抗性模仿学习预训练交易策略,从而缓解强化学习过程中的不稳定探索。预训练策略随后通过强化学习和真实市场奖励进行优化。为确保风险意识决策,BRaG集成了控制障碍功能,限制行动执行并规范政策学习以满足回撤限制。我们评估了拟议方法在包括美国、英国、印度和台湾等四大全球股市上的表现。在所有市场中,该方法比经典交易规则和近期的深度强化学习方法都更强,同时表现出更稳定的风险特征。
Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading
股票交易中稳定强化学习的自监督辅助任务发现
- Authors: Arishi Orra, Himanshu Choudhary, Manoj Thakur
- Subjects: Subjects:
Machine Learning (cs.LG); Computational Finance (q-fin.CP); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2608.15841
- Pdf link: https://arxiv.org/pdf/2608.15841
- Abstract
Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals. Auxiliary tasks are often used to improve representation learning and stabilize training, yet they are usually designed manually and depend heavily on prior assumptions about targets and prediction horizons. Such fixed designs may not remain suitable across changing market regimes. In this work, we propose a self-supervised framework that automatically discovers auxiliary tasks to support reinforcement learning for stock trading. The auxiliary tasks are formulated as General Value Functions so that their predictions enrich the learned state representation and assist policy optimization. The framework consists of two networks. The main network learns the trading policy along with the auxiliary predictions, while the secondary network generates the definitions of auxiliary tasks through learned cumulants and discount factors. These tasks are updated using a meta gradient mechanism that accounts for their long-term impact on trading performance and improves training stability. We evaluate the proposed approach across four major equity indices: DJI, FTSE, Sensex, and TAIEX. The empirical results demonstrate that automatically discovered auxiliary tasks lead to more robust learning and improved trading performance compared to existing baselines.
- 中文摘要
强化学习作为一种基于数据的股票交易方法,越来越受到关注。然而,由于市场行为不稳定和奖励信号嘈杂,学习既盈利又稳定的政策仍然具有挑战性。辅助任务常用于提升表征学习和稳定训练,但它们通常是手动设计的,且高度依赖于对目标和预测视野的先验假设。这种固定设计可能无法在不断变化的市场环境中保持适用性。本研究提出一个自监督框架,自动发现辅助任务以支持股票交易的强化学习。辅助任务被表述为一般价值函数,以丰富所学状态表示并协助策略优化。该框架由两个网络组成。主网络学习交易策略及辅助预测,而次级网络通过学习的累积量和贴现因子生成辅助任务的定义。这些任务通过元梯度机制进行更新,考虑其对交易表现的长期影响,并提升训练稳定性。我们评估了该提案方法在四大主要股票指数:DJI、富时、Sensex和TAIEX。实证结果表明,自动发现辅助任务能带来更稳健的学习和提升交易表现,相较于现有基线。
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
询问确认:信息丰富的互动,助你自信地推荐多回合大型语言模型
- Authors: Cedar Site Bai, Duanshun Li, Zhenyu Liao, Sheikh Sarwar, Huiyuan Chen, Yuan Chen, Changhe Yuan, Haiyang Zhang, Qilin Qi
- Subjects: Subjects:
Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.15949
- Pdf link: https://arxiv.org/pdf/2608.15949
- Abstract
Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained. We propose a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations. We apply this entropy reduction as a reward---without relying on ground-truth recommendations, which are often unavailable in real-world scenarios---to fine-tune the LLM, enabling strategic interaction generation. Empirical results with supervised fine-tuning (SFT) and direct preference optimization (DPO) on the INSPIRED and ReDial datasets show that our method improves both recommendation quality and conversational efficiency.
- 中文摘要
大型语言模型(LLMs)的最新进展使其能够作为会话推荐系统(CRS)使用,展现出强烈的推荐准确性和自然对话能力。然而,引导多回合交互以有效激发用户偏好仍然具有挑战性。现有方法要么使用带有模板交互的独立强化学习代理,要么优化由其他大型语言模型判断的交互性,而未测量实际获得多少有用信息。我们提出了一种新方法,通过减少助手的不确定性来量化每次交互的有效性,这一不确定性通过对推荐的熵来衡量。我们将这种熵减少作为奖励应用---不依赖现实场景中常常无法获得的地面真实推荐---来微调LLM,实现战略性交互生成。在 INSPIRED 和 ReDial 数据集上,采用监督微调(SFT)和直接偏好优化(DPO)的实证结果表明,我们的方法提升了推荐质量和会话效率。
DER Allocation without Load Prediction via Reinforcement Learning
通过强化学习实现无负载预测的DER分配
- Authors: Abed AlRahman Al Makdah, Aravind Ramana, Shaofeng Zou, Oliver Kosut, Lalitha Sankar
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2608.15977
- Pdf link: https://arxiv.org/pdf/2608.15977
- Abstract
The growing variability of renewable generation increases the need for fast and flexible grid-balancing mechanisms. Existing frameworks for distributed energy resource aggregations (DERAs) rely on short-term forecasts of net demand, making their performance highly sensitive to prediction errors. In this paper we present a forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data. We model the DERA dynamics as a deterministic linear system and the exogenous net load as a feature-based linear Markov process, capturing short-range temporal dependencies without explicit forecasting. We derive a closed-form expression for the optimal policy, which is learned through a least-squares value iteration (LSVI) algorithm using data collected across episodes. The proposed framework preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates. Numerical experiments on real California Independent System Operator (CAISO) net-demand data demonstrate that the learned controller achieves high tracking accuracy and stable regulation across heterogeneous DER aggregators without requiring any demand prediction.
- 中文摘要
可再生能源发电的变异性不断增加,增加了对快速且灵活的电网平衡机制的需求。现有的分布式能源资源聚合(DERA)框架依赖于短期净需求的预测,使其性能对预测误差极为敏感。本文提出了一个无预测强化学习(RL)框架,用于DERA分配,直接从运营数据中学习最优策略。我们将DERA动力学建模为确定性线性系统,外生净载荷则作为基于特征的线性马尔可夫过程,捕捉短期时间依赖性,无需显式预测。我们推导出最优策略的闭式表达式,该表达式通过最小二乘值迭代(LSVI)算法学习,数据涵盖多个集数。所提出的框架在通过数据驱动更新适应随机需求变化的同时,保持了DER模型的可解释性和约束满足性。对真实加州独立系统运营商(CAISO)净需求数据的数值实验表明,所学控制器能够在异构DER聚合器中实现高跟踪精度和稳定调控,无需任何需求预测。
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
学习剩下的,而非已掌握的:多奖励政策优化中的饱和优势重权重
- Authors: Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16072
- Pdf link: https://arxiv.org/pdf/2608.16072
- Abstract
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom. We introduce \textbf{Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization} (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels.
- 中文摘要
具有群体相对优势的强化学习(RL)已成为训练后语言模型推理者的事实标准。然而,在优化多个奖励目标时,现有方法通常会先用固定加权和标量奖励矢量,然后再进行组间标准化。我们表明,这种设计导致了两个根本性问题:具有不同奖励特征的推出可以获得相同的优势,且所有目标无论当前饱和程度如何,都采用固定的相对权重进行优化。因此,培训持续将梯度预算分配给已解决的目标,而非专注于剩余余量更大的目标。我们引入了\textbf{多重奖励政策优化的饱和感知优势重权重加权}(SA-MRPO),该方法独立标准化每个奖励目标,并根据批量级的客观饱和估计自适应折现其贡献。这动态地将优化努力重新分配到优化不足的目标上,同时在经验上保持对已经满意目标的性能。我们还进一步证明,感知饱和度的重权重可以逆转更新的符号,而不仅仅是重新调整其大小。在数学推理中,采用二目标和三目标奖励组合,SA-MRPO在15个基准比较中有12个提升了更难正确度目标相较GDPO,AIME24提升最高5%。在自适应推理方面,它在所有五项基准测试中平均提高了3.8美元,在AMC23上最高提升了9.2美元;在编码基准测试中,它提高了最多2.3美元,同时在所有设置下,较简单的目标都保持在已满足的水平附近。
US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina
US-VLA:一种针对具身腹部的超声视觉-语言-动作模型
- Authors: Cheng Zhang, Xingzheng Wu, Guihao Yan, Xifeng Hu, Zhi Liu, Mei Wu, Qing Cai
- Subjects: Subjects:
Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16074
- Pdf link: https://arxiv.org/pdf/2608.16074
- Abstract
Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing real-time guidance for standardized image acquisition and reducing operator dependence. However, existing reinforcement learning and learning-assisted ultrasound scanning methods typically rely on carefully designed reward functions or extensive interaction data, which limits their generalization ability and stability across different devices, patient populations, and complex clinical scenarios. To address these challenges, we propose an ultrasound vision-language-action model (US-VLA) for automated ultrasound scanning that explicitly encodes clinical semantic goals and generates sequential probe manipulation actions under real-time ultrasound feedback. In particular, we first design an ultrasound-aware expert fusion module to jointly integrate ultrasound observations with auxiliary contextual information, enabling semantic ultrasound feedback to effectively guide the scanning process. Then, we construct US-VLA-Data, a real-world dataset covering liver and kidney examinations, which includes five clinically defined standard planes and comprises 320 expert scanning trajectories with approximately 80,000 synchronized timesteps. Extensive experiments demonstrate that US-VLA achieves competitive performance in ultrasound probe manipulation tasks, indicating its effectiveness and promising generalization within the evaluated abdominal ultrasound setting. The source code is available at this https URL.
- 中文摘要
人工智能辅助超声扫描通过提供实时指导,实现标准化图像采集,提升诊断的可靠性和效率,减少操作员依赖。然而,现有的强化学习和辅助超声扫描方法通常依赖精心设计的奖励函数或大量交互数据,这限制了其在不同设备、患者群体和复杂临床场景中的泛化能力和稳定性。为应对这些挑战,我们提出了一种用于自动超声扫描的超声视觉-语言-动作模型(US-VLA),该模型明确编码临床语义目标,并在实时超声反馈下生成顺序探针操作动作。特别是,我们首先设计了一个超声感知的专家融合模块,将超声观测与辅助上下文信息联合集成,使语义超声反馈能够有效指导扫描过程。随后,我们构建了US-VLA数据,这是一个涵盖肝脏和肾脏检查的真实数据集,包含五个临床定义的标准平面,包含320条专家扫描轨迹,约8万个同步时间步长。大量实验表明,US-VLA在超声探针操作任务中具有竞争力,显示其有效性并在评估的腹部超声环境中具有推广潜力。源代码可在该 https URL 访问。
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
图神经辅助演员-批判者,用于延迟高效的边缘视觉系统
- Authors: Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Gutiérrez Gaitán, Eduardo Tovar
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16142
- Pdf link: https://arxiv.org/pdf/2608.16142
- Abstract
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities. The latency of video transmission has a profound impact on the effectiveness of the operator assistance. However, most techniques available for video transmission still incur significant latency costs. In this paper, we propose a graph convolutional neural network-assisted (GCN-Assisted A2C) deep reinforcement learning (DRL) system model to find the optimal pixel-correlated area of a suspicious object. We combine the Lagrangian dual form with gradient descent to prevent lack of convergence and over- and under-penalization constraint violation during latency optimization. The proposed system model sends a sub-group pixel-correlated area of the frame from the UAV to the server rather than the transmission of the whole video frame. The proposed framework utilizes the GCN model to explore hidden representations of feature-correlated groups of pixels. Moreover, the GCN supervises the A2C model, which selects a subgroup to enhance transmission latency, thus supervising the training of UAV actions in A2C. Experimental results show that GCN-assisted A2C reduces video frame transmission latency together with false detection rate in UAV vision systems over other DRL and state-of-the-art models.
- 中文摘要
无人机载视觉系统广泛应用于各种活动,包括禁飞区的监控。在这种情况下,配备视觉的无人机将视频传输到地面服务器,操作员协助其活动。视频传输的延迟对操作员辅助的有效性有深远影响。然而,大多数可用的视频传输技术仍然会产生显著的延迟成本。本文提出一种图卷积神经网络辅助(GCN辅助A2C)深度强化学习(DRL)系统模型,用于寻找可疑物体的最佳像素相关区域。我们将拉格朗日对偶形式与梯度下降结合,以防止延迟优化过程中收敛缺失以及过度惩罚和不足惩罚约束违规。所提系统模型将帧中一个子组像素相关区域从无人机发送到服务器,而不是传输整个视频帧。所提框架利用GCN模型探索特征相关像素群的隐藏表示。此外,GCN监督A2C模型,该模型选择一个子组以增强传输延迟,从而监督无人机在A2C中的训练。实验结果显示,GCN辅助的A2C相比其他日程学习和先进型号,在无人机视觉系统中降低了视频帧传输延迟和误探率。
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
TRCA:长视野LLM代理的过渡性评分标准学分分配
- Authors: Huan Zhang, Mingju Chen, Dongxu Zhou, Can Lv, Heng Chang, Sen Cui, Faguo Wu, Shiji Zhou
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16156
- Pdf link: https://arxiv.org/pdf/2608.16156
- Abstract
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from successful trajectories. However, successful trajectories are extremely scarce during early-stage reinforcement learning, substantially weakening anchor-based methods. We propose Transition-wise Rubric Credit Assignment (TRCA), which derives step-level supervision directly from action-induced transitions without learned evaluators or successful anchors. TRCA evaluates each transition using Evidence, Execution, and Invalidity rubrics to capture task-relevant information acquisition, valid task execution, and invalid or regressive behavior. From these judgments, Foundational Rubric Reward measures local transition quality, while Breakthrough Rubric Reward tracks newly covered Evidence and Execution conditions to reward incremental task progress. Combined with terminal outcomes, these signals produce fine-grained step-level advantages for policy optimization. Experiments on ALFWorld, WebShop, and seven search-augmented question-answering benchmarks show consistent improvements over the evaluated baselines. With Qwen2.5-7B-Instruct, TRCA improves the WebShop score by 6.0%-12.6%; with Qwen2.5-3B-Instruct, it improves the average SearchQA score by 1.9%-18.3%. These results demonstrate the effectiveness of transition-wise rubric credit assignment for long-horizon tasks with sparse successful anchors.
- 中文摘要
长视野大型语言模型(LLM)代理通常以稀疏的终端结果进行优化,这使得在多步交互中进行细粒度的信用分配变得困难。现有方法要么依赖过程评估器,产生注释和推断成本,要么从成功的轨迹中获得步骤级的信用。然而,在早期强化学习阶段,成功的路径极为稀少,极大削弱了基于锚点的方法。我们提出过渡性评分标准学分分配(TRCA),该方法直接从行动诱导的过渡中获得步骤级监督,无需学习评估者或成功锚点。TRCA利用证据、执行和无效性评分标准评估每个过渡,以捕捉任务相关信息获取、有效任务执行以及无效或倒退行为。基于这些判断,基础评分标准奖励衡量局部过渡质量,而突破评分标准奖励则追踪新涵盖的证据和执行条件,以奖励任务的增量进展。结合终极结果,这些信号为政策优化带来了细致的步骤级优势。在ALFWorld、WebShop以及七个搜索增强问答基准测试上的实验显示,相较于评估基线,持续有改善。通过Qwen2.5-7B-Instruct,TRCA将WebShop评分提升了6.0%-12.6%;配合Qwen2.5-3B-Instruct,平均SearchQA得分提升了1.9%-18.3%。这些结果证明了在长期任务中成功锚点稀疏的跨度评分标准分配的有效性。
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics
通过受控自助和受控价值动态理解并稳定深度Q学习
- Authors: Bozhou Chen, Yongyi Wang, Hanyu Liu, Xionghui Yang, Wenxin Li
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16182
- Pdf link: https://arxiv.org/pdf/2608.16182
- Abstract
Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously unstable. Existing studies often attribute instability to isolated factors such as overestimation bias or representation learning issues, lacking a unified understanding of how different sources of instability interact during recursive value estimation. In this work, we provide a systematic analysis of instability in deep Q-learning from three complementary perspectives: operator-level bias in Bellman bootstrapping, estimator-level sensitivity of greedy action selection to regression noise, and parameter-dynamics imbalance under aggressive data reuse. We identify a reward-triggered self-reinforcing trap and characteristic parameter spike dynamics, then derive stabilization principles for controlled bootstrapping, ensemble quantile estimation, and spike-based parameter regulation. Experiments on Atari-100K and Procgen demonstrate competitive performance and improved training stability.
- 中文摘要
深度Q学习(DQL)在强化学习方面取得了显著的实证成功,但其训练过程仍然以不稳定著称。现有研究通常将不稳定性归因于孤立因素,如高估偏差或表示学习问题,缺乏对递归价值估计中不同不稳定性来源相互作用的统一理解。本研究从三个互补视角系统分析深度Q学习中的不稳定性:Bellman自举法中的算子级偏差、贪婪动作选择对回归噪声的估计量级敏感性,以及激进数据重用下的参数动态失衡。我们识别了奖励触发的自我强化陷阱和特征参数尖峰动态,随后推导出受控自助法、集合分位数估计和基于尖峰参数调控的稳定原则。Atari-100K和Procgen上的实验显示出竞争性能和训练稳定性的提升。
RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing
RoboStriker:自动人形拳击的潜在空间战略游戏
- Authors: Kangning Yin, Kaige Liu, Zhe Cao, Wentao Dong, Weishuai Zeng, Tianyi Zhang, Qiang Zhang, Jingbo Wang, Jiangmiao Pang, Yang Li, Ming Zhou, Weinan Zhang
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.16195
- Pdf link: https://arxiv.org/pdf/2608.16195
- Abstract
Achieving human-level competitive intelligence and physical agility in humanoid robots remains a profound challenge, particularly in contact-rich and highly dynamic tasks such as boxing. While Multi-Agent Reinforcement Learning offers a principled framework for strategic interaction, its direct application to unstructured raw motor spaces inevitably leads to joint-level physical collapse, preventing the emergence of any viable combat tactics. To resolve this fundamental conflict between strategic exploration and physical feasibility, we formulate the humanoid combat task as a novel two-player latent-space zero-sum Markov game. Under standard regularity and approximate best-response assumptions, we show that the latent formulation induces an equivalent game over the decoder-reachable action manifold, providing an approximate-Nash interpretation of the resulting self-play dynamics. To instantiate this theoretical formulation, we propose RoboStriker, a hierarchical framework that decouples high-level reasoning from low-level execution. It first distills the tracking expertise of predefined boxing motions into a topologically bounded latent manifold. This structured latent foundation subsequently drives multi-agent co-evolution via Latent-Space Neural Fictitious Self-Play. Extensive experimental results demonstrate that gaming within this structured latent space substantially outperforms direct exploration. By constraining strategic exploration through a pretrained motion decoder, RoboStriker substantially reduces the catastrophic balance failures observed in raw action-space methods and achieves superior tactical performance in both competitive win rates and striking efficiency. Finally, we successfully deploy and validate our learned combat policies on real-world humanoid robots. Our code and video and supplementary materials are available at RoboStriker.
- 中文摘要
在人形机器人中实现人类水平的竞争智力和身体敏捷性仍是一项深刻挑战,尤其是在接触密集且高度动态的任务中,如拳击。虽然多智能体强化学习为战略互动提供了原则性框架,但其直接应用于无结构的原始运动空间,必然导致关节层面的物理崩溃,阻碍任何可行的战斗战术的出现。为了解决战略探索与物理可行性之间的根本冲突,我们将类人生物战斗任务设计为一种新颖的双人潜空间零和马尔可夫游戏。在标准正则性和近似最佳响应假设下,我们证明了潜在表述在解码器可达作用流形上诱导出等价博弈,从而对所得的自玩动态提供了近似纳什解释。为了实现这一理论表述,我们提出了RoboStriker,一个分层框架,将高层推理与低层执行解耦。它首先将预设拳击动作的跟踪技术提炼为拓扑有界的潜在流形。这一结构化的潜在基础随后通过潜在空间神经虚构自我游戏推动了多智能体的共进化。大量实验结果表明,在这一结构化潜伏空间中的游戏远远优于直接探索。通过预训练的运动解码器限制战略探索,RoboStriker大幅减少了原始动作空间方法中观察到的灾难性平衡失效,并在竞争胜率和打击效率上实现了更优越的战术性能。最后,我们成功地在现实世界的人形机器人上部署并验证了我们所学到的战斗策略。我们的代码、视频和补充资料可在RoboStriker获取。
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
BaT:迈向具有阶段评分标准的自我进化医学研究代理
- Authors: Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16211
- Pdf link: https://arxiv.org/pdf/2608.16211
- Abstract
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.
- 中文摘要
长期代理开始自动化完整的工作流程,生成代码、报告和研究成果。医学影像工作流程多阶段且数据敏感,而专家的经验仍然稀少且难以共享。结构化基准可以通过阶段级评分标准定位故障,但标准的培训后评估在下一轮培训前会舍弃这些诊断。我们介绍了基准即教师(Benchmark-as-Teacher,简称BaT),这是一个用于代理培训后递归自我提升的系统。BaT包含两个相关组件:异步阶段银行数据流水线和BiCuRL(双级课程强化学习),后者是其自我改进的训练后方法。Stage Bank 综合了策略更新循环之外的内容隔离训练状态。BiCuRL使用固定的预留评估来选择下一阶段课程,通过任务评分标准验证推广,使用GRPO更新策略,并将候选检查点返回评估。在AutoMedBench-Lite上,BaT-4B和BaT-9B的总分是Qwen Ininstruction基线的两倍多。BaT-9B特工整体评分为79.6,超过了克劳德作品4.6,克劳德代码为77.5。
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
Defake-o3:从推测性理由到可验证的AIGI检测证据
- Authors: Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, Jianfu Zhang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16259
- Pdf link: https://arxiv.org/pdf/2608.16259
- Abstract
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.
- 中文摘要
图像生成模型的快速发展要求AI生成图像(AIGI)检测器不仅准确,还具备可解释性和可靠性。虽然基于MLLM的探测器可以提供自然语言解释,但现有方法常常产生推测性推测:它们依赖模糊或幻觉的伪造物,遗漏最新生成器的细微局部缺陷,且无法提供可直观验证的证据。我们介绍Defake-o3,一种可解释的AIGI探测器,从推测性推测走向可验证的证据。它结合了交互式视觉搜索和验证者引导的证据对齐:模型会迭代放大到可疑区域,检查细致细节,而由人工验证注释训练的证据验证器则提供强化学习奖励,有利于有根据的证据并惩罚无根据的主张。为支持这一目标,我们构建了GroundFake数据集,旨在基于可解释的检测,包含局部边界盒证据、基于视觉基础和人工物特异性的人工验证、修正推理轨迹以及有效/无效证据监督。我们还进一步介绍了FakeFrontier,这是一个由10个近期生成器的真实图像和输出构建的非发行基准测试,以及基于MLLM的证据质量和说服力评估协议。GroundFake、FakeFrontier及其他非分发基准测试的实验显示,Defake-o3不仅提高了检测准确性,还提高了解释质量,从而产生了更具本地化、可验证性和说服力的证据。
TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation
TransAnyText:通过结构化视觉生成翻译电商图片中的任意文本
- Authors: Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He, Xiaoyan Su, Shaojie Guo, Hao Yang, Jingling Fu, Xiaolong Fu, Zhen Chen, Yu Guo, Fei Wang, Xinyi Liu, Yongjun Zhang, Ke Zhang, Junshi Huang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16284
- Pdf link: https://arxiv.org/pdf/2608.16284
- Abstract
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translation, faithful visual identity preservation, and easy-to-edit outputs, simultaneously. To address these challenges, we introduce TransAnyText, a structured visual code framework that reformulates image text translation as generating renderable HTML patches from source images and target languages. Our framework decouples semantic generation from pixel rendering: a vision-language model (VLM) handles visual understanding, cross-lingual translation, and structured visual generation, while a diffusion model performs background inpainting and pixel-level refinement, followed by deterministic rendering to synthesize the final image. Based on this formulation, we develop a three-stage post-training framework, where supervised fine-tuning (SFT) establishes the image-to-code mapping, privilege-gap weighted self-distillation (PWSD) improves the learning of style and layout tokens, and reinforcement learning with verifiable rewards (RLVR) further optimizes task-level performance. We further introduce TransAnyDataset and TransAnyBench, a multilingual dataset and benchmark for e-commerce image translation. Extensive experiments demonstrate competitive performance against cascaded pipelines, open-source end-to-end models, and closed-source image editing systems, providing an effective, controllable, and editable solution for cross-border e-commerce image translation.
- 中文摘要
跨境电子商务图片翻译对于全球零售至关重要,因为产品图片、横幅和详情页需要用不同语言制作。现有方法难以同时实现准确翻译、忠实的视觉身份保存和易于编辑的输出。为应对这些挑战,我们引入了TransAnyText,一个结构化的可视化代码框架,将图像文本翻译重新表述为从源图像和目标语言生成可渲染的HTML补丁。我们的框架将语义生成与像素渲染分离:视觉语言模型(VLM)负责视觉理解、跨语言翻译和结构化视觉生成,而扩散模型则进行背景修补和像素级细化,随后通过确定性渲染合成最终图像。基于这一表述,我们开发了一个三阶段的训练后框架,其中监督微调(SFT)建立图像到代码映射,特权差距加权自我蒸馏(PWSD)提升样式和布局标记的学习,带可验证奖励的强化学习(RLVR)进一步优化任务级表现。我们还进一步介绍了TransAnyDataset和TransAnyBench,这是一个多语言数据集和电子商务图像翻译的基准。大量实验证明了在与级联流水线、开源端到端模型和闭源图像编辑系统竞争中的性能,为跨境电商图像翻译提供了有效、可控且可编辑的解决方案。
PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster
海报文本:迈向电子商务海报统一的视觉文本生成与编辑
- Authors: Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He, Xiaoyan Su, Shaojie Guo, Jingling Fu, Xiaolong Fu, Hao Yang, Tongxuan Liu, Yu Guo, Fei Wang, Xinyi Liu, Yongjun Zhang, Junshi Huang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16289
- Pdf link: https://arxiv.org/pdf/2608.16289
- Abstract
Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end poster generation or follow multi-stage design pipelines, with limited capability for flexible and precise editing of existing posters. To enable unified generation and editing of e-commerce posters, we introduce Text Patch Generation and Editing, a unified task formulation that treats text patches as atomic units and covers four operations: poster generation, patch addition, patch deletion, and patch modification, with optional reference-guided style control. Based on this, we propose PosterText, a unified model trained with a four-stage curriculum, including text rendering pretraining, instruction-following training, reinforcement learning for preference alignment, and spatial guidance self-distillation for execution refinement. We further construct a large-scale dataset with patch-level annotations and a comprehensive benchmark for evaluation. Extensive experiments demonstrate that PosterText achieves competitive performance against existing generation and editing approaches, validating the effectiveness of the proposed framework.
- 中文摘要
自动化电子商务海报设计既需要高质量的海报生成,也需要对现有设计进行灵活编辑。然而,大多数现有方法要么面向端到端海报生成,要么采用多阶段设计流程,对现有海报的灵活和精确编辑能力有限。为了实现电子商务海报的统一生成和编辑,我们引入了文本补丁生成与编辑,这是一种统一的任务公式,将文本补丁视为原子单元,涵盖四项操作:海报生成、补丁添加、补丁删除和补丁修改,并可选地提供参考引导样式控制。基于此,我们提出了PosterText,这是一个统一模型,采用四阶段课程训练,包括文本渲染预训练、指令跟随训练、强化学习以匹配偏好,以及空间引导自蒸馏以实现执行细化。我们进一步构建了一个带有补丁级注释和综合评估基准的大规模数据集。大量实验表明,PosterText在与现有生成和编辑方法竞争中实现了竞争性能,验证了所提框架的有效性。
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding
StreamOPD:带有时空提示门控的培训后配方,用于流媒体视频理解
- Authors: Keming Wu, Baoyi Wang, Kaichen Zhang, Xiang An, Zuhao Yang, Sudong Wang, Haowei Zhu, Tingxuan Huang, Hongcheng Gao, Bin Wang
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16320
- Pdf link: https://arxiv.org/pdf/2608.16320
- Abstract
Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Existing systems add inference-time memory, retrieval, and compression, yet a training-free sliding-window baseline already matches them. We therefore fix a memory-free recent-window protocol and ask how far post-training alone can go. Reinforcement learning with verifiable rewards fits this regime poorly, encouraging long ``think-then-answer'' generations, while on-policy distillation (OPD) supplies dense token-level teacher supervision on student trajectories but is stable only when both models train in thinking mode. These observations lead to \textsc{StreamOPD}, a recipe combining verifiable streaming-video data, thinking-mode OPD, and instruct-mode deployment. It raises StreamingBench from $77.9\%$ to $83.9\%$---within $0.3$ points of the 9B teacher---and improves OVO-Bench excluding its hallucination-detection subtask (HLD) by $9.1$ points under unchanged inference. As a teacher-privilege extension, \emph{Spatio-Temporal CueGate (ST-CueGate)} aggregates cue-versus-no-cue teacher likelihood ratios into a group-relative response score that reweights OPD. It reaches $71.9\%$ on OVO-Bench (excluding HLD) and $64.9\%$ on Video-MME, and is the only variant that stays above the base model on all four benchmarks. Replacing the teacher with a frozen copy of the student's initial policy---on-policy self-distillation---retains most of these gains and lifts HLD to $57.0\%$, above both the untrained student and the 9B teacher, so abstention loss is not intrinsic to the recipe. We provide a transparent and reproducible reference for open-source streaming-video research.
- 中文摘要
流式视频理解需要对视频展开时因果观察到的前缀做出直接反应。现有系统增加了推理时间内存、检索和压缩功能,但无训练滑动窗口基线已经与它们匹配。因此,我们修正了一个无内存的最近窗口协议,并探讨仅靠训练后能达到多远。带有可验证奖励的强化学习不适合这种模式,鼓励长时间的“思考后回答”世代,而策略提炼(OPD)则提供了密集的代币级教师对学生轨迹的监督,但仅在两种模型都处于思考模式下训练时才稳定。这些观察结果促成了 \textsc{StreamOPD},这是一种结合可验证流媒体数据、思考模式 OPD 和指令模式部署的配方。它将StreamingBench的价值从77.9美元提升到83.9美元---比9B级教师差0.3美元以内---并且在不变推断下,去除幻觉检测子任务(HLD)后,OVO-Bench提升了9.1美元。作为教师特权扩展,\emph{Spatio-Temporal CueGate(ST-CueGate)}将提示与无提示教师的似然比汇总为一组相对反应评分,重新加权OPD。它在OVO-Bench(不含HLD)上达到71.9美元,在Video-MME上达到64.9%美元,并且是唯一一个在四个基准测试中都高于基础型号的版本。用学生最初的政策中---政策自我蒸馏---的冰冻副本替代教师,保留了大部分这些收益,并将HLD提升到57.0%美元,高于未受培训的学生和9B级教师,因此放弃率的损失并非配方的固有部分。我们为开源流媒体视频研究提供透明且可重复的参考资料。
KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning
KC-BFPRL:通过双级基于前指点的强化学习实现知识引导多无人机草地恢复协作
- Authors: Dongbin Jiao, Xianyi Wang, Yuchen Yuan, Weibo Yang, Peng Yang, Peng Zhao, Zhanhuan Shang, Shi Yan
- Subjects: Subjects:
Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2608.16326
- Pdf link: https://arxiv.org/pdf/2608.16326
- Abstract
Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.
- 中文摘要
多无人机系统为大规模环境任务(如草原生态系统修复)提供了可扩展的服务平台。然而,协调车队运营需要解决恢复区域最大化问题(RAMP)。这一非线性组合优化挑战因有效载荷依赖的能量动力学和异质生态退化而复杂化。我们提出了一种新的知识引导协作双级前指点强化学习框架(KC-BFPRL),以应对这一复杂性。通过层级范式,KC-BFPRL将RAMP分解为全局任务分配和局部恢复规划,后者进一步细分为上层轨迹规划和下层恢复区域分配。我们的专业架构对包括基于Transformer的编码器,将静态环境特征与动态无人机状态融合,以及通过稳健的actor-critic框架训练的指针网络解码器。通过嵌入生态优先规则和启发式逻辑,KC-BFPRL实现了结构化的热启动,解决了强化学习的冷启动问题,同时确保严格的约束满足。大量实验表明,KC-BFPRL持续优于最先进的基线,实现了更优的客观值和效率。它在最复杂场景中保持了0.00%%的最优差距,运行速度几乎是MAPDP的三倍,验证了其鲁棒性、可扩展性和实时适用性,适用于大规模自动化生态恢复。
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
PertMind:通过基于细胞扰动数据的强化学习,激发LLM中涌现的生物推理
- Authors: Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Quantitative Methods (q-bio.QM)
- Arxiv link: https://arxiv.org/abs/2608.16419
- Pdf link: https://arxiv.org/pdf/2608.16419
- Abstract
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.
- 中文摘要
大型语言模型可以描述机制,但可扩展的后训练仍依赖于昂贵且手动整理的生物推理痕迹。我们展示了细胞扰动图谱可以转变为强化学习环境,在那里测量的基因响应为生物学推理提供可计算的奖励。我们介绍PertMind,它结合了可信轨迹监督初始化与基因、通路和格式级强化信号。PertMind仅训练于前向扰动-响应预测,在未可见的细胞环境中提升了响应推断能力,同时保留了通用语言能力。它还在无特定任务的后期训练中转入逆向微扰识别、双重微扰推理、表型筛选优先级和生物过程解释。PertMind进一步生成了生物学档案,支持多尺度下游任务中基因、细胞和供体的竞争性表征。这些结果支持这样一个假说:对实验终点进行强化可以集中预训练模型已具备的可重复使用生物策略。更广泛地说,扰动衍生强化学习提供了一种可扩展的路径,将不断扩展的实验图谱转化为通用生物推理的训练环境。
Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination
验证大型语言模型在网络安全模拟中的实用性:全面考察
- Authors: Stylianos Kampakis, Fabio Rovai, Marcos Charalambides, Theodosis Mourouzis, Chris Hicks
- Subjects: Subjects:
Cryptography and Security (cs.CR)
- Arxiv link: https://arxiv.org/abs/2608.16422
- Pdf link: https://arxiv.org/pdf/2608.16422
- Abstract
Cyber threats continue to escalate in both frequency and sophistication, necessitating more adaptive and scalable defense strategies. This paper explores how Large Language Models (LLMs) can bolster cybersecurity simulations by automating the creation of synthetic environments and identifying latent vulnerabilities. We employ YAML as a structured representation format for simulating complex network configurations, thereby enabling Large Language Model-driven pipelines to support and improve reinforcement learning (RL) agent training. Comparative studies examine the advantages of LLM-based techniques over classical approaches such as Double Q-learning with Prioritized Experience Replay (PER), emphasizing increased efficiency, higher adaptability, and enhanced realism in cyberattack simulations. In empirical benchmarks across multiple synthetic topologies, LLM-instantiated Python agents achieved up to a 94.5% compromise rate while executing in 0.02-0.06 seconds per assessment---a ~25,000x to 50,000x speedup over traditional RL training cycles. Our findings underscore the transformative potential of integrating LLMs into cybersecurity research, ultimately paving the way for more intelligent and robust cyber-defense systems.
- 中文摘要
网络威胁的频率和复杂程度持续升级,需要更具适应性和可扩展性的防御策略。本文探讨了大型语言模型(LLMs)如何通过自动化创建合成环境和识别潜在漏洞来增强网络安全模拟。我们采用YAML作为结构化表示格式来模拟复杂网络配置,从而使大型语言模型驱动的流水线能够支持并提升强化学习(RL)代理训练。比较研究考察基于LLM技术相较于双Q学习(Double Q-learning)及优先体验重放(PER)等传统方法的优势,强调在网络攻击模拟中提升效率、更高的适应性和更真实性。在多种合成拓扑的实证基准测试中,LLM实例化的Python代理实现了高达94.5%的攻破率,每次评估执行时间为0.02-0.06秒---比传统强化学习训练周期快约25,000倍至50,000倍。我们的发现强调了将LLM整合进网络安全研究的变革潜力,最终为更智能、更强大的网络防御系统铺平道路。
Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics
通过不确定性引导混合动力学实现的稳定多步推广
- Authors: Andrei Maalberg, Axel Neumann, Jens Knobloch
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2608.16431
- Pdf link: https://arxiv.org/pdf/2608.16431
- Abstract
Multi-step rollouts are essential for model-based reinforcement learning (RL) and predictive control, yet learned dynamics models often become unstable when recursively applied, leading to divergence and unreliable policy updates. This paper proposes a model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law. The switching signal is derived from calibrated epistemic uncertainty and activates only when the system leaves the nominal region, ensuring that each model operates within its reliability regime. Under clearly stated smoothness and boundedness assumptions, we show that the resulting hybrid predictor yields globally bounded recursive multi-step rollouts: trajectories remain Lyapunov-stable in the nominal region and exhibit at most affine growth during excursions. To illustrate the theory in practice, we instantiate the hybrid dynamics framework within a model-based RL scheme that uses real one-step transitions for value learning and hybrid rollouts for policy improvement. Experiments on a nonlinear Duffing oscillator demonstrate stable long-horizon prediction and improved cost-effort trade-offs relative to a stabilizing baseline.
- 中文摘要
多步推广对于基于模型的强化学习(RL)和预测控制至关重要,但学习到的动力学模型在递归应用时常常变得不稳定,导致背离和策略更新不可靠。本文提出了一种模型无关的混合动力学框架,结合可证明收缩的名义模型与通过不确定性引导切换定律的灵活行程模型。切换信号源自校准的认知不确定性,仅在系统离开名义区域时激活,确保每个模型在其可靠性范围内运行。在明确的光滑性和有界假设下,我们证明所得混合预测器产生全局有界递归多步展开:轨迹在名义区域保持李雅普诺夫稳定,且在振幅期间最多为仿射增长。为了在实践中说明该理论,我们将混合动力学框架纳入基于模型的强化学习方案,该方案采用真实的一步转换进行价值学习,混合推广用于政策改进。非线性达夫振荡器实验显示了稳定的长视野预测,并且相较于稳定基线,成本与努力权衡得到了改善。
Drive, Pack, Fly: The Travelling Thief Problem with Drone
驾驶、打包、飞行:无人机的旅行小偷问题
- Authors: Kabir Murjani, Abhay Sobhanan
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2608.16435
- Pdf link: https://arxiv.org/pdf/2608.16435
- Abstract
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this penalty by retrieving outlying items, thereby shortening the makespan and increasing operational profit. However, travel time remains load-dependent, and each item collected by the ground vehicle shifts the arrival times that govern the drone's launch and rendezvous points. This paper introduces the Travelling Thief Problem with Drone (TTP-D), which maximises the collected profit, net of a time-based rental cost, by jointly optimising item selection, vehicle routing, and flight synchronisation. We formulate a mixed-integer linear program that solves small instances to optimality, and develop both metaheuristics and an attention-based Deep Reinforcement Learning (DRL) policy for larger instances. We further propose a learner-initialised hybrid solver, in which the DRL policy constructs an initial solution that a short annealing run subsequently refines. On two benchmark sets, this hybrid recovers most of the metaheuristic baseline's quality at a fraction of its computational budget, although the largest instances still require the baseline at its full budget. Finally, a sensitivity analysis reveals that the rental ratio is the primary driver of profitability, whereas the fleet parameters affect profit only at the margin.
- 中文摘要
在收集作业中,有效载荷的积累会逐渐减速,从而对路径效率产生累积的惩罚。机载无人机可以通过回收外部物品来抵消这一损失,从而缩短使用寿命并提高运营利润。然而,飞行时间仍依赖于负载,地面车辆收集的每件物品都会调整决定无人机发射和会合点的到达时间。本文介绍了无人机的旅行小偷问题(TTP-D),该问题通过联合优化物品选择、车辆路线和飞行同步,最大化扣除基于时间租赁成本的利润。我们制定了一个混合整数线性规划,将小实例求解至最优,并为较大实例开发元启发式和基于注意力的深度强化学习(DRL)策略。我们进一步提出了学习者初始化的混合求解器,其中DRL策略构建一个初始解,随后通过短退火运行进行细化。在两个基准测试集上,这种混合体能以极小的计算预算恢复元启发式基线的大部分质量,尽管最大实例仍需基线的全部预算。最后,敏感性分析显示,租赁比率是盈利能力的主要驱动因素,而车队参数仅在利润边际部分产生影响。
Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
ICU败血症血流动动力学管理的离线强化学习:结合双重离场策略评估的MIMIC-IV研究
- Authors: Marc Pérez-Roig, David Fernández-Narro, Carlos Sáez
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.16482
- Pdf link: https://arxiv.org/pdf/2608.16482
- Abstract
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.
- 中文摘要
败血症中静脉输液和血管抑制剂的剂量是一个在不确定性下做出的顺序决策,主要由临床判断指导,因此成为从历史护理中强化学习的自然目标。由于学到的政策无法在患者身上试验,其价值必须在政策外估算,而这种估计可能脆弱且乐观。本研究通过结合非策略估计、可靠性诊断和临床医生一致分析,在透明的验证框架中推进败血症治疗政策的可靠评估。我们以MIMIC-IV重症监护数据库中抽取的36,872例化脓性ICU住院患者为模型,采用离散化马尔可夫决策过程,包含1,000个状态和25个行动,定义为5×5的液体和血管增压剂水平网格,并通过政策迭代解决。临床医生的行为政策采用随机森林估计,这缓解了有效样本量(ESS 50.1对4.0平滑计数)的崩溃,从而削弱了重要性抽样估计的不稳定。所学政策采用两种估计量评估:加权重要性抽样(WIS)和拟合Q评估(FQE),ESS和临床医生的共识作为可靠性检验。一项实证变量选择发现,州的组成比其规模更为重要。两种估计指标都将学到的政策置于临床医生回报之上(WIS 50.8,FQE 46.8对38.2,ESS 50.1),但与实际实践仅有适度偏差(总变差0.18),更倾向于减少静脉输液。这些回顾性单一中心的非政策结果支持了该学出的政策作为对观察实践的临床合理改进,并促使其作为基于不一致的临床决策支持方法进行进一步评估。
FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning
FLEET:基于令牌的特征提取,用于基于事件摄像机的强化学习
- Authors: Tristan Gottwald, Maximilian Schier, Melanie Schaller, Bodo Rosenhahn
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16523
- Pdf link: https://arxiv.org/pdf/2608.16523
- Abstract
Event cameras generate asynchronous, high-frequency data streams offering spatially sparse information at lower latency than traditional this http URL principle, these properties should be ideal for the design of control this http URL, reinforcement learning research in this field remains limited as existing approaches fail to fully exploit the sensor's this http URL-based methods negate the sensors benefits by aggregating events into sparse grids. This couples compute cost to sensor resolution and blurs the temporal information. Meanwhile, existing generative baselines rely on the availability of trajectory data to pretrain the model. We propose FLEET (Feature Learning from Events via Efficient Tokenization), a feature extractor that processes event sequences directly. Leveraging random Fourier features and cross-attention, our architecture compresses variable streams into fixed-size latent representations. This decouples inference cost of the feature extractor's backbone from the sensor's resolution, enabling end-to-end learning without auxiliary losses. We validate FLEET on a new, high-throughput benchmark. The results demonstrate that our sequence-based approach surpasses SOTA performance and exhibits superior robustness to variations in observation frequencies.
- 中文摘要
事件摄像头生成异步高频数据流,以比传统更低的延迟提供空间稀疏信息。http URL 原则,这些属性应当非常适合控制设计。该 HTTP URL 的强化学习研究仍然有限,因为现有方法未能充分利用传感器。基于 http URL 的方法通过将事件聚合成稀疏网格,抵消了传感器带来的优势。这将计算成本与传感器分辨率耦合,同时使时间信息变得模糊。与此同时,现有的生成基线依赖于轨迹数据的可用性来预训练模型。我们提出FLEET(通过高效分词化从事件中学习特征),这是一种直接处理事件序列的特征提取器。利用随机傅里叶特征和交叉注意力,我们的架构将变量流压缩为固定大小的潜在表示。这使特征提取器骨干的推断成本与传感器分辨率得以实现端到端学习,无需辅助损耗。我们在新的高通量基准测试上验证了FLEET。结果表明,基于序列的方法超越了SOTA的性能,并且对观测频率的变化表现出更优异的鲁棒性。
Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning
问、条件还是弃权:缺失前提推理的强化学习
- Authors: Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2608.16554
- Pdf link: https://arxiv.org/pdf/2608.16554
- Abstract
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
- 中文摘要
仅答案强化学习(RL)训练推理模型来解决完全指定的问题,但许多现实的查询缺少唯一答案所需的前提。在这种情况下,有用的回答不总是拒绝:模型应要求缺失前提,以未知数为条件,或在没有信息性条件回答时弃权。我们介绍了\emph{Ask-Condition-Abstain Reinforcement Learning}(ACA-RL),这是一个针对该环境的数据增强强化学习框架。其推理图引导流水线将良好定题转换为带有局部缺口注释的缺失前提训练实例;ACA-RL随后对这些实例进行训练,并对五种可观察的反应行为进行结构化奖励。我们还介绍了\emph{缺失前提基准}(MPB),这是一个包含274个实例的人工验证基准,涵盖数学、逻辑和现实世界的应用题。在Qwen3和Llama模型中,ACA-RL在保持良好推理任务的竞争性能的同时,持续提升MPB性能。结合已发布的代码、MPB和训练数据,这项工作支持了NLP评估的新使命:衡量模型是否能识别任务未确定并处理不确定性,而不仅仅是能否回答完全指定的问题。
Interactive Whole Slide Images for RL-based Tumour Segmentation
用于基于强化学习的肿瘤切割的交互式全片图像
- Authors: Mohamad Mohamad, Francesco Ponzio, Maxime Gassier, Nicolas Pote, Xavier Descombes
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2608.16607
- Pdf link: https://arxiv.org/pdf/2608.16607
- Abstract
Whole-slide image (WSI) analysis remains computationally challenging due to the extremely large spatial resolution of slides and the sparse distribution of tumour regions. We propose an end-to-end reinforcement learning framework for sequential tumour segmentation directly on WSIs. Instead of treating the slide as a predefined collection of candidate patches, we formulate the WSI itself as a hierarchical multi-resolution environment through which an agent navigates using movement, zooming, and tumour selection actions. The agent jointly processes local observations and a global thumbnail representation within an actor-critic architecture trained using proximal policy optimization (PPO). Experiments on pulmonary adenocarcinoma WSIs demonstrate the feasibility of direct sequential tumour segmentation on full slides, achieving comparable coarse segmentation quality relative to patch-based approaches operating at similar magnification levels, while reducing inference time to a few seconds per slide. We further analyse the impact of environment design and action-space granularity. Our results suggest that modelling WSIs as interactive environments provides a promising direction for RL-based computational pathology
- 中文摘要
由于切片空间分辨率极高且肿瘤区域分布稀疏,整片图像(WSI)分析在计算上依然具有挑战性。我们提出了一个端到端强化学习框架,用于直接基于 WSI 进行的顺序肿瘤分割。我们不将切片视为预定义的候选贴片集合,而是将 WSI 本身定义为一个层级多分辨率环境,代理通过移动、缩放和肿瘤选择动作进行导航。智能体在通过近端策略优化(PPO)训练的演员-批评者架构中,共同处理局部观察和全局缩略图表示。肺腺癌 WSI 实验证明了在全切片上直接顺序肿瘤分割的可行性,相较于类似放大倍率下的贴片方法,粗分割质量相当,同时将推断时间缩短至每片几秒。我们还进一步分析了环境设计和动作空间细度的影响。我们的结果表明,将 WSI 建模为交互环境,为基于强化学习的计算病理学提供了有前景的方向
A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory
基于RFID支持的智能工厂的车间生产排程案例
- Authors: Zhihui Chen, Yize Sun, Yuhao Dong, Zeyu Xiao, Ray Y. Zhong
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16626
- Pdf link: https://arxiv.org/pdf/2608.16626
- Abstract
Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, which, in turn, can be used to support dynamic shop floor production planning and scheduling. Within such an environment, uncertainty in operation and production processes collectively contribute to the dynamicity in manufacturing, thereby hampering the scheduling system from achieving maximal utility. To highlight the importance of handling such uncertainty, this paper addresses the problem of dynamic shop floor scheduling for a real-life case smart factory equipped with RFID technology. Feasible production sequence mining and real-time processing rate estimation are conducted on RFID-collected production data to quantify the operation and production uncertainties. A deep reinforcement learning approach based on the RFID data analysis is then presented for shop floor production scheduling. Simulation studies based on real-life case data have demonstrated the feasibility and practicality of the proposed dynamic production scheduling framework. Specifically, it is observed that the proposed framework outperforms existing dispatch methods in terms of minimizing operation makespan, including first in first out (FIFO), last in first out (LIFO) and deep Q network (DQN).
- 中文摘要
射频识别(RFID)技术已被广泛应用于制造车间的实时数据收集,进而支持动态的车间生产计划和排程。在这样的环境中,运营和生产过程的不确定性共同加剧了制造的动态性,从而阻碍了排程系统实现最大效用。为了强调处理这种不确定性的重要性,本文探讨了配备RFID技术的真实案例智能工厂中动态车间排程问题。通过RFID收集的生产数据进行可行的生产序列挖掘和实时加工速率估算,以量化运营和生产不确定性。随后,基于RFID数据分析的深度强化学习方法被用于车间生产排度。基于真实案例数据的模拟研究已证明所提动态生产调度框架的可行性和可行性。具体来说,所提框架在最小化运算完成时段方面优于现有派遣方法,包括先入先出(FIFO)、后进先出(LIFO)和深度Q网络(DQN)。
Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents
Chronocooked:强化学习代理隐式区间时序基准
- Authors: Amrapali Pednekar, Alvaro Garrido-Perez, Yara Khaluf, Pieter Simoens
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16666
- Pdf link: https://arxiv.org/pdf/2608.16666
- Abstract
This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspired by Overcooked, the suite comprises cooking scenarios that require temporal decision making. The tasks and reward functions are designed such that temporal information is unobserved yet critical for optimal performance. The environment is intentionally kept simple to enable controlled experiments and support biologically plausible models. Evaluation metrics are designed to expose limitations in timing abilities of RL agents, and we report baselines using a non-recurrent, a recurrent, and a biologically plausible model. This work ultimately aims to underscore the need to incorporate time perception and temporal processing in artificial agents designed for human robot interaction and deployment in time dependent human societies.
- 中文摘要
本文介绍了Chronocooked,一套用于研究强化学习代理隐性区间时序的基准测试套件。该套装灵感来自《过度烹饪》,包含需要时间决策的烹饪场景。任务和奖励函数的设计使得时间信息虽未被观察,但对最佳表现至关重要。环境被有意保持简洁,以便进行受控实验并支持生物学上合理的模型。评估指标旨在揭示强化学习代理时间能力的局限性,我们采用非重复模型、重复模型和生物学合理模型报告基线。这项工作最终旨在强调,在为人类机器人交互和部署而设计的人工智能体中,将时间感知和时间处理纳入必要的必要性。
Le Critique: Privileged Value Functions for LLM Reinforcement Learning
Le Critique:LLM强化学习中的特权价值函数
- Authors: Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.16739
- Pdf link: https://arxiv.org/pdf/2608.16739
- Abstract
Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. Group-relative methods like GRPO reduce gradient variance by sampling multiple rollouts per prompt, but provide only sequence-level credit. Training is also blocked by straggler rollouts, reducing throughput and increasing off-policyness. Learned value functions theoretically address both problems, providing token-level advantages without requiring large groups. However, additional infrastructure engineering challenges combined with the practical success of critic-free methods have made it difficult to justify their inclusion in RL pipelines. We propose two complementary strategies to improve the performance of value function RL: 1) Privileged Value Functions (PVF) which provide an elegant mechanism to inject additional task-relevant token-level signal without biasing the policy objective; 2) TETHER, a baseline that adaptively interpolates between group-relative and value baselines depending on the value function accuracy. Across several reasoning tasks, both strategies consistently improve over the standard value function baseline, and are competitive with or outperform mean-baseline GRPO.
- 中文摘要
大型语言模型(LLMs)的强化学习算法主要区别于其方差缩减策略。像GRPO这样的组相对方法通过每个提示抽样多个展开来降低梯度方差,但只能提供序列层级的积分。培训还因落后部署而受阻,降低吞吐量并加剧了离策状态。学得的价值函数理论上解决了这两个问题,提供代币层面的优势,而无需大规模群体。然而,额外的基础设施工程挑战加上无批评方法的实际成功,使得将其纳入强化学习流程变得困难。我们提出了两种互补策略以提升价值函数强化学习的性能:1)特权价值函数(PVF),提供一种优雅的机制,在不偏袒策略目标的情况下注入额外的任务相关代币级信号;2)TETHER,一种基线,根据价值函数的准确性,在群体相对基线和价值基线之间自适应插值。在多个推理任务中,这两种策略均持续优于标准值函数基线,且在与平均基线GRPO的竞争性或优于平均值基线GRPO。
ClawGym II: Exploring Black-Box RL on Agent Harness
ClawGym II:探索黑盒强化游戏中的特工安全带
- Authors: Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.16798
- Pdf link: https://arxiv.org/pdf/2608.16798
- Abstract
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.
- 中文摘要
智能体机束通过协调智能体与环境的交互,显著提升了长期任务的性能。然而,通过复杂工具进行强化学习仍大多未被探索,因为将此类训练扩展到长期代理任务会带来根本性挑战。本研究提出了一个统一的黑箱强化学习框架,用于通过复杂工具实现通用代理的稳定且可扩展的优化。具体来说,我们首先构建一个基于沙盒的执行基础设施,将任务环境隔离,并在临时沙盒中利用这些资源,以实现大规模并发部署。然后我们将策略优化与不透明的机束执行解耦,并在模型边界放置服务代理以捕获模型调用。为了重建多回合轨迹并提高训练效率,我们将捕获的调用组织为前缀树,并进一步调整基于批评者的PPO和无批评的GRPO,以优化恢复后的树结构。同时,我们在整个优化过程中保持训练-推理一致性。最后,我们引入了混合束训练,允许单一模型通过异构束进行联合优化。通过Qwen3-30A3B,Black-box RL在ClawGym-Bench上分别Pass@1提升了9.98点和14.81分,同时在200-400个优化步骤内保持稳定。此外,该框架在更具挑战性的任务如JobBench和OfficeQA上持续带来收益。总体而言,我们的框架通过黑箱工具实现通用代理的有效、稳定和可扩展优化,支持跨异构执行系统的统一训练。
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
HAF:通过层级动作流和光谱潜在强化学习,将通用VLA适应为类人生物全身机动操作
- Authors: Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang
- Subjects: Subjects:
Robotics (cs.RO); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2608.16837
- Pdf link: https://arxiv.org/pdf/2608.16837
- Abstract
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: this https URL .
- 中文摘要
类人机器人作为以人为中心的环境中的通用智能体具有巨大潜力,但通用的视觉-语言-行动(VLA)基础模型并不容易适用于人形全身机车操作。高维度和类人生物动作的相互依赖性使传统单级VLA架构难以有效协调运动、腰部姿势和双臂操作。此外,通过离线行为克隆训练的策略在实际部署中可能保持次优状态。尽管在线强化学习可以通过现实世界互动来优化策略,但直接调优大型VLA骨干需要大量计算,并在真实机器人探索过程中可能带来安全风险。为解决这些瓶颈,我们引入了HAF(类人适应框架),这是一个由HAF-VLA和HAF-Steer组成的两部分框架,将现成的通用VLA基础模型转移到人形全身机车操作中。HAF-VLA是一个基于预训练流量匹配VLA构建的分层动作流生成器。它将全身动作去噪分为三个顺序阶段,通过阶段嵌入和跨阶段的KV缓存保留运动学依赖,避免单次生成时出现的非相干全体动作。在冻结的HAF-VLA之上,HAF-Steer是一个潜在的离线到在线强化学习流水线,利用流量匹配可逆性和基于DCT的降维技术,将强化优化限制在紧凑的噪声子空间内,并训练正则化的SAC策略。这避免了更新大型VLA骨干网,并实现高效的实际策略优化。HAF基于七项真实世界类人机车操作任务进行评估,超越了普通单阶段VLA基线,提升了全身协调和任务表现。项目网站:此 https 网址。
Q-based Variational Inverse Reinforcement Learning
基于Q的变分逆强化学习
- Authors: Ondrej Bajgar, Peter Tisnikar, Alessandro Abate, Konstantinos Gatsis, Maike Osborne
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2608.16888
- Pdf link: https://arxiv.org/pdf/2608.16888
- Abstract
The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, from expert behaviour. We introduce Q-based Variational IRL (QVIRL), a novel Bayesian IRL method that recovers a posterior distribution over rewards from expert demonstrations via primarily learning a variational distribution over optimal Q-values. Unlike previous approaches, QVIRL combines scalability with uncertainty quantification, important for safety-critical applications as well as active learning. We demonstrate QVIRL's strong performance in apprenticeship learning across various tasks, including gridworlds, Lunar Lander, the Highway Environment, and two ATARI games both with static expert data and with active learning. It is the first method for Bayesian IRL that demonstrates training from raw pixel observations.
- 中文摘要
安全且有益的人工智能的发展要求系统能够根据人类偏好学习和行动。然而,手动明确指定这些偏好往往不可行。逆强化学习(IRL)通过从专家行为推断偏好(以奖励函数表示)来应对这一挑战。我们介绍基于Q的变分IRL(QVIRL),这是一种新颖的贝叶斯IRL方法,通过主要学习Q值的变分分布,从专家演示中恢复奖励的后验分布。与以往方法不同,QVIRL结合了可扩展性和不确定性量化,这对安全关键应用和主动学习都至关重要。我们展示了QVIRL在学徒学习中多项任务的强劲表现,包括网格世界、月球着陆器、公路环境以及两款ATARI游戏,这些游戏均采用静态专家数据和主动学习。它是第一个展示从原始像素观测中训练的贝叶斯真实学习方法。
Keyword: diffusion policy
Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration
计划者条件扩散用于协调多智能体探索
- Authors: Marcus Yu Siong Teo, Jeric Lew, Tanishq Duhan, Guillaume Sartoretti
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.16229
- Pdf link: https://arxiv.org/pdf/2608.16229
- Abstract
Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant coverage across agents over extended planning horizons. Conventional approaches rely on hand-crafted coordination rules, while end-to-end multi-agent learning methods are difficult to scale and train. Diffusion-based planners such as DARE offer a promising alternative by generating long-horizon trajectories instead of single-step actions, but existing methods are trained on a narrow planner distribution, limiting behavioral diversity and inference-time controllability. We propose a Planner-Conditioned Diffusion Policy (PCDP) for graph-based multi-agent exploration. PCDP is trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same observation. Rather than learning coordination end-to-end, we reuse this multimodal single-agent policy across all agents and introduce coordination through local reranking, in which nearby agents jointly select the trajectory combination with minimal predicted overlap. We evaluate PCDP against classical and diffusion-based baselines on 100 held-out maps in a four-agent simulation setting. PCDP matches the perfect success rate of the diffusion-based baselines while improving mean max-agent travel, total team travel, and agent imbalance. Crucially, reranking alone over a single-planner baseline yields only marginal gains, indicating that planner-conditioned multimodality is the main contributor to improved coordination. Qualitative simulation results and real-robot experiments with two agents further validate that diverse long-horizon trajectory generation produces emergent spatial separation between agents without any explicit repulsion mechanism.
- 中文摘要
协调多代理探索不仅需要高效的个别覆盖,还需要跨代理在较长规划时间内实现非冗余覆盖。传统方法依赖手工定制的协调规则,而端到端多智能体学习方法则难以扩展和训练。基于扩散的规划器如DARE通过生成长视野轨迹而非单步行动,提供了有前景的替代方案,但现有方法训练于狭窄的规划器分布,限制了行为多样性和推理时间的可控性。我们提出了一种用于基于图的多智能体探索的规划者条件扩散策略(PCDP)。PCDP基于多种规划者风格的演示进行训练,规划者身份作为显式条件输入,使单一共享模型能够学习多模态轨迹分布,并从同一观测中生成多样且可控的轨迹候选。我们不再从端到端学习协调,而是在所有代理中重复使用这种多模态单代理策略,并通过局部重新排序引入协调,即邻近代理共同选择预测重叠最小的轨迹组合。我们在四剂模拟环境中,结合经典和基于扩散的基线,在100张保留图谱中评估PCDP。PCDP与基于扩散的基线完美成功率相匹配,同时改善平均最大代理旅行、总团队行程和代理不平衡。关键是,单独在单一规划者基线上重新排序仅带来边际收益,表明规划者条件多模态是提升协调性的主要因素。定性模拟结果和两位智能体的真实机器人实验进一步验证了多样的长视角轨迹生成能够在智能体之间产生无明确排斥机制的空间分离。
MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning
MatchingPolicy:对应感知策略支持跨对象上下文学习
- Authors: Qijin She, Hanyang Yu, Zeming Li, Ping Tan
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2608.16715
- Pdf link: https://arxiv.org/pdf/2608.16715
- Abstract
In-context imitation learning enables few-shot policy generalization but struggles to maintain performance on unseen objects and novel scenarios. To address this, we introduce MatchingPolicy, a correspondence-driven framework that explicitly decouples demonstration-to-scene matching from policy learning. Central to our method is a correspondence-aware diffusion policy that conditions robotic actions directly on dense semantic correspondences. This architectural separation resolves the inherent conflict between correspondence identification and action adaptation, enabling robust out-of-distribution transfer. Our framework integrates vision foundation models with a novel two-stage matching algorithm to dynamically establish reliable correspondences. Extensive evaluations on RLBench and real-world manipulation tasks confirm that MatchingPolicy achieves superior few-shot performance, generalizing reliably across unseen object instances and semantic categories.
- 中文摘要
上下文模仿学习实现了少数策略泛化,但在看不见的物体和新颖场景下难以维持性能。为此,我们引入了MatchingPolicy,一个以对应为驱动的框架,明确将演示到场景匹配与策略学习解耦。我们方法的核心是对应感知扩散策略,直接将机器人动作条件化为密集语义对应。这种架构分离解决了对应识别与动作适应之间的固有冲突,实现了稳健的分发外传输。我们的框架将视觉基础模型与一种新型两阶段匹配算法集成,动态建立可靠的对应关系。对RLBench和现实操作任务的广泛评估证实,MatchingPolicy实现了更优的少数样本性能,能够可靠地推广到未见对象实例和语义类别。