生成时间: 2026-10-08 23:20:12 (UTC+8); Arxiv 发布时间: 2026-10-08 20:00 EDT (2026-10-09 08:00 UTC+8)
今天共有 55 篇相关文章
Keyword: reinforcement learning
Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation
自适应工作流程智能:基于上下文的企业自动化认知架构
- Authors: Sreedevi Pandiyath Viswambaran
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2610.08793
- Pdf link: https://arxiv.org/pdf/2610.08793
- Abstract
Enterprise systems increasingly rely on automated workflows, yet many AI-driven solutions remain brittle under non-stationary conditions, evolving policies, and delayed operational feedback. While reinforcement learning and large language model (LLM) agents offer partial adaptability, they do not by themselves provide persistent reflection mechanisms or straightforward integration with policy-constrained enterprise operations. This paper introduces Adaptive Workflow Intelligence (AWI), a cognitive architecture for context-driven enterprise agents organized around a four-layer Perception-Cognition-Action-Reflection (PCAR) loop. AWI treats reflection as a mechanism for continuous policy refinement and combines hybrid reasoning with reflective memory and feedback-driven adaptation to support decision making under environmental drift and operational constraints. We evaluate AWI in a simulated enterprise decision workflow characterized by delayed outcomes and a controlled regime shift. In a drift-and-delay stress test, guardrail-constrained adaptive approaches recover more rapidly than static automation while maintaining policy compliance. Within this setting, AWI's reflective components modestly reduce behavioral oscillation and feedback variance, illustrating the stability-agility trade-off introduced by reflective policy adaptation.
- 中文摘要
企业系统日益依赖自动化工作流,但许多AI驱动的解决方案在非平稳条件、策略演变和操作反馈延迟下仍显脆弱。虽然强化学习和大型语言模型(LLM)代理提供部分适应性,但它们本身并不能提供持久反射机制或与策略限制企业运营的直接集成。本文介绍了自适应工作流智能(AWI),这是一种基于四层感知-认知-行动-反思(PCAR)循环的上下文驱动企业智能体认知架构。AWI将反思视为持续策略细化的机制,结合混合推理、反思记忆和反馈驱动的适应,以支持环境漂移和运营约束下的决策。我们在一个以延迟结果和受控体制转变为特征的模拟企业决策工作流程中评估AWI。在漂移延迟压力测试中,受保护栏约束的自适应方法在保持策略合规性的情况下恢复速度比静态自动化更快。在此设定下,AWI的反射组件适度减少了行为振荡和反馈方差,说明了反射策略适应带来的稳定性与敏捷性权衡。
HydroSphere: A Framework for Governed, Self-Healing Wastewater Infrastructure
HydroSphere:一个治理自愈污水基础设施框架
- Authors: Prabu, Fancy C, Suresh A, Srini Ramaswamy
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.08819
- Pdf link: https://arxiv.org/pdf/2610.08819
- Abstract
Rapid industrialization and urban growth are increasing pressure on water quality and wastewater treatment systems, while conventional treatment plants often rely on static monitoring and control strategies that cannot easily adapt to changing pollutant conditions. This paper presents HydroSphere, a governed, data-driven framework for real-time water quality monitoring, forecasting, treatment optimization, and fault recovery. HydroSphere is evaluated using 2.82 million water-quality measurements collected between 1940 and 2023. The framework integrates three main components. First, a hybrid TCN-LSTM model performs multi-step forecasting across seven water-quality parameters, achieving an RMSE of 0.1417, MAE of 0.1047, and R2 of 0.3596. Second, the Adaptive Dosage Optimization Module uses PPO reinforcement learning to adjust chemical dosing, achieving a mean step reward of 1.059 compared with 1.017 for a fixed-dose baseline. The results also show that unconstrained reward optimization can lead to excessive dosing, demonstrating the need for explicit operational safeguards. Third, the SHADE anomaly detection module uses a deep autoencoder to identify sensor and process anomalies, achieving an F1 score of 0.651 under controlled fault injection. HydroSphere combines these capabilities with tiered governance, deterministic safety bounds, and human oversight to support safer and more adaptive water infrastructure. The framework provides a scalable foundation for intelligent wastewater management and supports the objectives of UN Sustainable Development Goals 6 and 13.
- 中文摘要
快速工业化和城市发展正在加剧水质和污水处理系统的压力,而传统处理厂往往依赖静态监测和控制策略,难以适应不断变化的污染物条件。本文介绍了HydroSphere,一个受控、数据驱动的实时水质监测、预测、处理优化和故障恢复框架。HydroSphere基于1940年至2023年间收集的282万次水质测量数据进行评估。该框架整合了三个主要组成部分。首先,混合TCN-LSTM模型在七个水质参数上进行多步预测,实现RMSE为0.1417,MAE为0.1047,R2为0.3596。其次,自适应剂量优化模块利用PPO强化学习调整化学剂量,实现平均步进奖励为1.059,而固定剂量基线为1.017。结果还表明,无限制奖励优化可能导致过量投药,表明需要明确的操作保障措施。第三,SHADE异常检测模块使用深度自动编码器识别传感器和处理异常,在受控故障注入下获得0.651的F1评分。HydroSphere将这些能力与分层治理、确定性安全界限和人工监督相结合,支持更安全、更具适应性的水资源基础设施。该框架为智能污水管理提供了可扩展的基础,支持联合国可持续发展目标6和13的目标。
JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation
JoyAI-Voice 2.0:一个全连续自回归语音生成模型,具备语义-声学联合表示
- Authors: Yafeng Chen, Boya Dong, Yankun Huang, Hao Li, Jingdong Li, Xiangyu Liang, Hao Ni, Wenchao Wang, Yuxuan Wang, Zhangyu Xiao, Wei Deng, Nan Duan, Yu Gu, Wenhao Guan (Intern), Weisheng Han, Yabin Li, Yuan Liu, Jiaxin Ye (Intern), Fan Yu, Lin Zhu
- Subjects: Subjects:
Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
- Arxiv link: https://arxiv.org/abs/2610.08834
- Pdf link: https://arxiv.org/pdf/2610.08834
- Abstract
We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.
- 中文摘要
我们介绍JoyAI-Voice~2.0,一种端到端拟人化语音生成模型,构建在完全连续的双编码器架构之上。原始语音被编码成连续潜在部分并划分为补丁。每个补丁由语义-声学双编码器分解为语义纯化的表示和声学表示,这些结合并联合送入因果自回归变换器进行规划。变换器预测下一个补丁的条件,局部扩散变换器渲染其全潜在值以实现48,kHz合成。模型通过联合流匹配和停止预测目标训练,随后通过DiffusionNFT进行监督式微调和强化学习,以提升模型性能。在Seed-TTS上,平均词误率为2.51%,相较最强基线下降14.9%,在InstructTTSEval中中文和英语属性准确度均达到最先进,10个感知维度中有5个领先,MDVD-Eval总分最高为0.893。
Autonomous Droplet Navigation via Model-Based Reinforcement Learning: Zero-Shot Transfer and Emergent Dynamics
基于模型的强化学习实现自主液滴导航:零射点转移与涌现动力学
- Authors: Rajneesh Anand, Mayuresh V. Kothare
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.08852
- Pdf link: https://arxiv.org/pdf/2610.08852
- Abstract
Self-driving laboratories (SDLs) are transforming chemical and materials discovery through closed-loop automation, yet automated infrastructure for physical manipulation of soft, deformable matter remains beyond current robotic platforms. A critical instance is autonomous droplet transport on an open surface, where contact-angle hysteresis, capillary pinning, and surface heterogeneity produce partially observable dynamics that pose significant challenges for classical model-based controllers. We introduce the first robotic platform for closed-loop autonomous liquid droplet navigation on an open, unconfined surface using model-based reinforcement learning. A two-axis tilting board coated with a thin silicone oil film drives the droplet, while an overhead camera provides real-time feedback. A learned policy was trained on just 50 to 150 physical episodes depending on geometric complexity, without simulation or analytical models. Beyond performance alone, the platform demonstrates three capabilities of interest to the SDL community: it robustly transfers zero-shot to unseen geometries; it autonomously discovers an oscillatory depinning strategy to free the droplet when it sticks; and it completes its full training pipeline in under 90 minutes. These results extend reinforcement-learning manipulation from rigid microrobots to deformable soft-matter systems for next-generation SDLs.
- 中文摘要
自动驾驶实验室(SDLs)通过闭环自动化正在改变化学和材料的发现,但用于物理操作软体、可变形物质的自动化基础设施仍超出现有机器人平台的范围。一个关键例子是开放表面上的自主液滴传输,接触角滞后、毛细固定和表面异质性产生部分可观测的动力学,给经典模型控制器带来了重大挑战。我们推出了首个基于模型的强化学习,用于开放、无约束表面的闭环自主液滴导航机器人平台。涂有薄硅油膜的双轴倾斜板驱动液滴,而顶置摄像头提供实时反馈。根据几何复杂度,仅在50至150个物理事件上训练了一套学习策略,无需模拟或分析模型。除了性能之外,该平台还展示了三项令SDL社区感兴趣的能力:它能稳健地将零射向未知几何体;它自主发现振荡去钉策略,在液滴粘附时释放;并且在90分钟内完成完整训练流程。这些结果将强化学习操作从刚性微型机器人扩展到下一代SDL的可变形软物质系统。
AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
AdaGuard:通过支持推理的LLM作为评判者保护栏,提升安全和政策合规性
- Authors: Melissa Kazemi Rad, Sihui Dai, Isha Slavin, Kushal Chawla, Mann Patel, Jian Ni, William M. Campbell, Stephen Rawls, Sambit Sahu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.08923
- Pdf link: https://arxiv.org/pdf/2610.08923
- Abstract
Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency
- 中文摘要
企业生成式AI应用需要强大的安全机制,能够适应多样化的风险态势、不断演变的策略和不同的延迟约束。当前的护栏解决方案常常存在僵化问题,依赖固定的策略集,且透明度和推理灵活性有限。我们介绍Adaguard,一个自适应的LLM即法官框架,旨在通过动态策略执行和自适应推理预算分配来应对这些挑战。AdaGuard采用监督微调(SFT)和强化学习(GRPO)构建,能够在运行时推广到用户定义的安全与合规策略,无需频繁更新模型。我们方法的核心创新之一是能够动态推断输入-策略对的复杂性,使模型能够在高速黑箱推理和可解释、推理支持的调节之间切换。这种灵活性使开发者能够在严格的延迟要求与可操作的透明度之间取得平衡。这种自适应能力使 AdaGuard 能够与体积数倍的其他护栏和前沿型号竞争,而其自动推理模式则以极低的延迟恢复了始终在线推理的准确性
On KL-Regularized Policy Optimization
关于KL正则化策略优化
- Authors: Yifan Zhang
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.08963
- Pdf link: https://arxiv.org/pdf/2610.08963
- Abstract
Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler's own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO. The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.
- 中文摘要
对于大型语言模型(LLM)代理的异步强化学习(RL)将一种策略训练于另一种策略生成的轨迹:展开来自过时检查点,推理引擎的概率即使参数相同,也与训练器不同。标准的解决方法是通过剪辑重要性比(Clip Importance ratio)来偏向更新,或者像GRPO一样,每个提示采样一组响应,当集数较长时成本较高。我们提出了KL正则化策略优化(KLPO)框架,将KL正则化器锚定在采样器上。正则化改进步骤采用封闭式Gibbs解,KLPO通过采样器自身轨迹的最小二乘来满足其对数比最优条件,因此采样器概率通过对数比进入,不需要重要权重。剖析回归截距将难以处理的对数划分函数替换为信号采样均值加上采样器到训练器间的KL散度。对于令牌级策略镜像下降目标,我们证明所得梯度可通过终端返回计算,无需批评者,通过采样器中心分数或单一轨迹残差计算,即使在随机工具输出下。我们进一步证明,独立的蒙特卡洛估计的KL项保持这些梯度无偏,推导出更便宜的顶峰$K美元和二元近似的精确KL差距,并证明SPPO、GPO、REBEL和BPO是KLPO的特例。结果是一个无批评的更新,每个提示使用一次rollout,且不需学习归一化器或一组响应。
HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids
浩克:学习人形生物全身强制操控
- Authors: An Dang, Arturo Flores Alvarez, Yu-Ming Chen, Conor Mc Gartoll, Helen Sun, Aaron Ames, Nima Fazeli, Manikantan Nambi
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.08970
- Pdf link: https://arxiv.org/pdf/2610.08970
- Abstract
Humanoid loco-manipulation of large, heavy objects demands forceful interaction across the entire body. However, such payloads shift a humanoid's center of mass and impose sustained loads across the upper body, challenging balance and command tracking. We present HULK, a whole-body control framework for forceful loco-manipulation. Using model predictive control (MPC) to guide reinforcement learning with predictions of the loaded dynamics, we train two teachers: one tracks arm motions under wrist forces, and the other locomotes while holding large objects against the body. A capture-point control barrier function augments the wrist-force teacher during training to improve balance under load. We distill both teachers into a single policy. Evaluation spans simulation and the Unitree G1. In simulation, the teacher with the barrier function achieves the lowest forward and lateral velocity tracking errors at 10 kg per arm among evaluated controllers and reduces aggregate divergent component of motion (DCM) excursion magnitude by 35.7% relative to MPC-guided reinforcement learning alone. Our wrist-force teacher withstands torso push disturbances of up to 130 N.
- 中文摘要
类人机车操控大型、重物需要全身的强力互动。然而,此类载荷会移动类人生物的质心,并在上半身施加持续负荷,挑战平衡和指令追踪。我们介绍HULK,一种全身控制框架,用于强制机车操作。利用模型预测控制(MPC)引导强化学习,预测负载动力学,我们培训两位教师:一位在手腕受力下追踪手臂动作,另一位在手持大型物体时跟踪机车。捕获点控制屏障功能在培训中增强腕力教师,改善负荷平衡。我们将两位教师整合成一个统一策略。评估涵盖模拟和Unitree G1。在模拟中,使用障碍函数的教师在评估控制者中实现了每臂10公斤的前进和横向速度跟踪误差最低,相较于单独MPC引导强化学习,将总发散运动分量(DCM)的振幅幅度降低了35.7%。我们的腕部力量教师承受了高达130牛顿的躯干推力干扰。
LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL
LASER:支持约束熵正则化离线强化学习的潜在空间伴随匹配
- Authors: Songyuan Zhang, Oswin So, Eric Yang Yu, Matthew Cleaveland, Peter Crowley-Dolen, Chuchu Fan
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Optimization and Control (math.OC); Machine Learning (stat.ML)
- Arxiv link: https://arxiv.org/abs/2610.08989
- Pdf link: https://arxiv.org/pdf/2610.08989
- Abstract
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: this https URL.
- 中文摘要
虽然离线强化学习(RL)能够在不昂贵的在线交互的情况下从静态数据集中进行策略优化,但仍存在执行分布外(OOD)动作的风险。近期方法通过流匹配学习行为克隆策略,然后在其受限的潜在空间内执行强化学习来缓解这一问题。然而,过于简单地优化潜在策略很容易使策略崩溃为脆弱模式,或利用学习批评者的尖锐伪影。本研究发现熵正则化在潜在空间强化学习中对于应对这些挑战至关重要。我们引入了LASER,一种新型离线强化学习算法,应用潜在空间伴随匹配实现具有表达性流策略的熵正则化潜空间强化学习,同时避免时间反向传播。通过对40个具有不同数据集质量的挑战性OGBench任务进行全面实验,我们证明LASER实现了最先进的性能。值得注意的是,LASER在所有任务中使用固定方法特有的超参数,表现优于评估基线,包括针对任务和数据集调整的基线,这凸显了LASER的强大适用性。项目网站:此 https URL。
TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery
TAP:通过轨迹锚定恢复实现高效的长视野代理剪枝
- Authors: Yuanzhe Li, Pengxin Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Song Wang, Jingdi Chen, Huanrui Yang
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09074
- Pdf link: https://arxiv.org/pdf/2610.09074
- Abstract
Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student's response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents' task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.
- 中文摘要
新兴的长期代理任务需要反复的模型调用,进一步加剧了本已昂贵的语言模型的推理成本。虽然狭窄的代理任务暗示了激进的模型剪枝而不会导致性能下降,但实证结果显示,现有的问答任务方法在应用于代理模型时严重降低了任务性能。我们将这一失败归因于两个决策:修剪哪些内容以及如何恢复。对于修剪,一次性重要性估计无法追踪修剪模型的适应情况。对于恢复,离线蒸馏仅覆盖教师前缀,而全轨迹策略蒸馏则使学生错误在回合间叠加。本研究提出轨迹锚定剪枝(TAP),这是首个针对强化学习(RL)训练代理的结构修剪框架。TAP 将结构修剪与高效的策略恢复相结合,将互动锚定于教师轨迹,同时允许学生生成每一个推理-行动响应。冻结的密集教师监督学生的响应前缀,解决响应内的训练-推理不匹配,同时防止学生引发的偏差跨越训练回合传播。TAP 不是一次性修剪,而是利用恢复目标的梯度重新评分通道,将迭代通道选择与策略演变连接起来。移除 60% FFN 通道后,TAP 保留了 ALFWorld 和 WebShop 上密集 7B 代理任务成功率的 99.2% 和 88.0%,同时每个成功任务的 GPU 时间减少约 22% 和 17%。这些结果表明在有限的恢复预算下,长期视野代理的结构压缩有效。
Convex-Concave Reinforcement Learning
凸凹强化学习
- Authors: Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09108
- Pdf link: https://arxiv.org/pdf/2610.09108
- Abstract
Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[\pi/\pi_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.
- 中文摘要
策略学习推动了当今许多最重要且投资最重的强化学习应用。然而,其核心优化问题(最大化期望收益)以非凸性著称,即使在直接策略参数化下,领域也大多避免凸性:在信任区域约束(NPG、TRPO、PPO、AWR)下优化收益的凸代理近似。我们证明了这一看似非结构化的问题实际上并非无结构。在对数密度比坐标$y := \log[\pi/\pi_n]$中,精确的每次迭代目标(可通过每次决策重要性抽样(PDIS)计算,是一个凸差约束凸差值(DC-约束DC)程序。该结构使我们超越代理近似:它将CPI、NPG、TRPO和AWR作为可解释轴的特例恢复,并开启一个多步轴 $k$,将连续决策耦合。我们用序列凸规划(SCP)求解逐迭代程序,SCP是凸差问题的标准求解器,并在温和条件下提供收敛保证,桥接凸差优化与强化学习文献。经验上,多步凸凹RL(CCRL)在诊断MDP中胜出,在信用必须跨越视野传播(其优势随依赖长度增长)的MDP中获胜,在经典对照下可与调优PPO竞争,且在现实随机的中期视野医疗领域收敛速度明显快于调优PPO达到相同的近似最优存活率,训练曲线下面积增加11.3%。
The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning
欺骗性强盗问题:探索性耦合与多智能体学习的脆弱性
- Authors: Michael Tang, Mahmoud Abdelgalil, Jorge I. Poveda
- Subjects: Subjects:
Machine Learning (cs.LG); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.09120
- Pdf link: https://arxiv.org/pdf/2610.09120
- Abstract
Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these properties are critical for security purposes and demonstrate how an adversarial agent can exploit privileged information on another agent's exploration. We analyze a deceiver-victim pair in the minimal two-player strongly monotone setting, where a deceptive player obtains leaked signals that are merely correlated with the victim's exploration. We show that, by coupling their own exploratory action with this information, the deceptive player injects an externality that steers the learning dynamics to a new steady state, called the deceptive Nash equilibrium (DNE). We prove that the deceptive bandit learning (DBL) dynamics converge to an arbitrarily small neighborhood of the DNE while retaining optimal convergence rates. Interestingly, our analysis attains these optimal rates while relaxing second-order smoothness conditions from standard bandit optimization literature. We characterize conditions under which deception strictly shifts the steady state and its effect on the deceiver's cost, illustrating the results in a resource-allocation game.
- 中文摘要
随机探索是盗贼学习、多智能体强化学习和零阶策略搜索的核心,但其独立性和隐私通常仅被视为技术假设。我们证明这些属性对安全至关重要,并展示了对抗智能体如何利用另一代理探索中的特权信息。我们分析了在极小双人强单调环境中的欺骗者-受害者配对,欺骗者玩家获得的泄露信号仅与受害者的探索相关。我们表明,通过将自身的探索行为与这些信息耦合,欺骗玩家注入了一种外部性,引导学习动态进入一种新的稳态,称为欺骗性纳什均衡(DNE)。我们证明欺骗性盗贼学习(DBL)动态在保持最优收敛率的同时收敛到DNE的一个任意小邻域。有趣的是,我们的分析在放宽标准盗贼优化文献中的二阶平滑性条件的情况下,达到了这些最优速率。我们描述了欺骗严格改变稳态及其对欺骗者成本影响的条件,展示了资源分配博弈中的结果。
Mission-critical spectrum sharing with decentralized Multi-Agent Reinforcement Learning
通过去中心化多智能体强化学习实现关键任务频谱共享
- Authors: Dimitrios Pylorof, Imtiaz Nasim, Humberto E. Garcia, Vivek Agarwal, Jasni A. Mannil, Mingyue Ji
- Subjects: Subjects:
Systems and Control (eess.SY); Information Theory (cs.IT); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2610.09213
- Pdf link: https://arxiv.org/pdf/2610.09213
- Abstract
Motivated by emerging mission-critical applications and an increasingly congested spectrum, we develop a decentralized multi-agent reinforcement learning (MARL) model for dynamic spectrum access. The model enables secondary users to learn effective transmission strategies across shared frequency bands while minimizing collisions with high-priority primary users and among themselves. We design the agent-level learners following a Markov potential game approach, connecting independent local updates to system-level improvement. We instantiate this design using lightweight linear actor-critic learners suitable for resource-constrained edge devices, rather than computationally intensive centralized or deep multi-agent architectures. Across spectrum environments with different incumbent activities, the learned policies adapt their transmission policy and waiting behavior to preserve throughput while greatly reducing transmission collisions relative to random and forecast-aware heuristic baselines. The results establish the value of decentralized MARL and shows up to 96.8% reduction in overall collisions.
- 中文摘要
受新兴关键任务应用和频谱日益拥挤的推动,我们开发了一种去中心化的多智能体强化学习(MARL)动态频谱访问模型。该模型使次级用户能够在共享频段学习有效的传输策略,同时最大限度地减少与高优先级主用户及内部的碰撞。我们采用马尔可夫潜在博弈方法设计代理级学习器,将独立的本地更新与系统级改进连接起来。我们采用适合资源受限边缘设备的轻量级线性actor-critic学习器实现该设计,而非计算密集型集中式或深度多代理架构。在不同既有活动的频谱环境中,学习策略会调整传输策略和等待行为以保持吞吐量,同时大幅减少相较于随机和预测感知启发式基线的传输碰撞。结果确立了去中心化MARL的价值,并显示整体碰撞减少了高达96.8%。
AGAR: a reinforcement learning substrate for LLM program evolution
AGAR:用于LLM程序演进的强化学习基底
- Authors: Haoran Li, Zengle Ge, Xiaomin Yuan, Yui Lo, Haoxin Li, Songlin Zhou, Qianhui Liu, Jiahua Ying, Yuanhang Liu, Mingju Chen, Annan Li, Jianmin Wu, Dawei Yin, Dou Shen
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
- Arxiv link: https://arxiv.org/abs/2610.09215
- Pdf link: https://arxiv.org/pdf/2610.09215
- Abstract
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, rather than the program it emits. Credit assignment, value estimation, adaptive exploration, and experience memory can then attach to distinct components. AGAR (Algorithm Generation As RL) provides the resulting substrate: any estimator can be replaced or switched off without changing the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends, and three seeds under one harness, AGAR improves on the stronger of two published baselines on most tasks, with gains concentrated in the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice, but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.
- 中文摘要
给定一个任务和一个评估器,语言模型可以重写候选程序,而搜索循环决定哪些重写得以存活,这为算法发现提供了实用的路径。但该循环由五个常数手动控制:选择哪个父节点、突变难度、如何保持多样性、记忆内容,以及一个标量分数,但从不说明程序哪个部分获得了该分数。强化学习已经为每个部分提供了估计量。障碍在于程序演化通常不会被写成决策过程。我们将它形式化为马尔可夫决策过程,其作用是模型所依赖的模块前缀,而非它发出的程序。积分分配、价值估计、自适应探索和经验记忆随后可以附加到不同的组件上。AGAR(算法生成即强化学习)提供了最终的基底:任何估计器都可以替换或关闭,无需更换控制器,使传输可逐一机制进行审计,无需对后端模型进行梯度训练。在19个任务、两个后端和三个种子的框架下,AGAR在大多数任务上优于两个已发布基线中较强的,且收益集中在竞争性编程领域。形式化还提供了对先前工作的可检验解读:这些系统隐含为零折现,这并非出于选择,而是因为适应度对个体而言是外生的,而非对继任者的回报,因此折扣因子无可操作。
RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
RLDISCOVER:基于LLM驱动的强化学习算法共进化
- Authors: Haoran Li, Zengle Ge, Xiaomin Yuan, Yui Lo, Songlin Zhou, Jiahua Ying, Haoxin Li, Qianhui Liu, Yuanhang Liu, Jiaqun Liu, Guokai Chen, Mingju Chen, Ruinan Wang, Annan Li, Jianmin Wu, Dawei Yin, Dou Shen
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09218
- Pdf link: https://arxiv.org/pdf/2610.09218
- Abstract
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
- 中文摘要
LLM引导的程序演进推动了数学和计算优化领域的发现,带来了强化学习(RL)算法自我演化以改善智能体学习能力的前景。然而,实现这一前景面临两个障碍。耦合算法组件上的联合搜索难以扩展:同时变化会干扰学习,而孤立变化则忽视其依赖关系。评估候选算法也需要昂贵的训练,随机种子间的适应度仍不确定。我们介绍RLDiscover,一个用于无模型深度强化算法自我演化的框架。渐进共进化从目标组件编辑向联合进化推进,而渐进概率评估通过分阶段训练和反复评估平衡搜索广度和评估准确性。在四个基准套件上,SAC、PPO和DQN的实验显示平均回报显著提升,每家族中位数增益为32%-84%,峰值回报率约为363倍,且较接近零的基线值提升约363倍。这些提升包括从失败学习到任务成功完成的转变,且当更强的开源实现开始演进时,改进依然存在。在测量的SAC移动运行中,评估使用约为完整评估同一候选池所需计算量的十五分之一。令人惊讶的是,独立搜索反复发现可解释的自适应稳健损失、进度依赖值目标和运行统计的组合,且部分程序会迁移到未被发现的任务。这些发现指向人工智能自我演化的更广泛作用:发现可解释的算法,改善智能体的学习能力。
FLoRa: Flight-Assisted Data Collection from Duty Cycling LoRa Nodes under Energy Constraints
FLoRa:在能量约束下,从值班循环LoRa节点的飞行辅助数据收集
- Authors: Naresh Babu Kakarla, V. Mahendran
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09226
- Pdf link: https://arxiv.org/pdf/2610.09226
- Abstract
Data collection using Unmanned Aerial Vehicles (UAVs) is challenging when LoRa IoT Devices (IoTDs) duty-cycle to conserve battery. Under energy constraints, a UAV must decide which IoTDs to visit, in what order, where to hover, and how many times to probe each node, while time-based data freshness decays. Tractably solving this problem requires a multi-level optimization architecture: discrete combinatorial optimization for routing, continuous global optimization for spatial positioning, and sequential decision-making under uncertainty. We propose FLoRa, a Flight-assisted LoRa data collection architecture using Simulated Annealing (SA) for path planning, Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for hover positioning, and Partially Observable Markov Decision Processes (POMDPs) for probing IoTDs. To quantify collection utility from duty-cycling nodes, we introduce the Value of Information for Pull-based systems (VIP), a metric that rewards fresh data and penalizes failed probes, imposing well-posedness and preventing indefinite probing when an IoTD is off. Tracking hard battery constraints on every POMDP sample path requires state augmentation, worsening the curse of dimensionality. For tractability, SA and CMA-ES work on the hard battery constraints, while at the POMDP layer we relax them into soft average constraints via Lagrangian relaxation. Since solving the network-wide POMDP is computationally complex, we decompose it into node-level POMDPs by approximating inter-node time dependency using a forward-decomposition technique. Evaluation shows FLoRa outperforms metaheuristic, greedy, and deep reinforcement learning baselines by 30.6%, 27.8%, and 15.2% in total expected VIP, while increasing node coverage by 24.5%, 29.2%, and 8.8%, and successful collections by 24.3%, 25.6%, and 15.2%, respectively.
- 中文摘要
当LoRa物联网设备(IoTDs)采用工作周期以节省电池时,使用无人机进行数据收集具有挑战性。在能量限制下,无人机必须决定访问哪些IoTD、顺序、悬停地点以及探测每个节点的次数,同时基于时间的数据新鲜度会逐渐下降。解决这一问题需要多层次优化架构:离散组合优化用于路由,连续全局优化空间定位,以及在不确定性下顺序决策。我们提出了FLoRa,一种飞行辅助的LoRa数据收集架构,采用模拟退火(SA)进行路径规划,协方差矩阵适应演化策略(CMA-ES)进行悬停定位,部分可观测马尔可夫决策过程(POMDPs)用于探测IoTD。为了量化占空循环节点的收集效用,我们引入了基于拉取系统的信息价值(VIP),这是一个奖励新数据、惩罚失败探针的指标,施加良定性,防止在IoTD失效时进行无限探测。追踪每个POMDP采样路径的硬电池约束需要状态增强,这加剧了维度的诅咒。为了可解性,SA和CMA-ES处理硬电池约束,而在POMDP层,我们通过拉格朗日松弛将其松弛为软平均约束。由于解决网络范围的POMDP计算复杂,我们通过前向分解技术近似节点间时间依赖性将其分解为节点级POMDP。评估显示,FLoRa在预期VIP总量中分别领先元启发式、贪婪和深度强化学习基线30.6%、27.8%和15.2%,节点覆盖率提升24.5%、29.2%和8.8%,成功采集率分别提升24.3%、25.6%和15.2%。
An Informational Curse of Horizon in Goal-Conditioned Policy Learning
目标条件政策学习中地平线的信息诅咒
- Authors: John L. Zhou, Yuxuan Dong, Jonathan C. Kao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.09247
- Pdf link: https://arxiv.org/pdf/2610.09247
- Abstract
The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.
- 中文摘要
学习目标达成策略的难度通常归因于“视野诅咒”,表现为时间差备份中的偏见积累和噪声优势估计。本研究指出目标条件政策学习中视界的另一个信息诅咒,即提高目标重新标记视野可显著降低策略泛化和性能。通过一系列与预言机规划器的受控实验,我们将训练期间抽样的目标视野与策略测试时要求达到的目标视野解耦。即使仅基于一系列邻近子目标进行评估,目标条件行为克隆(BC)策略仍存在严重的训练视野依赖性能下降,而强化学习(RL)目标可予以缓解。我们将这一现象解释为行动与事后重标目标之间条件互信息的相近性下降,并实证发现,无论是基于较长视野目标训练的BC和强化学习策略,都表现出从目标信息到状态信息的敏感性转变,这一点通过策略输入的雅可比式来衡量。基于这一观察,我们发现将短期视野政策的输入雅可比量提炼为长视野策略,在组合操作任务中,性能显著提升。综合来看,我们的结果凸显了目标重新标记视野在从离线数据中学习通用策略时的重要考量。
Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks
动态OWC网络中的合作对抗DQN SAC学习以实现能源效率
- Authors: Walter Zibusiso Ncube, Ahmad Adnan Qidan, Taisir El-Gorashi, Jaafar M. H. Elmirghani
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.09266
- Pdf link: https://arxiv.org/pdf/2610.09266
- Abstract
Growing wireless traffic is increasing pressure on the congested radio-frequency spectrum. Optical wireless communication (OWC) provides a complementary solution by using the abundant unlicensed optical spectrum. However, indoor OWC networks are dynamic: users move, enter or leave the network, and subscribe to different services. Poorly coordinated resource allocation can consequently waste subcarriers, require excessive transmission power and cause frequent AP reassignments. Energy efficiency (EE), defined as the total delivered data rate divided by the total network power consumption, therefore requires the serving AP, number of allocated subcarriers, and transmission power to be jointly adapted while maintaining QoS. Optimising these in a dynamic time series, multi-service OWC environment produces a complex sequential EE optimisation problem. To address this problem, this work proposes Dual-Agent Resource Allocation using Deep Reinforcement Learning (DARA-DRL). DARA-DRL combines a branching duelling deep Q-network for association and subcarrier allocation with a conditional soft actor-critic agent for continuous power control. The agents are coupled through a common reward and a cooperative value update that evaluates each discrete allocation together with its corresponding power decision. Simulation results show that DARA-DRL remains within 5\% of the optimal solution and, compared with state-of-the-art benchmarks, it improves EE by 28.8\% and QoS satisfaction by 10.5\%, while reducing online decision time by 14.5\%. Results demonstrate that agent specialisation simplifies mixed-action learning, and cooperation outperforms independently trained agents.
- 中文摘要
不断增长的无线流量正在加剧拥塞的射频频谱压力。光纤无线通信(OWC)通过丰富的无许可光谱提供了互补解决方案。然而,室内OWC网络是动态的:用户移动、进入或离开网络,并订阅不同的服务。资源分配协调不当可能导致子载波资源浪费,需要过多的传输功率,并导致频繁的AP重新分配。能量效率(EE)定义为总传输数据速率除以网络总功耗,因此需要在保持服务质量的同时共同调整服务AP、分配的子载波数量和传输功率。通过动态时间序列优化这些,多服务OWC环境产生了一个复杂的顺序EE优化问题。为解决这一问题,本研究提出了利用深度强化学习(DARA-DRL)实现双代理资源分配。DARA-DRL结合了分支对抗深度Q网络用于关联和子载波分配,以及条件软行为者-批评代理实现连续功耗控制。代理通过共同奖励和协同价值更新耦合,后者评估每个离散分配及其对应的功率决策。模拟结果显示,DARA-DRL的精度仍保持在最优解的5%范围内,与最先进基准相比,它提升了28.8%的EE和10.5%的QoS满意度,同时在线决策时间减少了14.5%。结果表明,智能体专业化简化了混合动作学习,协作表现优于独立训练的智能体。
Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense
专家问答:基于LLM的自主网络防御强化学习
- Authors: Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen
- Subjects: Subjects:
Cryptography and Security (cs.CR); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09337
- Pdf link: https://arxiv.org/pdf/2610.09337
- Abstract
Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.
- 中文摘要
基于策略的强化学习(RL)方法已为自主网络防御取得了有前景的成果;然而,在防御者必须在延迟、部分观测下响应大型行动空间的环境中,它们的样本效率较低。虽然大型语言模型(LLMs)可以语义性推理安全状态空间,但高延迟和信任假设阻碍了吸引人的内联部署模型。我们介绍了“问专家”(Ask the Expert),这是一个训练时间指导框架,先总结硬性网络防御状态,然后通过受限动作接口间歇性查询LLM的主机级防御建议,最终将这些建议转化为用于PPO的分层奖励塑造。由于LLM在训练后被丢弃,部署是纯粹的强化学习策略。在TTCP CAGE CC1和CC2以及两种攻击者类型中,这种非对称设计提高了相对于PPO的样本效率,并优于评估的基于势能的奖励整形(PBRS)基线,同时保持最强的终端均值,且部署时无需依赖LLM。
Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization
通过最优性和反事实正则化学习未知约束而不使用不安全数据
- Authors: Zhouyu Zhang, Chih-Yuan Chiu, Glen Chou
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09350
- Pdf link: https://arxiv.org/pdf/2610.09350
- Abstract
Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.
- 中文摘要
从演示中学习(LfD)提供了一个框架,用于从局部最优、满足约束的专家行为推断未知约束。现有方法主要分为两种范式:受限逆最优控制(CIOC)和逆受限强化学习(ICRL)。CIOC利用了如Karush-Kuhn--Tucker(KKT)条件等最优条件,但通常假设已知动力学和结构化约束表示。与此同时,ICRL兼容复杂的未知约束和未知过渡动态,但通常需要大量在线探索,期间可能发生不安全的约束违规。本研究介绍了反事实KKT(CF-KKT),这是一种约束学习框架,利用学习动力学和局部最优演示,在无需已知动力学或额外风险探索的情况下恢复未知约束,结合了CIOC的数据效率和安全优势与ICRL的灵活性。首先,我们使用局部学习的可微动力学模型,直接在演示中施加受KKT启发的最优条件。其次,利用学习到的动力学在演示附近生成提升奖励的反事实行为,揭示在未知约束缺失时更优的行为,从而提供合成不可行数据。当约束参数化已知时,同一学习动力学框架支持基于CIOC的直接参数恢复,并表征其对动力学错误的敏感性。在高维机器人控制任务中,我们的方法相较于最先进的离线ICRL基线,学习神经约束表示并提升安全性和数据效率。
Learning Stability of Replay-Based Co-Optimization for Transmission Expansion under Strategic Bidding
基于重放的协同优化在战略竞价下传输扩展的学习稳定性
- Authors: Tomonari Kanazawa, Hikaru Hoshino, Eiko Furutani
- Subjects: Subjects:
Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.09366
- Pdf link: https://arxiv.org/pdf/2610.09366
- Abstract
This paper investigates the behavior of learning-based co-optimization for transmission expansion under strategic bidding in electricity markets. In this framework, transmission capacities are updated while market participants simultaneously learn their bidding strategies through deep reinforcement learning, resulting in coupled and non-stationary learning dynamics. We show that transient policy degradation of bidding agents can generate inconsistent cost-capacity samples, which bias the transmission-capacity update and prevent the co-optimization process from converging to the desired solution. To mitigate this, replay-based capacity updates with nearest-neighbor filtering are introduced to exclude inconsistent samples from the update data. We then analyze a new oscillatory behavior that appears when the replay memory is enlarged. Numerical results on the IEEE 30-bus system demonstrate that enlarged replay memories improve robustness against transient policy degradation but can introduce a temporal lag, leading to oscillations unless the capacity-update learning rate is appropriately reduced. These results reveal the trade-off between replay-memory size and learning rate and provide practical guidelines for stable co-optimization.
- 中文摘要
本文探讨了基于学习的协同优化在电力市场战略竞价下输电扩展的行为。在该框架下,输电容量在市场参与者通过深度强化学习同时学习竞标策略的同时进行更新,从而形成耦合学习和非平稳学习动态。我们表明,投标代理的暂时策略降级会产生不一致的成本-容量样本,从而偏向输电容量更新,阻碍协优化过程趋于理想解。为缓解这一问题,引入了基于重放的容量更新,采用最近邻滤波,以排除不一致的样本进入更新数据。随后,我们分析了回放内存扩大时出现的新振荡行为。IEEE 30总线系统的数值结果表明,放大的重放记忆增强了对瞬态策略退化的鲁棒性,但可能引入时间滞后,除非适当降低容量更新学习率,否则会导致振荡。这些结果揭示了回放记忆大小与学习率之间的权衡,并为稳定协优化提供了实用指导。
TutorLoop: Regulating Student Learning Behaviors via Sensor-in-the-Loop Generative Feedback
TutorLoop:通过传感器在环生成反馈调节学生学习行为
- Authors: Songlin Xu, Xinyu Zhang
- Subjects: Subjects:
Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09400
- Pdf link: https://arxiv.org/pdf/2610.09400
- Abstract
We present TutorLoop, a sensor-in-the-loop system that regulates student learning behaviors by delivering adaptive feedback based on real-time cognitive states. Unlike prior large language model (LLM) tutors that directly depend on scenario-specific content, TutorLoop operates on sensor-derived signals captured via webcams. Moreover, unlike direct cognitive-to-feedback mappings that are short-sighted, the system employs a deep reinforcement learning (DRL) agent to optimize the feedback type across the entire learning process. Finally, another LLM tutor refines feedback into human-like, context-aware messages. We evaluate TutorLoop in a large-scale user study (N=187), where a model trained offline is directly applied to a new learning task without retraining. Results show that TutorLoop provides less frequent yet more effective interventions, improving attention, reducing workload, increasing engagement, and ultimately enhancing learning outcomes. These findings highlight the potential of closed-loop, sensor-driven feedback for scalable human-AI integrated systems to support learning.
- 中文摘要
我们介绍了TutorLoop,一个传感器在环系统,通过基于实时认知状态的自适应反馈来调节学生的学习行为。与以往直接依赖场景特定内容的大语言模型(LLM)导师不同,TutorLoop基于摄像头捕捉的传感器信号运行。此外,与短视的直接认知到反馈映射不同,该系统采用深度强化学习(DRL)代理,优化整个学习过程的反馈类型。最后,另一位LLM导师将反馈细化为类人、具上下文感知的信息。我们在一项大规模用户研究(N=187)中评估了TutorLoop,该研究将离线训练的模型直接应用于新的学习任务,无需重新训练。结果显示,TutorLoop提供了更少频率但更有效的干预,提升注意力、减轻工作负荷、提高参与度,最终提升学习成果。这些发现凸显了闭环、传感器驱动反馈为可扩展的人机集成系统支持学习的潜力。
TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding
TiTok:多段时间接地视听大型语言模型
- Authors: Eunji Shin, Dahyun Choi, Seungyeon Jo, Yejin Hong, Jiyoung Lee
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.09408
- Pdf link: https://arxiv.org/pdf/2610.09408
- Abstract
Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.
- 中文摘要
未修剪视频中的视听多段接地(AV-MSG),即对视听证据进行推理并预测查询的多个片段,是一个根本性问题,但仍然具有挑战性。仅视觉模型忽视互补的声学线索,而视听模型常常无法校准事件数量——我们称之为计数校准错误。我们介绍TiTok,一种视听大型语言模型(AV-LLM),它为每个查询定位任意数量的时间事件段。为了精确预测边界,我们引入时间令认交(TTI)方法,该方法明确向视听流注入特殊时间令码,使输入端的时间感知与输出端时间预测对齐。我们还提出了解耦、多段导向的强化学习奖励,包括全局、局部、计数、精确和格式奖励,并采用组奖励解耦规范策略优化(GDPO)进行优化。为评估AV-MSG的性能,我们建立了基于UnAV-100的新评估协议,并提出CountF1指标用于量化重叠指标无法捕捉的计数错误校准。TiTok达到65.7 mIoU,CountF1达到0.58,实现了最先进的性能。我们的代码可在此链接获取。
Precise SE(3) End-Effector Tracking in Whole-Body Humanoid Control
全身类人生物控制中的精确SE(3)末端效应器跟踪
- Authors: Joohwan Seo, Xiaofeng Guo, Jinkun Cao, Roberto Horowitz, Rocky Duan, Guanya Shi, Koushil Sreenath
- Subjects: Subjects:
Robotics (cs.RO); Systems and Control (eess.SY)
- Arxiv link: https://arxiv.org/abs/2610.09479
- Pdf link: https://arxiv.org/pdf/2610.09479
- Abstract
Precise end-effector tracking during humanoid whole-body motion is challenging due to floating-base oscillations, gravity, dynamic coupling, and locomotion-induced disturbances. We propose ResGAC, a whole-body humanoid controller for precise end-effector pose tracking that combines geometric admittance control (GAC) with residual reinforcement learning. GAC provides structured $\SE$ task-space feedback and generates nominal arm joint-position targets, while residual RL compensates for unmodeled dynamics and coordinates locomotion and balance in the shared joint-position action space. The left-invariant geometric formulation allows the same GAC law to be used across manipulation reference frames. This enables the use of a ground-attached heading frame that preserves planar locomotion while removing pelvis roll, pitch, and heave from the manipulation reference, thereby reducing reference-induced end-effector motion during locomotion. ResGAC is validated on a real Unitree G1 humanoid. Across four standing end-effector tracking benchmarks, ResGAC consistently outperforms representative baselines, including SONIC, achieving lower translational and rotational errors. Real-world experiments further demonstrate reduced propagation of pelvis motion to the desired end-effector pose using the proposed ground-attached heading frame. ResGAC achieves $90\%$ success in a standing peg-in-hole task compared with $50\%$ for SONIC, and accurate world-frame $\SE$ end-effector pose tracking during lower-body motion. Experimental videos are included in the supplementary material and are also available on the project website: this https URL.
- 中文摘要
由于漂浮基底振荡、重力、动态耦合和运动引起的扰动,在人形全身运动中精确跟踪端部执行器具有挑战性。我们提出了ResGAC,一种全身型人形控制器,用于精确的末端执行器姿态跟踪,结合了几何导纳控制(GAC)和残差强化学习。GAC提供结构化的$\SE$任务空间反馈,并生成标称臂关节位置目标,而残差RL则补偿未建模的动力学,并在共享关节位置作用空间中协调运动与平衡。左不变几何表述允许在操作参考系间使用相同的GAC定律。这使得使用地面连接的航向框架能够保持平面运动,同时消除操作参考中的骨盆滚动、俯仰和隆重,从而减少运动过程中参考引起的末端执行器运动。ResGAC在真实的Unitree G1类人生物上得到了验证。在四个站立末端执行器跟踪基准测试中,ResGAC持续优于代表性基线,包括SONIC,实现更低的平移和旋转误差。真实实验进一步展示了使用拟建地面连接航向框架,骨盆运动传播到期望末端执行器姿态的传播减少。ResGAC在站立钉入洞任务中成功率为90美元,而SONIC为50%%,且在下半身运动时实现了准确的世界帧端执行器姿态追踪。实验视频包含在补充材料中,也可在项目网站获取:此 https URL。
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
MIMESIS:学习用户模拟器作为交互代理的训练环境
- Authors: Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng, Wancen Mu, Yue Cao, Shengjie Bi, Yun He, Changdae Oh, Deren Lei
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09484
- Pdf link: https://arxiv.org/pdf/2610.09484
- Abstract
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.
- 中文摘要
训练和评估交互式语言代理通常需要丰富的用户互动,但收集人类反馈既昂贵又难以扩展。模拟用户提供了可扩展的替代方案,但它们必须既接近真实用户行为,又必须为代理提供有用的学习体验。相比之下,大多数代理训练框架依赖现成的辅助LLMs,其实用性使其与真实用户过于合作、显式且行为趋于同质化。我们介绍了MIMESIS,这是一个专门构建的用户模拟器,基于人类对话训练,并有明确的推理监督和13种基于真实用户互动的真实行为模式训练。从实证角度看,我们的9B模型实现了65.7的SOUL-Index,超过了最强的前沿模型。与RealUserSim和SimulatorArena上最强基线Claude-Opus-5相比,MIMESIS分别提升了13.4个点的行为忠实度,并减少了3.6个点的图灵距离。然后,我们冻结模拟器,并通过多回合强化学习与冻结模拟器交互来训练代理。在八种环境中,使用MIMESIS训练的代理表现优于在所有九个未见用户模拟器下用GPT-5.5训练,展示了对新用户模拟器的更强泛化能力。此外,我们提出了教练策略自我蒸馏(CSD),利用模拟器生成的私有推理痕迹和后续话语作为对智能体满足用户需求的反馈。教练将这些信息转化为简明的教练笔记,描述代理如何更好地预判用户需求并在交互过程中调整行为。CSD将这些反馈转化为密集的代币级监督,超越稀疏的任务奖励,带来九个评估用户模型的进一步收益。
Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?
平均来说安全,尾巴不安全:什么时候可以控制分段消耗的尾巴?
- Authors: Samuel Tetteh, Cody Fleming
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09508
- Pdf link: https://arxiv.org/pdf/2610.09508
- Abstract
Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using $\mathrm{CVaR}{0.1}$, the average cost of the worst $10\%$ of episodes. We classify a policy as tail-safe when $\mathrm{CVaR}{0.1}$ is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.
- 中文摘要
安全强化学习寻求在满足累计成本约束的同时最大化回报的策略。大多数方法对预期的情节成本施加这些约束。因此,标准评估报告的是平均情节成本,却未描述成本如何在剧集中分布。满足平均成本标准的策略因此在其最差剧集中可能仍然不安全。平均成本报告既未识别这一尾部违规,也无法显示其是否能在保持回报的同时控制在预算内。本研究中,我们使用 $\mathrm{CVaR}{0.1}$(最差10%$集数的平均成本)来衡量情节成本尾部。当$\mathrm{CVaR}{0.1}$在安全预算内时,我们将该策略归类为尾部安全。这使我们能够首先识别平均安全但在尾部不安全的政策,然后研究其尾部违规行为是否可控制且保持回归。为识别尾部不安全政策,我们评估了三个安全体育馆导航任务中的五个标准算法。随后,我们考察了四个高密度危险导航的约束家族,并评估四个导航任务和四个移动任务中的尾部控制。
It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank
它没有意识到危险:一个冻结的视觉语言安全评分衡量了它的字幕库
- Authors: Samuel Tetteh, Cody Fleming
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09517
- Pdf link: https://arxiv.org/pdf/2610.09517
- Abstract
Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.
- 中文摘要
冻结视觉语言模型越来越多地为强化学习提供安全信号。它们的使用假设与描述危险的语言相似性表明了危险本身。然而,策略返回率和碰撞率无法揭示得分是否检测到危险或响应场景相关特征。基于VLM的方法报告通过将图像文本相似度转换为奖励、成本或置信权重,在驾驶和安全强化基准中取得了进展。此类信号有望减少对人工设计反馈的依赖。它们也可能反映提示结构、嵌入几何或摄像头视角,导致其安全意义尚未验证。为弥补这一空白,我们提出了对冻结CLIP提示-边际安全评分的受控评估。我们将该分数应用于从未收到该评分的策略生成的轨迹,将接触前观测与具有相似危害几何形状的无接触观测匹配,并调整字幕、编码器和摄像头视角。在三种策略、180集和130次孤立接触起始中,分数在接触前约二十步下降。机制控制显示,分数主要追踪与其字幕共享场景的相似度,以及随着字幕分离和摄像机视角的变化。恒定置信控制保留较低的灾难率点估计,因此策略的提升不构成危险感知。
EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs
EvoSignal:基于LLM的模块化交通信号控制程序的进化设计
- Authors: Leizhen Wang, Peibo Duan, Zhenlin Qin, Yancheng Ling, Jian Xu, Yue Wang, Hao Wang, Zhenliang Ma
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09563
- Pdf link: https://arxiv.org/pdf/2610.09563
- Abstract
Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rules for a target network. Large language models (LLMs) can automate this process, but directly using them to select signal phases leaves decision rules embedded in black-box models and incurs recurring inference costs and latency. This paper formulates TSC as a modular program design problem and proposes EvoSignal, an LLM-guided evolutionary framework using traffic knowledge and performance feedback. The modular representation separates traffic feature extraction, local phase prioritization, and optional network-based priority adjustment. Starting from several established strategies, EvoSignal improves programs through feedback on congestion and signal operation, retaining strategies with different performance trade-offs. The resulting programs operate without online LLM inference. Simulation experiments across five scenarios on two real-world road networks show that the selected default EvoSignal program reduces waiting time by 16.8--49.2\% relative to the lowest waiting time achieved by the 20 conventional, reinforcement learning-based, and LLM-based baselines in each scenario. A program prioritizing travel time and queue length outperforms all 20 baselines on all three metrics in the search scenario and remains among the top three on each metric when transferred unchanged to the other four scenarios. These findings support automated design of inspectable control programs that transfer across the evaluated road networks and traffic this http URL is available at this https URL.
- 中文摘要
有效的交通信号控制(TSC)需要策略能够响应变化的交通需求和网络状况,同时满足不同的控制目标。然而,调整现有策略通常需要反复手动设计和调整,这使得系统性地探索目标网络更好的控制规则变得困难。大型语言模型(LLMs)可以自动化这一过程,但直接使用它们选择信号相位时,决策规则会嵌入黑箱模型中,并产生持续的推理成本和延迟。本文将TSC表述为模块化程序设计问题,并提出了EvoSignal,这是一个基于LLM的进化框架,利用交通知识和性能反馈。该模块化表示将交通特征提取、局部相位优先级和可选的基于网络的优先级调整分离开来。基于几种既定策略,EvoSignal通过对拥堵和信号运行的反馈改进程序,保留具有不同性能权衡的策略。最终的程序无需在线LLM推断即可运行。在两个真实道路网络的五个场景中进行的模拟实验显示,所选默认的EvoSignal程序相较于每个场景中20个常规、基于强化学习和基于LLM基线的最低等待时间,减少了16.8-49.2%的等待时间。优先考虑行车时间和队列长度的程序在搜索场景中的三个指标上均优于所有20个基线,且在转移至其他四个场景时仍位列前三。这些发现支持自动设计可检测控制程序,这些程序能跨评估的道路网络和交通进行传输。该http URL可在此 https URL 访问。
SAPD: Step-Aligned Privileged Distillation
SAPD:阶梯对齐特权蒸馏
- Authors: Tianle Wang, Jiayu Liu, Ruizhi Zhao, Ning Miao
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.09665
- Pdf link: https://arxiv.org/pdf/2610.09665
- Abstract
On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at this https URL.
- 中文摘要
政策后培训可以通过学习大型语言模型自身轨迹来改进,但需要代价高昂的推广生成。我们探问固定示范是否能通过更好的监督支持竞争性非策略学习。我们的前提是,固定示范的有用性不仅取决于训练轨迹,还取决于督导是否能在延续中提供信息偏好,并将该指导与所学推理决策连接起来。我们介绍了步对齐特权蒸馏(SAPD),这是一种无需推广的自蒸馏方法,将演示转化为步进对齐分布式监督。其关键见解是利用参考解的已知进展,将每个推理转换与目标特权引导联系起来,而非将解算视为未区分的上下文。在数学推理基准测试中,SAPD平均优于监督微调和标签平滑,同时在策略内强化学习和自我蒸馏中保持竞争力。分析支持上下文依赖分布指导的价值以及将特权信息与当前步骤对齐的好处。SAPD还在很大程度上保留了域外编码性能,并实现了约2倍的训练循环加速,相较于策略内基线。这些发现表明,精心构建的监督可以使完全非策略的后训练成为一种具有竞争力且计算效率高的替代方案。我们的代码可在此 https 网址获取。
CERO: Where and When to Allocate Rollouts for RL Post-Training
CERO:在哪里以及何时分配强化学习后的推广时间
- Authors: Yiming Zong, Yige Wang, Xing Hu, Jiashuo Jiang, Zuo-Jun Max Shen
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09679
- Pdf link: https://arxiv.org/pdf/2610.09679
- Abstract
Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.
- 中文摘要
自适应展开方法用于群体相对强化学习,通常在提示之间分配固定的每次更新预算。我们转而研究如何在整个训练时间段协调有限的推广预算。我们利用累积提示暴露的凹替代工具来表述该问题,并引入了CERO——一种用于提示录取和预算节奏的在线原始双调度器。在我们的实验中,每个被接受的提示都获得一个固定大小的响应组。CERO则调整了选择哪些提示、每轮重复使用频率以及每轮生成的组数。紧凑的Fenchel表示线性化了对累计暴露的依赖,而预测的在线梯度下降则通过奖励-变异反馈和预算偏差,更新提示特定的支持斜率和共享预算价格。我们建立了针对固定速率和同路径时变基准的替代分配目标的路径保证,明确了代理差异和速率变异的术语。在匹配训练-响应预算下,CERO在五个数学推理基准测试中,三个骨干链均达到了最高的宏观平均avg@16。机制分析将CERO的提示选择与组内奖励对比联系起来,而多种子消融显示,适应性节奏相较于均匀和预设支出计划均有优势。
Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving
超越政策支持:互动受限的离线强化学习对自动驾驶
- Authors: Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.09763
- Pdf link: https://arxiv.org/pdf/2610.09763
- Abstract
Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: this https URL
- 中文摘要
离线强化学习使得从固定数据集中实现奖励驱动的策略改进,无需在线探索,因此在安全关键领域尤具吸引力。然而,一个核心挑战是分布转移:策略优化可能偏向离线数据支持较弱的行动,使价值估计变得不可靠。现有方法主要控制策略自身行动空间中的这种变化。在自动驾驶等交互环境中,这可能不足:候选自我轨迹可能在边际行为分布下得到良好支持,但与记录互动中观察到的周围代理行为支持较差。我们将这种交互支持的退化称为 \emph{交互分布转移}(IDS),并引入 \emph{交互约束驱动策略}(ICDP),一种显式控制交互级分布转移的离线强化学习框架。从自我和环绕代理未来对联合数据分布出发,我们证明联合支撑退化精确分解为自我支持成分和残余交互支持成分。我们通过对比密度-比估计恢复后者,隔离交互兼容性,无需显式联合密度建模、周围代理预测或在策略优化过程中在反应模拟器或学习世界模型中展开。对nuPlan、Interplan和真实卡车实验的闭环评估表明,ICDP抑制高价值但无交互支持的轨迹选择,并提升交互关键驾驶场景的性能。项目网页:此链接
BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
BoT-GRPO:通过代币袋聚合进行推理的高效过程-奖励强化学习
- Authors: Yingxiang Yang, Weihang Xiao, Zhunxuan Wang, Joshua Flashner, Niresh Agarwal
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09804
- Pdf link: https://arxiv.org/pdf/2610.09804
- Abstract
Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches $80\%$ compile rate up to $1.9\times$ faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@$k$ gains up to $8.1\%$ over GRPO in half the steps. For both tasks we compare the algorithm's performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.
- 中文摘要
强化学习现已成为大型语言模型中引发推理的核心,而在流行的算法Group Relative Policy Optimization(GRPO)中,部署中的每个代币都获得相同的优势。我们探讨如何使流程监督高效:加速收敛并提升最终质量,同时不承担价值网络的成本。我们提出了Bag-of-Tokens群相对策略优化(BoT-GRPO),它通过长度不变的“代币袋”聚合将GRPO扩展到代币级奖励模型:它收集所有跨推广的代币级奖励,按其源序列长度的倒数加权,并计算相对于加权组统计的每个代币优势。BoT-GRPO无批评,是代币级奖励可用时使用GRPO的临时替代。在React前端代码生成中,BoT-GRPO的编译速率比GRPO快达$1.9\times,收敛速度也快于现代GRPO变体(GSPO、DAPO、PURE),同时实现更高的最终编译和VLM判定胜率。在第二个任务AIME数学推理中,BoT-GRPO在半步内实现了绝对Pass@$k$的提升,高达8.1%$。在这两项任务中,我们比较了算法在推理与非推理基模型族(Qwen2.5-3B、SmolLM3-3B、Phi-4-mini-推理)上的表现。我们的实验还得出了奖励模型本身的实用配方:奖励稳定性比丰富度更重要:干净、有界、稳定的细粒度信号能持续加速学习,而噪声较大的替代方案则停滞。
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
SkillForge:通过动态技能生命周期共同进化技能与特工
- Authors: Yuyao Ge, Yiwei Wang, Yuchen He, Baolong Bi, Lingrui Mei, Jiayu Yao, Lizhe Chen, Shenghua Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09832
- Pdf link: https://arxiv.org/pdf/2610.09832
- Abstract
Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.
- 中文摘要
记忆增强强化学习增强LLM代理解决复杂长期任务的能力。技能就是其中一种记忆形式,将指令与任务类型适用条件结合起来。然而,随着策略改进,随意保留所有技能,会让过时或有害的条目积累并误导智能体。我们提出了SkillForge,一种智能化强化学习方法,通过适应度驱动的技能生命周期(试用状态、主动状态、稳定状态和退休状态)编译和演进技能库,使技能与模型在训练过程中共同进化。前RL评估阶段首先利用基础模型自身的推广预淘汰低适应度技能,产生过滤库,随后引入监督微调。强化学习接管该检查点,每次迭代中选择性退休、稳定化和LLM引导突变继续锻造技能库,同时策略优化。在多个交互代理基准测试中,SkillForge实现了最高的综合成功率,相较最强基线提升高达7.8%,同时在整个训练过程中保持技能库紧凑。我们引入SkillFurnace,这是一个包含5k+注释记录的数据集,包含退休过滤的SFT轨迹、带适应度注释的进化技能库,以及带有人工注释失败类别的退休事件,以支持技能质量和生命周期管理的研究。
Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents
自我演化,参考文献:工具集成代理的锚定训练
- Authors: Wenjie Liao, Liangjie Zhao, Zehong Cao
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09856
- Pdf link: https://arxiv.org/pdf/2610.09856
- Abstract
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5\% on mathematical reasoning and 2.8\% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.
- 中文摘要
自我演化的工具集成代理从自身训练循环内生成的任务和反馈中学习。课程代理生成任务,而执行者代理通过强化学习从自我一致性信号中学习。然而,仅依赖当前执行者获取反馈存在两个局限性:在完全共识下群体相对优势消失,而基于不确定性的课程奖励则支持分歧,但不显示生成任务是否支持进一步学习。这些局限促使当前执行者之外增加了一个额外的参考。我们提出 \textit{AnchorLoop},它引入前一迭代执行者的冻结副本作为历史参考,并在训练循环的两侧重复使用。对于执行者,锚点提供了交叉引用优势,能够将当前输出与当前和历史多数答案进行比较。对于课程,它基于抽样多数意见的差异提供基于协议的参考。由于执行者和锚点在课程培训中参数相同,该比较作为任务选择的代理指标,而非版本间改进或正确性的证据。在13个推理基准测试中,AnchorLoop在数学推理方面比Agent0提升2.5%,在一般推理任务上提升2.8%。它还保持更高的有效优势方差,并在后续迭代中持续提升,因为无锚基线的收益递减。这些结果表明,在无外部任务或答案监督下,向自我演化的工具集成代理引入轻量级历史引用的好处。
Training Advisors for LLM Agents from Task Outcomes
任务成果中的LLM代理培训顾问
- Authors: Sergei Polezhaev, Barys Liskavets, Ori Press, Alexander Golubev
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09858
- Pdf link: https://arxiv.org/pdf/2610.09858
- Abstract
Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $\tau^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
- 中文摘要
大型语言模型代理通过将推理和工具调用与环境观察交织来处理多步任务。先前研究表明,自然语言反馈可以帮助这些代理在任务执行时修正决策。我们介绍了Caddie,一种训练批评者在执行任务时提供自然语言分析和建议的方法。与依赖步骤级标签或引用批评的方法不同,Caddie通过代理在收到批评者反馈后是否最终成功来学习。我们通过强化学习优化批评者,同时保持基础模型冻结。通过单一基础模型进行多跳问答训练,我们的Qwen3-4B批评者在四个不同尺度和架构的基础模型上提升成功率,其中包括三个未被批评者训练中的模型。在MuSiQue基准测试中,训练有素的批评者将Qwen3-4B的成功率提升了25个百分点以上,超过了无批评者的Kimi K3。同一批评者在域外交互基准测试(如$\tau^3$和DeepDive)中也能获得提升,无需额外训练。我们的结果表明,代理可以在推理时决定何时寻求批评者的帮助,基于结果的批评者训练可以产生跨基模型和任务域转移的指导。
Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks
基于TSN的车辆边缘网络的截止时间感知多智能体强化学习
- Authors: Bernardo A. C. Pereira, Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Andreas Gavrielides, Johann M. Marquez-Barja, Daniel F. Macedo
- Subjects: Subjects:
Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09870
- Pdf link: https://arxiv.org/pdf/2610.09870
- Abstract
Vehicular edge computing (VEC) enables latency-sensitive applications by bringing computing and networking resources closer to vehicles. However, existing approaches often overlook network contention among co-located services with heterogeneous and dynamic latency requirements. While time-sensitive networking (TSN) provides bounded-latency communication, conventional and reinforcement learning-based schedulers struggle to adapt to highly dynamic vehicular environments and inter-queue dependencies. To address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC. Each TSN queue is assigned an autonomous agent that jointly learns the queue service order and time-slot duration to minimize deadline misses under speed-dependent latency requirements. We employ multi-agent proximal policy optimization (MAPPO) to enable coordinated yet autonomous scheduling decisions. Evaluation against single-agent, multi-agent, and non-learning-based baselines shows that MAPPO provides robust performance across different traffic profiles. Compared with centralized single-agent methods, it reduces service latency by up to 66.2% and improves reliability by up to 271.8%. Furthermore, unlike urgency-based heuristics, MAPPO ensures balanced scheduling while achieving lower inference times compared to other MARL methods.
- 中文摘要
车载边缘计算(VEC)通过将计算和网络资源更接近车辆,使延迟敏感的应用得以实现。然而,现有方法常常忽视了同址服务之间具有异构和动态延迟需求的网络争用问题。虽然时敏网络(TSN)提供了有边界延迟通信,但传统和基于强化学习的调度器难以适应高度动态的车辆环境和队列间依赖。为解决这些限制,我们提出了在TSN支持的VEC中用于队列级调度的多代理强化学习(MARL)方法。每个TSN队列都被分配一个自主代理,协同学习队列服务顺序和时隙持续时间,以在速度依赖延迟需求下最小化截止日期遗漏。我们采用多代理近端策略优化(MAPPO)实现协调且自主的调度决策。针对单代理、多代理和非学习型基线的评估显示,MAPPO在不同流量轮廓中提供了稳健的性能。与集中式单代理方法相比,MAPPO将服务延迟降低了高达66.2%,可靠性提升了高达271.8%。此外,与基于紧急性的启发式不同,MAPPO确保调度平衡,同时实现了比其他MARL方法更短的推断时间。
A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration
关于人机协作中基于人类反馈强化学习的范围范围综述与实验研究
- Authors: Alexandra Coroiu, Andrea Vogt, Viktor Werbilo, Andreas Poppele, Johann Christensen, Sven Hallerbach
- Subjects: Subjects:
Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.09891
- Pdf link: https://arxiv.org/pdf/2610.09891
- Abstract
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.
- 中文摘要
人机协作(HRC)可促进工业4.0中的大规模定制化,其中“人类反馈强化学习”(RLHF)代表了开发安全AI机器人的有前景方法。在人工智能开发过程中的安全、人类反馈质量及双向人机适应方面仍面临实际挑战。我们对HRC系统中的RLHF进行了范围综述,绘制了应对这些挑战的方法。遵循PRISMA指南,我们筛选了199条记录,纳入了20篇跨HRC领域(2020-2025年)同行评审的论文。据我们所知,这是首篇聚焦RLHF双向闭环设计的综述。我们的综述发现了多种反馈模式,支持多种反馈格式的数据收集。收集的数据可以整合到AI训练的不同阶段,形成多步骤开发过程。试点实验常用于基于人类和机器人指标评估HRC系统。为实证检验综述中的关键空白,我们进行了受试者间虚拟现实实验,比较系统与用户对机器人近端行为的反馈以实现安全导航。利用贝叶斯模型,我们分析了收集到的反馈与安全指标之间的关系:心理安全(实验后问卷)与物理安全(碰撞时间反比)。结果显示,用户主动反馈比系统主动反馈更能捕捉感知安全,表明反馈时机直接影响反馈质量。我们的综述和实验结果表明,RLHF依赖适当的反馈方法确保HRC中的AI安全性,未来的RLHF研究应优先考虑真实的HRC实验,评估反馈收集方法对相关人机指标的影响。
RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning
RollVerify:桥接长尾推广强化学习的效率与准确性
- Authors: Yongqiang Yao, Jinru Tan, Kaihuan Liang, Zixin Yin, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.09914
- Pdf link: https://arxiv.org/pdf/2610.09914
- Abstract
Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
- 中文摘要
强化学习对于提升大型语言模型的推理和泛化至关重要。它依赖于大规模的推展,随着上下文窗口的扩大,这些推展长度会变得越来越长尾。在策略内训练中,这些长尾推展可能导致GPU泡沫,降低系统利用率并限制强化学习的可扩展性。异步或部分展开方法通过放松同步来提高吞吐量,但不可避免地引入过时的非策略样本(轨迹),可能损害最终准确性。现有方法主要通过在训练中重新加权非策略样本来缓解这种非策略问题,但与完全策略内训练相比,这些方法仍可能存在性能差距。在本研究中,我们提出了RollVerify,这是一个基于部分部署的轻量级强化学习框架,能够主动验证和修复样本,使其进入训练前进行。具体来说,它引入了非策略偏移指标OPS,用于量化部分生成轨迹的非策略偏差。在OPS约束的指导下,RollVerify执行序列级和令牌级验证,以识别并截断轨迹的无效后缀。这产生了高质量的样本,既保护模型的准确性,又保持部分展开带来的效率提升。数学和工具辅助数学推理的实验表明,RollVerify在降低训练成本的同时,实现了与策略训练相当的准确性。额外的代码生成结果提供了超越数学的初步证据。
Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
成功方式多样:基于多样性驱动的强化学习微调以实现VLA泛化
- Authors: Haoru Li, Jinmei Liu, Zhiyong Wang, Xiaoming Li, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09943
- Pdf link: https://arxiv.org/pdf/2610.09943
- Abstract
Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $\pi_0$ and 2.0 points on $\pi_{0.5}$. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.
- 中文摘要
强化学习(RL)微调通过闭环体验改善视觉-语言-行动(VLA)策略,但超出微调分布的泛化仍然有限。我们的分析揭示了探索的选择性重塑:强化学习在全球范围内收缩行为,但多样化成功轨迹,以更少的展开引发成功,且比监督微调覆盖更多潜在任务有效解空间。更广泛的成功模式覆盖可能为分布转移提供替代策略。受此启发,我们引入了DRIVE(VLA能量化的多样性驱动RL fIne-tuning),将成功行为多样性转化为明确的强化学习目标。DRIVE在匹配任务条件下展开,比较其轨迹与时间对齐,并从相对行为多样性中获得成功条件的内在奖励。这种设计鼓励更广泛地覆盖可行解决方案,同时不奖励多样化的失败或表面时间差异。在 LIBERO-Plus、ManiSkill3 和 RoboTwin 2.0 中,DRIVE 在原版 RL 微调上平均域外(OOD)性能提升了 $\pi_0$ 5.3 点,$\pi_{0.5}$ 提升 2.0 点。在双臂 AgileX PiPER-X 平台上,DRIVE 将平均 OOD 成功率从 64.1% 提升至 73.3%(+9.2 分),展示了在实体部署下持续的提升。
Decoding Neural Population Dynamics through Robotic Analog
通过机器人模拟解码神经群体动力学
- Authors: Wenhui Chen, Jiyue Tao, Yitao Cheng, Yutong Shi, Feitian Zhang, Xitong Liang, Ke Liu
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.09977
- Pdf link: https://arxiv.org/pdf/2610.09977
- Abstract
Animal evidence shows that precise voluntary movements arise from rotational neural population dynamics in motor cortex, but their physical effects remain unknown. We developed a robotic analog of biological motor systems with artificial muscles, multimodal sensors, and a neural network controller trained via reinforcement learning. The robotic analog exhibited accurate movements, robustness to damage, and neural population dynamics akin to animals. This task-driven, embodied model illuminates the causal link between neural population dynamics and motor outcomes. We discovered that neural rotations generate oscillatory maneuvers orthogonal to the reaching direction, optimizing trajectory adjustments, which is confirmed by primate neural data. The model also revealed counterintuitive neural energy principles under sensor and motor redundancies, and striking Eureka moments during motor learning, bridging biological and artificial systems. These findings provide new perspectives on how neural dynamics contribute to accurate and flexible movement, inspiring future intelligent robots with animal-like mobility.
- 中文摘要
动物证据表明,精确的自愿运动源于运动皮层中的旋转神经群体动态,但其物理效应尚不清楚。我们开发了一种生物运动系统的机器人类比,配备人工肌肉、多模态传感器和通过强化学习训练的神经网络控制器。该机器人模拟体展现出准确的运动、抗损伤能力以及类似动物的神经群体动态。这一任务驱动的具身模型揭示了神经群体动态与运动结果之间的因果联系。我们发现神经旋转产生与伸展方向垂直的振荡动作,优化轨迹调整,这一观点得到了灵长类动物神经数据的证实。模型还揭示了传感器和运动冗余下的神经能量原理,以及运动学习中的“尤里卡时刻”,连接了生物系统与人工系统。这些发现为神经动力学如何促进准确且灵活的运动提供了新的视角,激励未来拥有动物般机动性的智能机器人。
Learning to Accumulate Knowledge with Mutual Information
学习通过互信息积累知识
- Authors: Yuyang Zhao, Lizi Liao, Leyang Shen, Xiaoyan Zhao, Yang Zhang, Fuli Feng, Xiangnan He
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10042
- Pdf link: https://arxiv.org/pdf/2610.10042
- Abstract
Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge. Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories. We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation. Together, these signals encourage the curator to preserve distinct information from experience and produce entries that improve task success when added to existing knowledge. Standalone success rewards also favor entries that are useful on their own. On ALFWorld and WebShop, Knowledge Weaver achieves mean success rates of 54.0\% and 42.0\% with k=10 retrieved entries, exceeding GRPO by 16.9 and 18.7 percentage points, respectively. Its knowledge banks also outperform the evaluated prompt-based and established banks, including human-written banks, in overall ALFWorld success rate and WebShop score with the executor frozen. Our codebase is available at this https URL.
- 中文摘要
大型语言模型(LLM)代理可以通过重复利用从过去互动中提炼出来的知识来提升性能。然而,将新经验整理成随着知识库增长而变得更有价值仍然具有挑战性。有效的知识积累应限制条目间的重复重叠,并确保新知识贡献超出银行现有的部分。然而,通过小组相对策略优化(GRPO)培训策展人进行独立任务成功的培训,即使重复了现有知识,也能强化通用指导。因此,我们提出了知识编织者(Knowledge Weaver)这一强化学习框架,训练语言模型以策划主体轨迹中的可复用知识。我们将基于代币互信息(MI)的反馈与边际成功奖励相结合,指导知识积累。这些信号共同鼓励策展人保留经验中的不同信息,并在现有知识中生成提升任务成功的条目。独立成功奖励也优先于单独有用的条目。在ALFWorld和WebShop上,Knowledge Weaver在k=10条检索条目时平均成功率为54.0%和42.0%,分别比GRPO高出16.9%和18.7个百分点。其知识库在整体ALFWorld成功率和WebShop得分上也优于已评估的基于提示和成熟的银行,包括人工编写的银行,执行者冻结时的评分。我们的代码库可在该网址访问。
RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
RewardWeaver:通过自我演化的奖励适应,为语言代理提供长期交互式学习
- Authors: Hengbo Xiao, Boyao Zhang, Purui Liu, Yuxuan Zheng, Haoran Yin, Haibo Liu, Fan Zhang
- Subjects: Subjects:
Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10120
- Pdf link: https://arxiv.org/pdf/2610.10120
- Abstract
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.
- 中文摘要
带可验证奖励的强化学习(RLVR)在任务结果可可靠评估领域取得了显著进展,但由于终端反馈稀疏和学分分配困难,长期交互仍具挑战性。过程奖励提供了更密集的监督,但与培训最相关的能力可能随着政策演进而变化:易于评估或常缺失的行为不必成为当前限制任务成功的瓶颈。我们介绍RewardWeaver,一个面向长期视野交互语言代理的自我演进奖励适应框架。RewardWeaver维护了一个经过验证的能力空间,在该空间中已承认的评分标准语义保持固定,并闭合了策略优化、任务评估、失败归因和奖励适应之间的循环。每个培训阶段结束后,它对低结果轨迹进行基于结果的逆向归因,汇总重复性和策略控制的能力瓶颈,并动态选择下一阶段对应的流程奖励。未被现有能力空间覆盖的重复性失败触发一个独立的受控扩展过程。我们在SOTOPIA、Amazon?HistoryPrice以及新构建的销售基准上评估了REWARDWEAVER。在社交互动、双边谈判和领域特定销售方面,REWARDWEAVER 建立了新的最先进(SOTA)结果。消融进一步展示了动态奖励分配、基于失败的归因和稳定语义对承认能力的重要性。
Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
超越结果奖励:构建与分配检索代理的检索信用
- Authors: Wenyu Huang, Xinyu Hou, Pavlos Vougiouklis, Ruofei Lai, Jeff Z. Pan
- Subjects: Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10179
- Pdf link: https://arxiv.org/pdf/2610.10179
- Abstract
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
- 中文摘要
搜索代理使大型语言模型(LLMs)能够迭代检索并利用复杂多跳问题的信息。带可验证奖励的强化学习(RLVR)为此类代理的后训练提供了有前景的方法,但其对稀疏、基于结果的监督可能使学分分配变得困难并限制学习效率。本文系统地探讨了中间监督如何改善搜索代理的强化学习。我们研究了多种奖励塑造和学分分配策略,这些策略从中间检索步骤中获得学习信号。基于这些见解,我们开发了一个将中间信号与最终结果奖励结合的训练框架,以提升多步搜索轨迹中的学习能力。在匹配训练条件下,多个基准测试的实验显示了总搜索代理表现的提升,并表明中间信号的选择及其学分分配位置都会影响训练行为。这些发现表明,奖励设计和信用分配是培训有效搜索代理的重要设计维度。
VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
VideoEvolve:共进化记忆与检索以实现长视频理解
- Authors: Yongchao Xu, Bowen Ye, Jiefeng Gan, Junkai Ma, Wenzhao Li, Sen Tao, Yi Wei, Jiawei Liu
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10183
- Pdf link: https://arxiv.org/pdf/2610.10183
- Abstract
Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.
- 中文摘要
长视频理解越来越依赖外部记忆将庞大的视觉流组织成紧凑的表示。然而,大多数基于内存的方法动态调整信息检索方式以适应不同问题,同时在很大程度上修正记忆内容。这种不匹配使得缺失细节的恢复成本高昂,而存储的信息只有在能够可靠检索时才有价值。为解决这一问题,我们提出了VideoEvolve,一种新型自我演化框架,结合记忆和检索,实现长视频理解。具体来说,从粗略的低帧率概述出发,VideoEvolve将用于选择性记忆增强的记忆演化器与用于自适应检索的记忆演化器结合起来。随后,我们通过交替代理强化学习(代理强化学习)共同进化两个进化器,更新一个,冻结另一个。为引导这一交替演进,瓶颈感知进化反馈(BEF)识别当前瓶颈是内存还是检索,并将优化引导至更受限的一方。此外,VideoEvolve引入了能力感知演化反馈(CEF),缓解下游反馈因记忆过度专精而转向固定训练问题,将训练转向尚未成熟但可学习的视频能力。通过将代理强化学习与BEF和CEF整合,VideoEvolve将下游推理经验转化为可转移的能力更新,提供从静态长视频系统向体验驱动、自我提升多模态智能的具体路径。在多个长视频理解基准测试上的广泛实验展示了VideoEvolve的有效性。
Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains
通过分层强化学习实现节能步态适应,适用于跨越多样地形的四足行走
- Authors: Ammar Issa, Anubhav Singh, Anton Tsaritsin, Sergey Kolyubin
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10297
- Pdf link: https://arxiv.org/pdf/2610.10297
- Abstract
While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.
- 中文摘要
虽然能效是腿部机器人运动控制的关键目标,但在不同速度范围和地形条件下保持稳健性能的同时实现低能耗仍是关键挑战。这在端到端强化学习策略中尤为重要,步态生成、动作执行和能量优化紧密耦合,导致对奖励设计高度敏感。本研究提出一个层级强化学习(HRL)框架,将高频策略与低频步态适应区分开来,以实现稳定且稳健的关节级运动执行,并明确降低运输成本(CoT)。基于Isaac的三阶段训练过程实现零射击模拟到真实传输,提升跟踪精度、稳健性和能量效率。学习的层级表现出自动的速度依赖步态适应,能从低速配速过渡到高速慢跑。我们在代表性单策略和层级运动基线的模拟中验证了该方法,展示了在广泛指令速度范围内降低CoT的效果,同时在平坦、崎岖和倾斜地形上保持稳健的运动能力。我们还通过在物理Unitree AlienGo四足机上零发射部署,进一步证明了其可行性。
Continual Graph Multi-Agent Reinforcement Learning
持续图多智能体强化学习
- Authors: Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi
- Subjects: Subjects:
Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.10302
- Pdf link: https://arxiv.org/pdf/2610.10302
- Abstract
In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.
- 中文摘要
在持续多智能体强化学习(CMARL)中,智能体在任务序列中学习协作策略,旨在有效适应新任务,同时保持解决先前遇到任务的能力。在许多应用中,任务的底层结构不同,例如不同的操作条件或目标配置(例如电网中的不同网络拓扑或编队控制中的布局)。现有的CMARL方法缺乏专门机制来利用这些结构信息来学习新任务,未能促进转移和减少遗忘。为填补这一空白,我们提出了连续图多智能体强化学习(CGMARL),这是一种针对CMARL问题的新框架,该框架将任务序列映射为一系列带属性的图,每个图建模任务特定结构。在CGMARL中,每个图决定环境动态(下一状态和/或奖励)及相应任务的代理数量。随后,我们介绍了基于图的形成(GRAFO),这是首个CGMARL基准测试,并展示了遗忘在该环境中的产生机制。最后,为解决这一限制,我们提出了Frozen Graph Encoder(FROG)方法,该方法依赖于冻结的图骨干来保存基于图的CMARL策略中的过去结构信息。GRAFO上的实验表明,将FROG与现有CL方法配对能显著提升多种CGMARL场景下的性能。
Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach
多链多层次多重计算平台的平均奖励强化学习:一种层级分解方法
- Authors: Huizhen Yu, Isaiah Heidt
- Subjects: Subjects:
Machine Learning (cs.LG); Optimization and Control (math.OC)
- Arxiv link: https://arxiv.org/abs/2610.10326
- Pdf link: https://arxiv.org/pdf/2610.10326
- Abstract
We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.
- 中文摘要
我们研究了在平均奖励多链马尔可夫决策过程(MDP)中学习最优策略,其中最优收益可能取决于初始状态,且递归结构在不同策略间存在差异,这给强化学习(RL)方法带来了挑战。我们提出了一种异步基于价值迭代的强化学习算法,除了MDP的转换图外,不要求任何模型知识,并利用Bather分解将状态空间分层划分为通信的子系统和瞬态状态。这种分解促使全局决策问题被重新定义为结构化子问题,我们的算法对此进行了利用。我们证明该算法在有限时间内收敛到最优增益并产生增益最优策略。基于该基础算法,我们开发了两个进一步算法:一个近似求解多链平均最优方程以获得近似收益最优策略,另一个通过近似最优偏置函数并用基础算法求解诱导平均奖励多链MDP,实现近偏向最优。我们为这三种算法提供了几乎确定的收敛保证,并通过实证比较它们的权衡,表明后两者相对于基础算法也持续提升瞬态性能。据我们所知,这些是首个基本无模型的通用多链MDP平均奖励强化学习算法,且不涉及折现问题的简化。
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
SOTA:股票期权交易代理,受期权隐含收益分布指导
- Authors: Yizhen Xie, Mengyang Liu
- Subjects: Subjects:
Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Portfolio Management (q-fin.PM); Trading and Market Microstructure (q-fin.TR)
- Arxiv link: https://arxiv.org/abs/2610.10407
- Pdf link: https://arxiv.org/pdf/2610.10407
- Abstract
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.
- 中文摘要
随着期权市场的增长和人工智能的进步,期权交易中的代理系统正受到越来越多的关注。基于语言模型的代理可以基于上下文信息(如新闻)进行推理,但期权交易带来了一个特别具有挑战性的决策难题:同一只股票可能拥有数千个合约,代理必须决定交易哪些合约以及如何组合它们。现有方法通常通过将政策限制在固定的策略结构(如跨式策略)中来规避这一复杂性,从而限制其在市场变化时切换策略的能力。我们介绍了SOTA(股票期权交易代理),这是一个用于结构化期权-策略选择的代理交易框架。SOTA将庞大的期权宇宙抽象为策略级决策,而确定性解析器则负责投资组合的实施。我们通过对Qwen3.8-27B进行后期训练,进行监督微调,随后进行强化学习,从而开发SOTA。SOTA在同一交易环境中,基于九只大型股美国股票和SPY的期权,与基于规则的策略选择器和机器学习策略选择器进行评估。在六个月的样本外期内,SOTA获得18.3%的总回报,夏普比率为1.60,最大回撤率为8.96%。我们还记录了新闻的非对称作用:新闻改善了前沿教师的轨迹,但在强化学习期间保留新闻则使样本外回报率从18.3%降至-2.7%。
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
是哪次推广教会了它这一点?BehaviorTrace 与在线强化学习中训练数据归因的局限性
- Authors: Amit Nautiyal
- Subjects: Subjects:
Machine Learning (cs.LG); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.10422
- Pdf link: https://arxiv.org/pdf/2610.10422
- Abstract
When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.
- 中文摘要
当强化学习教一个语言模型一种新行为时,我们能否找到教过它的训练推广?当归因方法说可以时,我们如何知道答案是真实的?我们用GRPO学习在线强化学习微调这两个问题,使用一个已知原因的植入行为。我们发布了BehaviorTrace,一个开放评估工具,结合了全梯度草图、种植行为设置以及梯度大小、流畅度、余量以及种子和生成抽取间的变异控制。在Qwen2.5-1.5B的三个种子中,大部分表观归因信号来自混杂因素。一个仅根据梯度大小排序训练步骤且无行为目标的对照,概率达到4.2到4.5倍,并且在三个种子中有两个上达到或超过了最佳目标估计者。在饱和检查点,模型的流畅度预测行为标签至少与我们比较的所有梯度方法相当。一旦流畅度被控制,每次部署的结果会在种子间、一代抽取之间变化,因此单次运行无法解决问题。一个信号对三个种子均成立。触发标记的梯度与行为实际发生的目标相符。我们将这些发现转化为评估强化学习归因的检查表。我们测试现有估计器,包括GAS(重整化TracInCP)和TRAK风格估计器,不提出新估计器。
CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
CoTrace:利用束-模型共进化训练终端代理的数据配方
- Authors: Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Zeyuan Chen, Lianhui Qin, Chien-Sheng Wu
- Subjects: Subjects:
Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.10426
- Pdf link: https://arxiv.org/pdf/2610.10426
- Abstract
Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.
- 中文摘要
终端代理能力依赖于模型权重和运行时的束缚器,该束缚器负责格式化提示、绑定工具并处理错误恢复。现有的束缚-模型共进化方法提升了这两个组件,但通常将束缚搜索过程中产生的轨迹视为未差别的重放缓冲区。这种做法忽视了轨迹在模型训练中的价值取决于其生成的束体。为系统分析该接口,我们建立了交替共进化框架,通过组件级推广决策将束缚器搜索与策略训练解耦。在该框架下,我们引入了CoTrace,一种具约束力束的数据配方,明确管理轨迹路由、来源匹配和课程刷新。在CoTrace下,重复执行失败指导工具束综合,而策略训练严格依赖于已验证的部署,且与所采用运行时间匹配的监督微调(SFT)或强化学习(RL)的新鲜在线交互。在Tmax推广分段,CoTrace将Qwen3.5-9B在监督微调下从78个已解决任务推进到88个,而在线强化变体则达到90个。具体来说,紧凑的束缚匹配语料库在远低于跨兄弟框架的大型语料库池下,实现了稳定的模型提升。此外,对Terminal-Bench 2.1和SWE-bench Lite的评估显示,分布外转移根本依赖于束带兼容性,在训练和评估运行时间间保持一致性可防止在外部支架下观察到的过程执行崩溃。
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
一个优秀的自学者会从学生所在的地方出发:联合政策学习与教学
- Authors: Randy Ardywibowo, Arnav Dalal, Jiantao Jiao
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
- Arxiv link: https://arxiv.org/abs/2610.10447
- Pdf link: https://arxiv.org/pdf/2610.10447
- Abstract
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.
- 中文摘要
来自结果奖励的强化学习(RL)存在稀疏监督,尤其是在困难且视野较长的任务中,成功轨迹罕见且生成成本高昂。政策提纯(OPD)通过提供更强教师对学生自身世代的密集代币级监督,提供了有吸引力的替代方案。自我提炼方法进一步消除了对独立教师模型的需求,通过将同一策略设定为特权信息作为独立教师。然而,仅有特权条件反射并不能保证最终的提炼更新能提升学生。事实上,特权信息可能导致教师通过学生无法获得的捷径来完成任务,导致监督与学生当前行为不匹配。因此,即使是表现更好的教师也可能提供降低学生表现的指导。为此,我们分析了特权教师的选择如何影响学生的更新。我们推导出一个必要且充分条件,使教师的局部蒸馏更新为学生奖励梯度的正数倍。我们的分析表明,教师不仅应在任务中表现出色,还应提供适合学生当前能力的指导。这一特征激励了一个实用的教师培训替代品,将结果奖励与代币级的Kullback-Leibler(KL)正则化结合起来。基于该结果,我们提出了联合政策学习与教学(JOLT),即将单一策略分成两个角色:特权教师使用KL正则化目标,非特权学生使用密集的策略提炼。在数学推理、编码、工具使用和终端使用方面,JOLT提升了培训效率和表现,并进一步提升学生的奖励。
MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration
MORCA:视频扩散加速中自适应缓存重用的离线到在线强化学习
- Authors: Yuxiang Xiong, Ruiyan Wang, Wenqiang Wang, Teng Hu, Songhang Shen, Bohao Feng, Hongqian Deng, Ran Yi
- Subjects: Subjects:
Computer Vision and Pattern Recognition (cs.CV)
- Arxiv link: https://arxiv.org/abs/2610.10457
- Pdf link: https://arxiv.org/pdf/2610.10457
- Abstract
Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at this https URL.
- 中文摘要
扩散变换器(DiT)在视频合成中表现出色,但其迭代去噪过程存在较高的推理延迟。为此,缓存已成为一种有效的加速策略,利用去噪期间的步间冗余。现有的动态缓存方法通常估计缓存重用在每个去噪步骤中引入的误差(步误误差),以指导缓存决策,而我们关注的是缓存重用对最终生成视频造成的质量损失(终端错误)。我们表明步频误误不直接对应终端错误,潜在信息有助于捕捉它们的关系,从而指导缓存决策。此外,现有基于阈值的方法无法提供精确的加速控制,难以满足用户指定的加速目标的实际需求。为解决这些限制,我们引入了MORCA,一种通过离线到在线强化学习训练的缓存调度框架,用于在用户指定的加速目标下做出潜在感知的重用/重算决策。在不同视频生成模型中,跨多个目标加速比率的大量实验表明,MORCA在相当的计算预算下,比最先进的缓存方法实现了更好的生成保真度。代码可在此 https 网址获取。
HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion
HuMBLE:具身运动的人类运动驱动行为学习
- Authors: Mike Zhang, Dongho Kang, Kevin Bergamin, Nicola Burger, Robin Deits, Jonathan Foster, Bilal Hammoud, Katie Hughes, Francesco Iacobelli, Twan Koolen, M. Eva Mungai, Zach Nobles, Shane Rozen-Levy, Jean Pierre Sleiman, Fangzhou Yu, Yunbo Zhang, Alfred Rizzi, Jessica Hodgins, Scott Kuindersma, Yeuhi Abe, Sylvain Bertrand, Farbod Farshidian
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.10489
- Pdf link: https://arxiv.org/pdf/2610.10489
- Abstract
Despite recent advances in humanoid locomotion, controllers optimized for command tracking and robustness tend to produce mechanical gaits, whereas controllers tied to human motion data often fail to generalize to commands outside the data distribution. This work introduces a learning framework that balances these competing objectives to synthesize real-time steerable, robust, and biomimetic locomotion policies from human data. Using an in-house curated locomotion dataset covering diverse speeds and directions, we first learn a natural locomotion prior policy through a teacher-student distillation process. Specifically, we train a full-body reference-conditioned policy with Reinforcement Learning (RL), then distill it into a lightweight prior policy conditioned solely on proprioception and a planar torso-velocity steering command. Next, we fine-tune the prior policy with multi-task RL to expand command coverage and robustness beyond the data distribution, pairing a goal-conditioned task that tracks arbitrary commands with a reference-guided task that tracks the human data as an explicit style regularizer. We validate our framework on three humanoid robots: the Boston Dynamics Atlas R1, Atlas D1, and Unitree G1. Experimental results demonstrate robust performance across real-world scenarios, including direct user-controlled locomotion in indoor and outdoor environments, and integration as the locomotion layer within hierarchical control stacks. Benchmarks against Tabula Rasa RL policies trained without human data and ablation studies confirm that our framework yields a lightweight, deployable policy that reconstructs coordinated whole-body behavior from a steering command, retaining the human gait characteristics while remaining robust and fully steerable.
- 中文摘要
尽管类人移动技术最近有所进展,针对指令跟踪和稳健性优化的控制器往往会产生机械步态,而依赖人体运动数据的控制器往往无法推广到数据分布之外的指令。本研究提出了一个学习框架,平衡这些相互竞争的目标,从人类数据中综合实时可操控、稳健和仿生的运动策略。利用内部策划的移动数据集,涵盖多种速度和方向,我们首先通过师生提炼过程学习自然的先行策略。具体来说,我们通过强化学习(RL)训练一个全身参考条件策略,然后将其提炼为仅基于本体感觉和平面躯干速度引导指令的轻量级先行策略。接下来,我们用多任务强化学习微调先前策略,扩展命令覆盖范围和鲁棒性,超越数据分布,将跟踪任意命令的目标条件任务与跟踪人类数据的引用引导任务作为显式风格规范器结合。我们在三款类人机器人上验证了该框架:波士顿动力Atlas R1、Atlas D1和Unitree G1。实验结果显示,在实际场景中表现出稳健性能,包括室内外环境下用户直接控制的移动,以及作为层级控制栈中运动层的整合。基于Tabula Rasa RL策略的基准测试,在无人类数据和消融研究中训练,确认我们的框架提供了轻量化、可部署的策略,能够从转向指令重建协调的全身行为,既保留人类步态特征,又保持稳健且完全可操控。
Decoupling Exploration from Optimization in RLVR
RLVR中探索与优化的解耦
- Authors: Saif Punjwani, Micah Goldblum
- Subjects: Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
- Arxiv link: https://arxiv.org/abs/2610.10536
- Pdf link: https://arxiv.org/pdf/2610.10536
- Abstract
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
- 中文摘要
现代语言模型在已训练的检查点基础上进行可验证奖励(RLVR)强化学习。RLVR的一个关键承诺是发现新的推理策略。原则上,模型可以采样其先前训练数据中缺失的新想法。然而,实际上,用强新奇激励增强RLVR的效果有限,且可能降低模型质量。由于可验证奖励只监督模型知识和行为的一小部分,这种退化难以恢复。相反,我们将探索与优化分离,采用称为探索-蒸馏(ExpDis)的框架。我们训练一个或多个探索者策略,奖励中带有新颖性加成,过滤其轨迹以确定正确性和质量,并提炼成独立的学生策略。学生策略随后在不加新奇加成的情况下进行训练。我们重复上述程序进行多轮,交替进行探索和优化。这种解耦使我们能够积极扩展探索,同时不降低学生政策。在七个数学推理基准测试和两个模型家族中,ExpDis在相同的墙时钟预算下表现优于DAPO。此外,我们观察到pass@$k美元扩展有所改善,表明ExpDis产生的模型能生成更多样化的正确解。
Keyword: diffusion policy
Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment
不可混淆扩散策略:通过无标签噪声分配保持多模态机器人动作
- Authors: Xiao Zhang, Yuxin Chen, Zhixuan Liang, Guojian Zhan, Chenran Li, Chenfeng Xu, Masayoshi Tomizuka, Yiheng Li
- Subjects: Subjects:
Robotics (cs.RO); Machine Learning (cs.LG)
- Arxiv link: https://arxiv.org/abs/2610.09369
- Pdf link: https://arxiv.org/pdf/2610.09369
- Abstract
When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.
- 中文摘要
扩散策略刚引入时,预期能恢复多模态动作分布。然而,我们发现这一期望并不总是成立,因为扩散策略往往会崩溃为单一模态,即使我们保证了数据集模态的平衡和批内精确对称性。我们的分析表明,独立的动作-噪声配对导致了这一失败,因为它增加了扩散路径之间的混合和交叉,从而产生平均的去噪响应,抑制模态特定行为。这个问题在机器人规划中尤为严重,因为动作空间密集且低维,显著增加了混合和交叉。为缓解此问题,我们提出了不可混淆扩散策略,这是一种无标签的训练时间扩展,利用动作-噪声分配保持相对不同的噪声到动作路径,而无需修改策略架构或推理过程。在五个模拟任务和两个真实世界的人形操作任务中,涵盖状态、RGB和点云观测,我们的方法显著提升了策略对动作模态的保留,同时保持了强劲的任务性能。它将三个双模态任务中非主导模态比例增加了6.0倍至14.6倍,并恢复了两种四模态任务中完全缺失的演示模态。这些结果表明,不可混淆扩散策略为一般机器人学习任务中保持动作多模态提供了简单而稳健的方法。
RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input
RoboPrompt:直观的机器人政策引导,人力稀疏
- Authors: Yanwen Zou, Chenyang Shi, Guoxuan Xu, Wenye Yu, Wendi Chen, Ye Pan, Cewu Lu, Chuan Wen
- Subjects: Subjects:
Robotics (cs.RO)
- Arxiv link: https://arxiv.org/abs/2610.10534
- Pdf link: https://arxiv.org/pdf/2610.10534
- Abstract
End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, $\pi_{0.5}$, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5\% for $\pi_{0.5}$ across three tasks and by 21.3\% across three policies(Diffusion Policy, $\pi_{0.5}$, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0\% (2.86 to 1.60) and 81.9\% (2.60 to 0.47), respectively.
- 中文摘要
通过模仿学习训练的端到端机器人策略仍受限于有限的数据多样性,使得在现实环境中可靠的零机会部署变得困难。共享自治方法通过远程操作实现人工纠正,但专业硬件和操作员培训阻碍了大规模部署。其他方法则将人工指导作为额外策略输入,通常需要架构调整和专门的引导训练,限制了其跨策略的适用性。我们介绍RoboPrompt,一种通用的轻量级机器人策略引导系统,使用户通过直观、稀疏的输入(包括绘制的轨迹、目标点和粗略方向指令)来引导策略行为。RoboPrompt将人类意图翻译与底层策略解耦:一个可复用模块将人类指导转换为动作草稿,并通过基础策略的扩散或流量匹配动态进行细化。通过控制噪声空间中的动作生成,RoboPrompt在不修改基础策略架构或微调以实现可引导性的情况下,平衡了人类意图与策略。实验证明了扩散策略、$\pi_{0.5}$和FastWAM之间的有效引导。我们还通过DAgger使用引导式推导进行在线策略改进。经过2-3轮迭代后,三个任务的$\pi_{0.5}$平均成功率提升15.5%,在插入面包任务中,三个策略(扩散政策、$\pi_{0.5}$、FastWAM)平均成功率提升21.3%,而平均人类干预次数分别下降44.0%(2.86降至1.60)和81.9%(2.60降至0.47)。