生成时间: 2026-10-08 23:20:12 (UTC+8); Arxiv 发布时间: 2026-10-08 20:00 EDT (2026-10-09 08:00 UTC+8)

今天共有 55 篇相关文章

Keyword: reinforcement learning

Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation

自适应工作流程智能:基于上下文的企业自动化认知架构

HydroSphere: A Framework for Governed, Self-Healing Wastewater Infrastructure

HydroSphere:一个治理自愈污水基础设施框架

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

JoyAI-Voice 2.0:一个全连续自回归语音生成模型,具备语义-声学联合表示

Autonomous Droplet Navigation via Model-Based Reinforcement Learning: Zero-Shot Transfer and Emergent Dynamics

基于模型的强化学习实现自主液滴导航:零射点转移与涌现动力学

AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails

AdaGuard:通过支持推理的LLM作为评判者保护栏,提升安全和政策合规性

On KL-Regularized Policy Optimization

关于KL正则化策略优化

HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids

浩克:学习人形生物全身强制操控

LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL

LASER:支持约束熵正则化离线强化学习的潜在空间伴随匹配

TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery

TAP:通过轨迹锚定恢复实现高效的长视野代理剪枝

Convex-Concave Reinforcement Learning

凸凹强化学习

The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning

欺骗性强盗问题:探索性耦合与多智能体学习的脆弱性

Mission-critical spectrum sharing with decentralized Multi-Agent Reinforcement Learning

通过去中心化多智能体强化学习实现关键任务频谱共享

AGAR: a reinforcement learning substrate for LLM program evolution

AGAR:用于LLM程序演进的强化学习基底

RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms

RLDISCOVER:基于LLM驱动的强化学习算法共进化

FLoRa: Flight-Assisted Data Collection from Duty Cycling LoRa Nodes under Energy Constraints

FLoRa:在能量约束下,从值班循环LoRa节点的飞行辅助数据收集

An Informational Curse of Horizon in Goal-Conditioned Policy Learning

目标条件政策学习中地平线的信息诅咒

Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks

动态OWC网络中的合作对抗DQN SAC学习以实现能源效率

Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

专家问答:基于LLM的自主网络防御强化学习

Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization

通过最优性和反事实正则化学习未知约束而不使用不安全数据

Learning Stability of Replay-Based Co-Optimization for Transmission Expansion under Strategic Bidding

基于重放的协同优化在战略竞价下传输扩展的学习稳定性

TutorLoop: Regulating Student Learning Behaviors via Sensor-in-the-Loop Generative Feedback

TutorLoop:通过传感器在环生成反馈调节学生学习行为

TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding

TiTok:多段时间接地视听大型语言模型

Precise SE(3) End-Effector Tracking in Whole-Body Humanoid Control

全身类人生物控制中的精确SE(3)末端效应器跟踪

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

MIMESIS:学习用户模拟器作为交互代理的训练环境

Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?

平均来说安全,尾巴不安全:什么时候可以控制分段消耗的尾巴?

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

它没有意识到危险:一个冻结的视觉语言安全评分衡量了它的字幕库

EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs

EvoSignal:基于LLM的模块化交通信号控制程序的进化设计

SAPD: Step-Aligned Privileged Distillation

SAPD:阶梯对齐特权蒸馏

CERO: Where and When to Allocate Rollouts for RL Post-Training

CERO:在哪里以及何时分配强化学习后的推广时间

Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

超越政策支持:互动受限的离线强化学习对自动驾驶

BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation

BoT-GRPO:通过代币袋聚合进行推理的高效过程-奖励强化学习

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge:通过动态技能生命周期共同进化技能与特工

Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

自我演化,参考文献:工具集成代理的锚定训练

Training Advisors for LLM Agents from Task Outcomes

任务成果中的LLM代理培训顾问

Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks

基于TSN的车辆边缘网络的截止时间感知多智能体强化学习

A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

关于人机协作中基于人类反馈强化学习的范围范围综述与实验研究

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

RollVerify:桥接长尾推广强化学习的效率与准确性

Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

成功方式多样:基于多样性驱动的强化学习微调以实现VLA泛化

Decoding Neural Population Dynamics through Robotic Analog

通过机器人模拟解码神经群体动力学

Learning to Accumulate Knowledge with Mutual Information

学习通过互信息积累知识

RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

RewardWeaver:通过自我演化的奖励适应,为语言代理提供长期交互式学习

Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

超越结果奖励:构建与分配检索代理的检索信用

VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

VideoEvolve:共进化记忆与检索以实现长视频理解

Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains

通过分层强化学习实现节能步态适应,适用于跨越多样地形的四足行走

Continual Graph Multi-Agent Reinforcement Learning

持续图多智能体强化学习

Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

多链多层次多重计算平台的平均奖励强化学习:一种层级分解方法

SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions

SOTA:股票期权交易代理,受期权隐含收益分布指导

Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

是哪次推广教会了它这一点?BehaviorTrace 与在线强化学习中训练数据归因的局限性

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

CoTrace:利用束-模型共进化训练终端代理的数据配方

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

一个优秀的自学者会从学生所在的地方出发:联合政策学习与教学

MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration

MORCA:视频扩散加速中自适应缓存重用的离线到在线强化学习

HuMBLE: Human Motion-Driven Behavior Learning for Embodied Locomotion

HuMBLE:具身运动的人类运动驱动行为学习

Decoupling Exploration from Optimization in RLVR

RLVR中探索与优化的解耦

Keyword: diffusion policy

Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment

不可混淆扩散策略:通过无标签噪声分配保持多模态机器人动作

RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input

RoboPrompt:直观的机器人政策引导,人力稀疏