生成时间: 2026-07-23 18:25:11 (UTC+8); Arxiv 发布时间: 2026-07-23 20:00 EDT (2026-07-24 08:00 UTC+8)

今天共有 25 篇相关文章

Keyword: reinforcement learning

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation

从轨迹到前缀:通过重放前缀和在线续写重用教师轨迹

HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions

HyGRL:多实体问题的自适应混合图推理

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

利用离线监督实现大规模视觉-语言-行动模型中高效且可推广的强化学习

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

REGEN:基于离线强化学习的专家到通才提炼的重放循环

CHMAS: A Coupled Hierarchical Framework for Multi-Agent Reinforcement Learning

CHMAS:多智能体强化学习的耦合层级框架

CRB-Driven Beamforming and Trajectory Optimization for UAV-assisted ISAC System

CRB驱动的波束成形与无人机辅助ISAC系统的轨迹优化

The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning

机制很重要:知识图谱何时有助于强化学习

HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems

HypEMBER:基于超网络的集合,用于参数化动力系统的稳健策略学习

A Human-AI Teaming Framework for Deep Reinforcement Learning-Based Voltage Regulation in Distribution Networks

基于深度增强学习的配电网络电压调节人机团队框架

SLPO: Scaling Latent Reasoning via a Surrogate Policy

SLPO:通过替代政策扩展潜在推理

The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL

《世界模特记,演员遗忘:持续模型强化的梦境彩排》

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Trace:一个多领域视觉推理的分类引导环境

Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

Dreamer-CPC:基于世界模型进行的信息学习,用于去中心化多智能体强化学习

Rewarding Better Thinking for LLM Preference Alignment

奖励更好的思维以实现LLM偏好对齐

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness

EA-Nav:学习具身意识的安全视觉导航政策

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

MOF-Sleuth:工具基础奖励对齐,实现可解释的细粒度MOF审计

Towards Reliable C-to-Rust Translation with Rule-Guided Reasoning and Reinforcement Learning

迈向可靠的C到Rust翻译,结合规则引导推理和强化学习

Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing

记忆协调:动态制造中多智能体适应中的图结构体验重用

Generalized Kalman filter based temporal difference reinforcement learning

基于广义卡尔曼滤波器的时间差分强化学习

Active Inference as a Convex Markov Decision Process

主动推断作为凸马尔可夫决策过程

OLEDLM: A Unified Language Model for OLED Molecular Design

OLEDLM:OLED分子设计的统一语言模型

Notes to Self: Can LLMs Benefit from Experiential Abstractions?

给自己的笔记:大型语言模型能从体验式抽象中受益吗?

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

PercepCap:具备结构化时空感知的视频字幕

Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning

迈向利用虚拟现实和强化学习的微型人形远程机车操控

Keyword: diffusion policy

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

扩散重叠:可修正去噪用于机器人顺序预测