生成时间: 2026-09-18 20:49:03 (UTC+8); Arxiv 发布时间: 2026-09-18 20:00 EDT (2026-09-19 08:00 UTC+8)

今天共有 43 篇相关文章

Keyword: reinforcement learning

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

CovR:通过推理引导强化学习实现覆盖感知硬件验证

AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

AUDITPLAN:承诺,然后负责可审计的安全对齐

Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation

通过选择性奖励刺激改善离线目标条件强化学习

Winning a Won Game: Strict Reach-Avoid-Stay Control Barrier Functions for High-Dimensional Black-Box Systems

赢得胜利:高维黑匣子系统的严格伸手-避开-停留控制屏障函数

GLAMDRING: Gait Learning And Morphology co-Design via Reinforcement LearnING of CPGs

GLAMDRING:步态学习与形态学通过强化学习共同设计消费品

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

强化学习后语言模型中的组合推理

Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks

低空无线网络中异构无人机系统的代理人工智能网络

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

通过离线隐藏状态蒸馏恢复被激进修剪的视觉-语言-行动模型

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

伸手还是解决?通过检查点交接归因于能动强化学习的收益

EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence

EmbodiedMind:自适应数据管理与前缀树强化学习,实现高效的具身智能

Towards High-DoF Dexterous Manipulation through VLA Post-Training

通过VLA后训练实现高景深灵巧操作

UniExo: Unified Multi-Skill Policies for Musculoskeletal Locomotion and Co-Adaptive Exoskeleton Control

UniExo:肌肉骨骼运动与共适应外骨骼控制的统一多技能政策

Region-Level Policy Optimization for Fine-grained MLLM Perception

区域级政策优化以实现细粒度MLLM感知

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

DeliveryGym:一个具备长期具身代理规划的强化学习环境,结合自适应课程

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

LLM代理的双轴策略优化:贝叶斯反馈归因与轨迹质量归一化

GR2PO: Group Relative Return Policy Optimization for Continuous Robot Control

GR2PO:连续机器人控制的组相对返回策略优化

Learning Reliable Parking Policies via Offline Reinforcement Learning with Quantized Action Representations

通过带有量化行动表征的离线强化学习学习可靠的停车政策

VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes

VERA:Kubernetes 中 HPC 工作负载动态内存扩展的强化学习

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG树:用于梯度高效强化学习的计算最优分支

MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution

MAGMA-GEN:通过反事实重执行验证的模糊失败恢复监督

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

匹配:模型感知工具学习,结合课程安排和分层奖励

Safety-Critical Scenanrio Emerges from Initial Scene

安全关键的场景从初始场景中浮现

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

AnyViewDex:RGB观测中的视野不变灵活操作

Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

直播语音合成的多维韵律判断

Robust Federated Q-Learning with Almost No Communication

几乎无交流的稳健联邦Q-学习

VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots

VLN即时运行:用于空中机器人的机载视觉语言导航协议栈

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

通过双向行为事先蒸馏提升在线强化学习

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

通过边界感知数据增强提升离线强化学习中的泛化性和鲁棒性

DDQN-MLP: An Explainable and Adversarially Robust DRL-Guided Adaptive Learning Framework for Ransomware Detection

DDQN-MLP:一个可解释且具对抗性强的DRL引导自适应学习框架,用于勒索软件检测

Learning Principal-Agent Contracts for Equitable Smallholder Carbon Farming under Moral Hazard and Adverse Selection

学习道德风险和逆向选择下公平小农碳农业的主要代理人合同

The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes

在恒定步长下非线性两时间尺度随机近似的偏置

Visual Sim-to-Real Learning for Robotic Insertion under Geometric Variations: Application to Rebar Installation

在几何变分下机器人插入的可视化模拟到真实学习:在钢筋安装中的应用

Mitigating Retaliatory Algorithmic Collusion in Repeated Games

缓解重复游戏中报复性算法串通的问题

Learning Slope-Adaptive Whole-Body Locomotion for Humanoid Robots in Roofing Construction

学习屋顶施工中人形机器人的坡度自适应全身运动

Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning

推理质量重要:应对基于LLM的嵌入式学习推理崩溃

UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

UniPolicy:生成式搜索广告的统一目标特定政策

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

不要掩盖环境:观察监督改变了强化学习下代理的探索方式

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

MILER:非结构化自动驾驶中模拟到现实强化学习的语义中级表示

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

被选:使用无渲染教师进行政策内端到端驾驶微调

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

退休OPD:能动强化学习的自我退役策略提炼

Score Centering Stabilizes Off-policy Reinforcement Learning

评分中心稳定了非策略强化学习

Keyword: diffusion policy

ULOHA: An Underwater Bimanual Robot System for Robot Learning

ULOHA:一种用于机器人学习的水下双手动机器人系统

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

学习3D扩散策略中没有显式轨迹的前瞻性