生成时间: 2026-09-17 21:08:34 (UTC+8); Arxiv 发布时间: 2026-09-17 20:00 EDT (2026-09-18 08:00 UTC+8)

今天共有 38 篇相关文章

Keyword: reinforcement learning

Composite-Gradient Learning for Shared Control Authority Between Deep Reinforcement Learning and Model Predictive Control

深度强化学习与模型预测控制共享控制权威的复合梯度学习

REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

反转工作台:用于测量无重置强化逻辑悬崖的可逆轴和重置神谕

Learning Market Competition in Shared Spectrum: A Multi-Agent Reinforcement Learning Approach

共享频谱中的学习市场竞争:多智能体强化学习方法

CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors

CALOS:Control-仿射李雅普诺夫流管安全层,用于四旋翼的深度强化安全学习

AgenTeeth: A Model-Agnostic Framework for Suppressing Hallucination in Frozen Vision-Language Models on Dental X-Rays via Tool Evidence Injection

AgenTeeth:一种模型无关框架,通过工具证据注入,在牙科X光片上抑制冻结视觉语言模型中的幻觉

The Free Inference Dimension: Complexity Measure for Zero-Collision Navigation under Hypothesis Mixtures

自由推断维度:假设混合下零碰撞导航的复杂度量

Adaptive hybrid coupling with operator inference, the overlapping Schwarz alternating method and reinforcement learning

自适应混合耦合与算子推断、重叠施瓦茨交替法与强化学习

SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

工具调用代理的SFT还是RL?一项跨数据、方法和规模的受控研究

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

锚定重要因素:一个基于视觉基础的多模态推理的双层次学习框架

Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration

掩码2D-3D:用于图像到点云注册的自适应双掩盖自编码网络

Fetch My Beer: Synthetic-to-real Hierarchical Policy for Smooth Pick-and-place

“取我的啤酒”:合成到真实的层级政策,实现顺畅的挑选与放置

DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning

DualSQL:多智能体强化学习的文本转SQL。

Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning

模型型强化学习中动态转变下的重放保留特性

Reinforcement Learning for Real-Time Vision-Language-Action Policies

实时视觉-语言-行动政策的强化学习

APGEM: Adaptive Policy-Guided Error Mitigation for Quantum Reinforcement Learning on a Real-World CVRP Case Study

APGEM:基于真实世界CVRP案例研究的自适应策略引导错误缓解量子强化学习

MiST: Mid-Training LLMs for Cybersecurity

MiST:网络安全中期培训LLMs

GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

GroundingVLN:基于接地的推理与行动,进行视觉语言导航

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

无标签引导:将测试时间强化学习压缩到仅有偏见的子空间

CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning

CoRe-MARL:利用循环多智能体强化学习在未知动力学下的合作再分布

STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution

拓展边界:一个统一的自学框架,促进渐进式大型语言模型演进

FIERCE: From Generalist Robot Policies to Fast Specialists via Progress-Failure Feedback

Fierce:从通才机器人政策到通过进步-失败反馈快速专家

Learning to Program Adaptive Non-Local Observables for Machine Learning

学习编程自适应非局域可观测量以用于机器学习

M$^3$P-R1: Reinforcement Learning for Large Language Model Guided Multi-Modal Motion Planning via MIP Code Generation

M$^3$P-R1:通过MIP码生成实现大型语言模型引导多模态运动规划的强化学习

Voice of Reason: Reinforcement Learning for Spoken Math

理性之声:口语数学的强化学习

WeaveRL: Weaving Reconstruction into Scene-Aware Fabrics for Perceptive Reinforcement Learning

WeaveRL:将重建编织进场景感知织物中,以实现感知强化学习

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

重新思考PPO中的批评者学习:理解并缓解价值平坦化

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

CERA-MoA:与持续学习的大型语言模型代理共同演进的路由机制

KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

KINO:用于人形机车操控中VLM规划和全身控制的关键帧接口

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

使用双足移动机械臂学习整体全身运动操作

FedGuide: Diffusion Prior Alignment and Value Baseline Guidance for Heterogeneous Federated Reinforcement Learning

FedGuide:异构联合强化学习的扩散先验对齐与价值基线指导

Loco-Loco-RL: Low-Cost Terrain Mapping for Humanoid Locomotion with Reinforcement Learning

Loco-Loco-RL:低成本的类人机动地形映射与强化学习

Integrated Optimization of Automated Warehouse Operations and Last-Mile Transport for Differentiated On-Demand Delivery

自动化仓库运营和最后一公里运输的综合优化,实现差异化的按需配送

RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

RLLBC-Lib:一个用于强化学习和基于学习控制的教育代码库

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

ScienceIDE:将全球科学代码库转变为可智能体学习环境

Keyword: diffusion policy

Vision-Language Grounded Task-Context-Aware Imitation Learning for Robotic Disassembly

视觉语言基础任务上下文感知模仿学习用于机器人拆解

Missing Bridges: Composition-Aware Active Imitation Learning

缺失的桥梁:构图感知主动模仿学习

A Comprehensive Review of Generative Physical Artificial Intelligence

生成式物理人工智能的全面综述

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

使用双足移动机械臂学习整体全身运动操作