生成时间: 2026-09-23 21:26:40 (UTC+8); Arxiv 发布时间: 2026-09-23 20:00 EDT (2026-09-24 08:00 UTC+8)

今天共有 38 篇相关文章

Keyword: reinforcement learning

Towards Adaptive Interaction Strategies for Human Companion Robot via Deep Reinforcement Learning

通过深度强化学习迈向人类伴侣机器人的自适应交互策略

RULER: Instance-aware Rubric Rewards for SVG Generation

RULER:SVG生成的实例感知评分奖励

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

人机协作中的规划学习:多模态强化学习用于自适应交互

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

TelecomGPT-R1:跨异构电信任务推理的统一后期训练

HOTICE: Whole-Body Humanoid Object Transportation in Cluttered Environments

HOTICE:在杂乱环境中实现全身类人生物物品运输

Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing

基于一维箱装箱的物品兼容性图的深度强化学习

Norm2Tex: Augmenting Visuo-Tactile Simulations with Texture

Norm2Tex:用纹理增强Visuo触觉模拟

WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning

WeightBridge:高效的强化学习权重转移库

Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

强化学习推理大型语言模型推广效率:分类法与未来方向

Learning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy

学习针对多样化推理攻击的智能电表隐私防御策略

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

套娃归因:学习将语言模型输出归因于表示和权重

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

基于零基LoRA的后强化学习LLM的保持推理的微调

PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning

PAKT:强化学习的身体对齐动觉教学

DynaForge: Planning-Guided Residual Learning for Dynamic Manipulation Demonstration Generation

DynaForge:动态操作演示生成的规划引导残差学习

Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform

向高中生教授强化学习与类人机器人:一项专家验证的低成本开放平台上的课程设计

Fully Byzantine-Resilient Multi-Agent Reinforcement Learning

完全拜占庭弹性多智能体强化学习

PLAT: Sparse Timed Keyframe Motion Tracking for Humanoid Control via Privileged Latent Transition Learning

PLAT:通过特权潜在过渡学习实现稀疏定时关键帧运动追踪,用于人形控制

Model-Free Current Control of Permanent Magnet Synchronous Motors via ESO-Based Disturbance Feedforward and Data-Driven H-infinity Residual Feedback

通过基于ESO的扰动前馈和数据驱动的H-Infinity残差反馈,实现永磁同步电机的模型无电流控制

Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models

视频跳链:多跳问题与视频推理模型的信心门槛探索

Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving

有时候你得先跑,才能走路:VLM自动驾驶的先跑后走的排班策略

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

知情掩蔽:扩散大型语言模型中强化学习中的结构感知扰动

MATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies

MATES:通过变换观察值以实现冻结单智能体策略学习多智能体交互

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

多层网络可行几何结构上的可微策略传输

Test-time Reinforcement Learning for Anomalous Video Understanding

异常视频理解的测试时强化学习

A Cross-Dataset based Zero-Day Intrusion Detection System by Integrating Siamese Network and Reinforcement Learning

通过整合暹罗网络和强化学习实现的跨数据集零日入侵检测系统

From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems

从风险评分到风险分配:多智能体系统中多样化监控的密度驱动框架

MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning

MGRL-RSCC:多粒度奖励强化学习,用于细粒度遥感变更字幕

ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

ForeDrive:前瞻性引导的端到端自动驾驶,结合规划相关潜在世界模型

PACT: From Credit Assignment to Critic Alignment

PACT:从署名分配到批评对齐

KwaiMind Technical Report

KwaiMind技术报告

RouteRLT: Learning When and Which RL Specialist Should Control a Vision-Language-Action Policy

RouteRLT:了解何时以及由哪位强化学习专家来控制愿景-语言-行动策略

Learning Air-Ground Motion Control with Temporal Mode Switching and Cross-Terrain Tracking

学习利用时间模式切换和跨地形跟踪的空地运动控制

GeoComposer: Geometry-Grounded Photographic Composition Instruction

GeoComposer:几何基础摄影构图教学

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

MAGIC:通过增量构建结合密集奖励强化学习实现混合粒度代理图

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

超越重复抽样:学习大型语言模型推理的搜索策略

Keyword: diffusion policy

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

JAMB:用于双手操作的联合动作-运动扩散

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

RoboMP-DINOv2:用于强健机器人操作的提示,而非过滤器

From Instrument-Mounted Demonstrations to In-Vivo Execution: Learning Bimanual Laparoscopic Appendectomy Without Robot-Collected Demonstrations

从器械演示到体内执行:在不依赖机器人收集演示的情况下学习双手腹腔镜阑尾切除术