生成时间: 2026-10-09 23:05:19 (UTC+8); Arxiv 发布时间: 2026-10-09 20:00 EDT (2026-10-10 08:00 UTC+8)

今天共有 66 篇相关文章

Keyword: reinforcement learning

Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

教机器人狗新把戏:通过强化与模仿学习结合对抗性任务选择,实现多样化四足动物技能

Sample-Efficiency of Kolmogorov-Arnold Networks

Kolmogorov-Arnold 网络的样本效率

Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction

医疗代币用于诊断预测的覆盖意识推理

Safe Learning of Adaptive Control Policies for Remote Patient Monitoring

远程患者监测自适应控制政策的安全学习

Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training

对话任务对表格数据的消歧:泄漏感知表述、基准套件与训练

Adaptive E2E Monitoring Path Selection with Deep Reinforcement Learning in Multi-domain Optical Networks

多域光网络中的自适应端对端E监测路径选择与深度强化学习

Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach

能源转型中价值创造的战略投资决策:强化学习方法

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

VICO:视觉环境为视觉语言模型推理而共同演进

NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

NavGPT-3:在分层导航运行时中利用上下文

On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

上班时间:迈向准时且高效的时间预算AI代理

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

SPLIT-RL:分阶段感知-语言推理训练,具索赔级优势

Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory

基于强化学习和博弈论的自适应多判别器WGAN框架,用于资源受限车辆互联网

World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning

目标条件强化学习的世界模型政策仲裁器

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

StoreBench:一个用于评估和培训自主运营代理的实时商务环境

Learning Multi-Step Query Rewriting via Corpus Feedback for Conversational Search

学习通过语料库反馈进行多步骤查询重写以实现会话搜索

ATLAS: Adaptive TDA-guided Landscape-Aware Transistor Sizing

ATLAS:自适应TDA导向景观感知晶体管尺寸

Measuring and Mitigating Solution Mode Collapse in RLVR

测量与缓解RLVR中解模坍缩

DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

DaCe-DT:通过自适应提示和轨迹修正的数据中心离线多任务强化学习,针对异构任务

FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment

重点:从特权状态到带有受控模态切换和表示对齐的RGB-D

When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction

当接口发声:数据感知生成式UI驱动主动交互

Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning

LLMs是否能从情境中的奖励中学习?:重新思考情境内奖励在情境强化学习中的作用

PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

PIVOT:基于困惑度的KD到RL垂直域少数精馏过渡调度

Higher-Order Action Supervision Makes A Strong Policy Class

高阶行动监督构成强有力的政策类别

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

RAG压力:探讨检索增强生成中证据依赖的极限

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

为什么策略提炼有时失败:学习信号的消失

How to post-train on a surrogate: Envelope sampling mitigates reward hacking

如何对代理进行后期训练:信封采样有助于缓解奖励黑客行为

SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning

SynCo:通过多智能体强化学习实现自演化大型语言模型的数据综合协同训练

RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty

RL-ARC:通过推理引导不确定性校准大型推理模型

From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning

从提示到词汇:不断演变的功能性词汇助力LLM持续学习

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

环境反馈建模的重要性:重新思考能动性回顾自我蒸馏中的反馈处理

DRL-Based AoI Minimization for RSMA in Finite-Blocklength MU-MISO

基于DRL的有限区块长度MU-MISO中RSMA的AOI最小化

SAIL: Scientific Agentic Intelligence via a Science-Aware Loop

SAIL:通过科学感知循环实现科学代理智能

Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization

Fed-GRPO:奖励信号驱动联邦集团相对政策优化

DAMP: Humanoid Locomotion via Denoised Belief Learning and Adversarial Motion Priors

DAMP:通过去噪信念学习和对抗运动先验实现类人运动

Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards

残余优势:学生-亲属教师强化学习指导,并提供可验证的奖励

When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning

何时干预?联邦强化学习中的状态感知稀疏操作

Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating

学习协调进化搜索:动态DE-CMA-ES优化与结构模型更新中的进阶感知深度强化学习

Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games

即时战略游戏中带有强盗策略选择的受限指令条件强化学习

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

自回归检索器:提升从项目反馈中获取查询的理解,实现通用多模态检索

Can Jev be Your Q or Policy in Reinforcement Learning?

Jev可以作为你在强化学习中的Q或政策吗?

A 3D Characterization Framework for Intelligent Sequential Decision Making

一个用于智能顺序决策的三维特征框架

What is the goal of unsupervised machine learning?

无监督机器学习的目标是什么?

Distributed Constrained Resource Management in 6G Networks: A Scalable Hybrid Model-Learning Framework

6G网络中的分布式受限资源管理:一个可扩展的混合模型学习框架

WAND: Learning Robust Navigation under Complex Wind Disturbances and Dense Obstacles for Quadrotors

WAND:在复杂风扰和密集障碍条件下学习四旋翼机的稳健导航

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

ConventionPlay:能力有限的培训,促进强有力的临时协作

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

GRPODropout:在线强化学习推广的“少即是多”

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

MiMo-v2.6:面向自我提升的扩展强化学习

CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

能力:通过行为潜在编码实现能力感知策略适应

When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation

代理者何时应思考?通过交叉回合估计的自适应推理

Predefined-Time Integral Reinforcement Learning for Unknown Nonlinear Systems via Inverse-Optimal Design

通过逆最优设计实现未知非线性系统的预定义时间积分强化学习

Credal Machine Learning for Risk-Averse Decision Making

Credal 机器学习用于风险规避决策

Q-Shaped Options for Hierarchical Reinforcement Learning

Q 形的层级强化学习选项

Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility

利用人机参与的演示在强化学习中实现数字孪生驱动的机器人灵活性

Sim-to-Real RL for ASVs using SysID

使用 SysID 的模拟到真实 ASV 强化学习

VibeEdit: Image Editing with Canvas Instructions

VibeEdit:使用Canvas操作说明进行图像编辑

Walking on Roofs: Exploring the Potential of Walking Robots for Construction Work on Roofs

屋顶行走:探索步行机器人在屋顶施工中的潜力

HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

HarnessSQL:现实数据库环境中的SQL代理本地开发培训

Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

人工智能代理能否通过学习登顶?在长期游戏代理竞赛中评估启发式学习

AgentGarten: Code Worlds for Evolving Agents

AgentGarten:进化特工的代码世界

A Unified Bellman Operator for Safety-Critical Reinforcement Learning

一个统一的贝尔曼操作员,用于安全关键强化学习

FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems

FAITH:可行性感知安全过滤强化学习,适用于高维系统

Generative Neural Retargeting for Human-to-Robot Dexterous Manipulation

生成神经重定向用于人对机器人灵巧操作

A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control

平衡数据饮食:解决机器人控制大规模强化学习探索瓶颈

Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration

Dex-One2Many:通过一次人类演示学习灵巧操作

Keyword: diffusion policy

Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams

诊断与恢复长视界技能层的观测空间转移

SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills

SkillWeave:将异质演示融入长远视野操控技能