生成时间: 2026-09-16 21:13:58 (UTC+8); Arxiv 发布时间: 2026-09-16 20:00 EDT (2026-09-17 08:00 UTC+8)

今天共有 26 篇相关文章

Keyword: reinforcement learning

ViCo: Visual-oriented Coding with Self-Reflection for Chart Replication

ViCo:以图表复制为基础的视觉化编码与自我反思

Managing Action Preconditions in Neuro-Symbolic RL: Three Placement Strategies for Embodied Agents

神经符号强化学习中的行动前置条件管理:具身代理的三种配置策略

Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation

迈向可扩展RLVR:数据综合与蒸馏后的多模态指令

World-Action Models for Robot Learning and Control: A Survey

世界行动机器人学习与控制模型综述

The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis

模仿游戏:当大型语言模型通过以代码为中心的推理数据综合学会像程序一样推理时

How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning

我如何学会停止担忧并热爱StopGrads:平稳性、收敛性以及一个关于流程图学习的案例研究

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

虚假工具的使用:当强化学习者学会错误的行动理由时

Symmetric solution of the Bellman optimality equation for repeated harmony game

重复和声博弈贝尔曼最优方程的对称解

Policy Gradient over History-Dependent Policy Classes for LQR with Domain Randomization

LQR与域随机化的历史相关策略类策略梯度

Autonomous Droplet Navigation via Model-Based Reinforcement Learning

通过基于模型的强化学习实现自主液滴导航

Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching

通过梯度聚类BS采样实现的快速收敛元强化学习(Meta-RL)用于边缘缓存

Register Tokens for Bounded-State Reasoning in Diffusion Language Models

扩散语言模型中有界状态推理的寄存器令牌

UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data

UniDex-ViTac:从人类视频数据学习统一的Visuo-触觉灵巧操作策略

A Cyber Range Evaluation of Autonomous Network Incident Response Agents

自主网络事件响应代理的网络范围评估

GrowMTP: Can RL Grow Its Own Draft Head?

GrowMTP:现实生活能培养自己的选秀负责人吗?

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

奖励推理,而非答案:医学质量保证中的测试时间强化学习修正与界限

TIAO: Token Importance-Aware Policy Optimization for Text Summarization

TIAO:令牌重要性感知策略优化文本摘要

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

不可可能的评分标准:压力测试生成的评分标准作为奖励信号

Interactive Memory Learning for Long-Term Conversations

长期对话的互动记忆学习

Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand

手指作为腿:用拟人化手学习自我支撑的运动与操作

MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation

MOCC-R1:强化推理-反应一致性以生成多模态咨询师反应

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

FluxVLA引擎:实现具身智能的一站式VLA工程平台

Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record

抓骗子容易,诚实骗子难:语言模型从验证记录诊断出奖励通道损坏

Calibrate Once, Fly Any Team: Residual-Grounded Low-Fidelity Training for Cooperative Drone Swarms

校准一次,任意飞行:合作无人机群的残余地面低保真训练

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

ScienceBuddy:交互式科学代理的递归中递归自我改进

Keyword: diffusion policy

Dissecting Motion-Prior Regularization for Data-Scarce Robotic Insertion

数据稀缺机器人插入的运动先验正则化剖析