生成时间: 2026-08-24 16:49:10 (UTC+8); Arxiv 发布时间: 2026-08-24 20:00 EDT (2026-08-25 08:00 UTC+8)

今天共有 20 篇相关文章

Keyword: reinforcement learning

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

RLVR中的多语言验证偏差:基准、推广诊断与跨语言选择瓶颈

World models of environment, agent and joint agent-environment systems

环境、代理及联合代理-环境系统的世界模型

From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing

从热偏好预测到适应性热干预:利用生理与环境感知的强化学习方法

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

注释作为推广:视频多层次营销的高效且可扩展的强化学习

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

AgentMercury:您的代理可以大规模综合可验证的业务场景环境

Why2Speak: Faithful Reasoning for Abstaining Action Policies

Why2Speak:忠实理由支持弃权行动政策

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

连续时间跳跃马尔可夫决策过程的强化学习及其在网络动态定价中的应用

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

CDRL:认证驱动的强化学习用于中微子味道模型发现

Towards Faithful Simulation of Human Shopping Behavior

迈向对人类购物行为的忠实模拟

CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting

CAS:通过自适应检索和策略加权实现的规范化代理搜索

Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision

通过高效的分段到视频监督,增强长视频理解的本地化推理

Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards

通过引导辅助课程和分阶段奖励,自然坐立动作合成

Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface

演示引导的人形单口喜剧,在模拟可变形表面上

Sharing the Control Authority Between Deep Reinforcement Learning and Model Predictive Control: Application to Multi-Class Transportation Networks

深度强化学习与模型预测控制之间的控制权共享:多类别运输网络的应用

Multi-Objective Deep Reinforcement Learning for Secure and Stable Power System Operation

多目标深度强化学习,实现电力系统安全稳定运行

Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning

教学是一个过程:TOSS框架用于模拟人类交互机器人学习中的人类教学决策

SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

SRL-MPC:形状感知强化学习模型预测控制

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

超越模仿:通过非策略Q规划实现自我改进的机器人政策

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

AUSO:从内化到利用的行动级统一技能优化

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Re$^3$Cap:通过强化学习实现图像字幕增强的检索引导精炼

Keyword: diffusion policy

There is no result