生成时间: 2026-08-19 16:39:47 (UTC+8); Arxiv 发布时间: 2026-08-19 20:00 EDT (2026-08-20 08:00 UTC+8)

今天共有 30 篇相关文章

Keyword: reinforcement learning

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

通过LLM引导强化学习打造高效的个性化AI导师

FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences

FetchMan:从模拟体验中学习可视化人形机车操作策略

Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation

Lambda保持控制:预测性肌肉骨骼模拟中,从最小任务奖励中产生类人运动

Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis

基于物理的强化学习用于随机距离-避免分析

Q-Learning With World Models

Q-世界模型学习

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

情境强化学习的任务专门化微调

Reinforcement Learning as (Discrete) Potential Theory

作为(离散)势理论的强化学习

An O-RAN-Assisted MARL Approach for Dynamic Sidelink and Infrastructure Selection in V2X Communications

一种基于O-RAN辅助的MARL方法,用于V2X通信中的动态侧链和基础设施选择

Safe Deep Reinforcement Learning for Energy-Efficient HVAC Control in Multi-Zone Residential Buildings

多区住宅建筑中节能暖通空调控制的安全深度强化学习

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Co-RL:多代理强化学习中多元群体中出现的无监督推理

A Hybrid End-to-End and Modular Control Architecture Toward Safe Vehicle Lateral Control: Combining Soft Actor-Critic with Model Predictive Control

一种混合端到端和模块化控制架构,实现车辆横向安全控制:软行为者-批评与模型预测控制的结合

The Road Less Traveled: Congestion-Aware NoC Placement and Packet Routing for FPGAs

少有人走的道路:FPGA的拥塞感知NoC布置与数据包路由

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

SignalReasoner:评估3B模型在信号数学推理中的上界

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

代理型ESOpt:以最小GPU需求微调长视野LLM代理

Robust Brachiation on a Life-Sized Dual-Arm Robot Using Waypoint-Guided Reinforcement Learning

利用航点引导强化学习对真人大小双臂机器人进行强健的断压

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

结合新奇与惊喜,以提升基于图像的强化学习中的体验优先级和探索

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

LEGO-RL:编码代理的绑架-原生强化学习

REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

REChart:大型推理模型的高效推理图表编辑

Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

棱镜-GRPO:通过拆分相同结果组实现更快的VLA策略优化

Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

迈向更优秀的多回合用户交互代理:下一次用户回合不仅仅是上下文

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

评估强化学习可解释性方法对修复代理漏洞的帮助程度

tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

tinyDSM:资源受限毫米机器人的技能建模与开发框架

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

迭代抓握姿态精炼:二维视觉的深度强化学习方法

rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment

rl-triton:用于强化学习学分作业的高性能Triton GPU内核

Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control

基于物理知情的世界模型进行合作混合交通控制的离线多智能体强化学习

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

用低资源语言思考:SFT构建了什么,强化学习修复了什么,准确性看不到什么

Debate Training Reduces Reward Hacking in RLAIF

辩论培训减少RLAIF中的奖励黑客行为

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

通过图结构在线难度估算实现高效的RLVR调度

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

迈向零样子任务转移,采用神经符号世界模型

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

基于LLM反馈的策略不变奖励塑造:混合强化学习代理框架

Keyword: diffusion policy

There is no result