生成时间: 2026-08-25 16:44:17 (UTC+8); Arxiv 发布时间: 2026-08-25 20:00 EDT (2026-08-26 08:00 UTC+8)

今天共有 55 篇相关文章

Keyword: reinforcement learning

AI Learning and Conceptual Transfer in the Game of Hidden Rules

隐藏规则游戏中的人工智能学习与概念转移

Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?

模型中的模型:何时由专业人员生产胜过关注、适应或调谐?

Runtime Action Interference for AI Control of AlphaStar in StarCraft II

星际争霸II中AI控制AlphaStar的运行时动作干扰

Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning

让学分跟随计算:大型语言模型强化学习的架构感知学分传输

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

权重中被遗忘,工具恢复:大型语言模型代理的代理工具去学习

Force/Torque-Based Kinematic Adaptation for Robotic Manipulation Tasks

基于力/扭矩的运动学适应用于机器人操作任务

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

扰乱思想,而非像素:潜在空间推广多样化以强化视觉语言模型学习

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

基于良性事实的强化学习加剧了记忆中的私人数据泄露

CounterAlign: Counterfactual Supervision for Vision-Language-Action Models

CounterAlign:视觉-语言-行动模型的反事实监督

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

SAFE-G:结构感知、忠实证据引导的知识型视觉问答生成

MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning

MCite-RL:通过引用增强的代理强化学习迈向可靠的多模态RAG

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

提示、批评与教师:视觉语言数学推理中稀疏奖励强化学习的先例注入

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

超越成功与失败:面向图形界面代理的长度感知对比学习

LLMs are Few-Shot Decision-Makers: Generalized Context-Aware Microgrid Frequency Control through Prompt Decision Transformer

LLMs是少数方决策者:通过即时决策变换器实现的通用上下文感知微电网频率控制

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

HiDiffTIR:多回合工具集成推理的层级难度感知策略优化

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

BioMed-Agent-RL:一种元学习,生物医学应用所需的一切

The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning

追逐是课程,捕获是功劳的锚点:追踪-回避自我游戏,为零数据大型语言模型推理

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

从求解器反馈到忠实计划:符号规划中的多角色强化学习

CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

CIDER:具身强化学习的持续互动蒸馏

Beyond Fixed Directions: Adaptive Representation Analysis of Reasoning and Memorization in LLMs

超越固定方向:大语言模型中推理与记忆的自适应表征分析

ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

ESCRAG-R1:情感支持对话中的检索增强强化学习

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

EDGE:引导探索代理强化学习的体验提炼

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

ToSCA:利用分层强化学习对会话代理的时间和战略抽象

DELTA: Deformable Elevation-Based Local Terrain Attention Encoder for Sparse-Terrain Quadrupedal Locomotion

DELTA:稀疏地形四足行走的可变形基于高程的局部地形注意编码器

Decoupled Physical Modeling and Execution for Physics Reasoning

物理推理中的解耦物理建模与执行

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

MCP-Universe RL:通过强化学习训练MCP工具使用代理的框架

MARL-Based Sequential RIS Auctions: A Physical-Layer Security Analysis

基于MARL的顺序RIS拍卖:物理层安全分析

UR$^{2}$-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

UR$^{2}$-MLLM:多模态大型语言模型中不确定性感知的再访推理以生成放射科报告

Risk-Sensitive Reinforcement Learning with Smoothed Quantile Objectives

风险敏感强化学习与平滑分位数目标

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

从测试中学习:深度强化学习代理的自指差分测试

WAM-OPD: On-Policy Distillation for World Action Models

WAM-OPD:世界行动模型的政策提炼

Small Reasoning Models are Instruction Followers in Function Calling

小推理模型是函数调用中的指令跟随者

Scaling Curriculum Learning For Autonomous Driving

自动驾驶课程学习规模化

Coalition-Aware Skill Reliability for Self-Evolving Agents

联盟感知技能可靠性,适用于自我进化的特工

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

DeepSAGE:结构化认知行为疗法咨询对话的阶段感知强化学习

Enhancing Sim2Real Transfer for Torque-Controlled Robots through Real2Sim Dynamics Estimation and Reinforcement Learning

通过Real2Sim动力学估计和强化学习,增强扭矩控制机器人的Sim2Real传输

Learning Generalizable Behaviors for Terminal Agents

学习终端代理的可推广行为

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

GeoRisk-RAG:一个层级感知的风险框架,通过选择性回答提升RAG可靠性

Spiking Neural Networks for Continuous Control: Neuromorphic Reinforcement Learning in Conventional Computing

持续控制的尖峰神经网络:传统计算中的神经形态强化学习

Can We Perform Online RL for Image Editing without Editing Rewards?

我们能否在不获得编辑奖励的情况下进行在线强化学习进行图像编辑?

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

TailSieve:用于LLM部署的部分展开引导尾路由

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

FinixDoc:重新思考财务文件解析,超越饱和基准

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

GSAR:具有自我演进数据综合的移动图形界面代理的目标-状态-锚点奖励

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

PsychJail:通过多回合说服LLM策略探讨心理越狱

Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction

人工同理心:迈向无监督机构检测与政策重建的框架

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

MobilePA-Bench:为移动规划代理在复杂现实任务中的基准测试

Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning

使用强化学习进行噪声定位下的宏动作拓扑导航

Distributed Trajectory Planning and Resource Allocation for Dynamic Multi-UAV Collaborative Computing

动态多无人机协作计算的分布式轨迹规划与资源分配

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

引导黎曼优化(GuRO):模型预测控制与决策变换器的桥接

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Agent-G$^2$:智能强化学习的高斯指导

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

超越视频的思考:统一视频推理与开放世界视频代理的深入研究

Temporal Property-driven Design Space Exploration with Reinforcement Learning for Cyber-Physical Systems

基于时间属性的设计空间探索与强化学习,适用于网络物理系统

Reward-Free Continual Adaptation for Resilient Space Robots

无奖励的持续适应,打造韧性空间机器人

How to Train a Critic Stably and Efficiently

如何稳定高效地培训评论家

Keyword: diffusion policy

ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD

ODG-NoMaD:俯视摄像头方向制导NoMaD