生成时间: 2026-09-03 20:39:36 (UTC+8); Arxiv 发布时间: 2026-09-03 20:00 EDT (2026-09-04 08:00 UTC+8)

今天共有 22 篇相关文章

Keyword: reinforcement learning

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

WMLLM:通过预测然后行动世界建模实现自我演化优化代理

DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving

DiDrive:一个风险感知的分层扩散框架,用于自动驾驶中安全离线强化学习

PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

PRISM:一种用于自动驾驶交通系统中主动安全的代理多模型架构

Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control

Sim2Signal:交通信号控制的模拟到现实基准测试

Reinforcement learning to choose optimizers

强化学习选择优化器

Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge

导入你需要的:学习何时以及如何用外部知识增强EHR图表

Thinking effort aligns between humans and reasoning models in abductive reasoning

在溯因推理中,人类与推理模型之间的思考努力是一致的

OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

OR-Transformer:实时决策规模提升至1000项

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

策略内提炼遇上非策略GRPO:培训紧凑的跟随教学者

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

多行动,少做决定:长视野大型语言模型代理的技能引导自适应行动分块

IDEEA: training-free Input-Dependent stEEring via Activation cluster matching

IDEEA:通过激活簇匹配实现无训练输入依赖的stEEring

DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

DMRL:文档介导强化学习,用于广告推荐技能优化

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

PhoenixNest视频:基于证据的多模态代理自动化视频访谈评估框架

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

PGPO:多回合代理任务的潜在引导策略优化

Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

远程目标条件强化学习中的递归价值学习

RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution

RideSkill:一种基于LLM驱动自动演进的泛化共乘分层算法

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

APEx:针对自适应深度研究问答的代理程序经验提炼

A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN

基于GNN的L2RPN电网控制图表示的比较研究

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

向谁学对的人:多领域大型语言模型的答案验证多教师提炼

GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design

GDB-奖励:从评估指标到平面设计培训奖励

Cliff: Learning Process Rewards from the First Mistake

悬崖:从第一次错误中获得的学习过程奖励

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

编程竞赛金牌表现的训练后语言模型

Keyword: diffusion policy

There is no result