生成时间: 2026-10-06 00:55:27 (UTC+8); Arxiv 发布时间: 2026-10-05 20:00 EDT (2026-10-06 08:00 UTC+8)

今天共有 45 篇相关文章

Keyword: reinforcement learning

Lexicographic Multi-Objective On-Policy Distillation

词典序多目标政策提炼

Energy Saving in 5G and Beyond Networks: A Quantum Reinforcement Learning Approach

5G及更远网络中的节能:一种量子强化学习方法

Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments

核物理散射实验中靶极化优化的强化学习技术

Tropical Reinforcement Learning

热带强化学习

Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning

多保真度策略梯度稳定数据稀缺的强化学习

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

学习下一步调查内容:长远研究代理的元推理

Reward Inflation: A Healthy Stimulus for Reinforcement Learning

奖励膨胀:强化学习的健康刺激

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

通过分解值梯度流进行测试时间多代理协调

Test-time Calibration Learning for Large Language Model Reasoning

大型语言模型推理的测试时校准学习

Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

从演进错误中学习:策略上蒸馏的自适应迭代修复

Bellman Error Minimization Via Linear Programming Normalization

通过线性规划归一化实现贝尔曼错误最小化

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

前瞻性回顾:通过预测与现实差距进行自我校准的强化学习

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

强化学习前的门诊:基于评分标准的暖启动强化学习,配合政策提炼

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

VIGOR:基于模型的强化学习中通过潜在空间一致性实现零样样视觉推广

Text-Centric Post-Training for Omni-Modal Reasoning

全模态推理的文本中心后期培训

MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

元评分标准:学习基于评分标准的强化学习奖励

All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR

只工作不玩耍让杰克变得无趣:理解并防止RLVR中灾难性策略崩溃

Understanding Enrichment in Reinforcement Learning

理解强化学习中的丰富化

Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning

开放团队多智能体强化学习中的交替正交学分分配

FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization

FARM:多任务无线网络优化的基础代理奖励模型

RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction

RASPER:临床记录的奖励对齐摘要,用于EHR结局预测

hacktrace: behavior-supervised detection of reward hacking during code generation

Hacktrace:代码生成过程中的行为监督奖励黑客检测

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

HARPO:忠实且富有创造力的语言生成的幻觉感知强化学习

Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings

通过动态嵌入学习无作用时间序列的可转移策略

Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

超越单一视频:基于基准测试与主动证据寻求电子商务跨视频推理

How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning

如何在终身强化学习中寻找并重复使用持续适应的策略

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

政策提炼的进展与崩溃:强化学习视角

Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case

除非证据表明:教大语言模型调查员何时结案

Toward SLM-based agentic task-tool intent matching

迈向基于SLM的代理任务工具意图匹配

EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation

EVOL:模拟器引导的进化专家综合,用于无部署学习路径推荐

VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction

VenusRL:一个完全拆分的代理强化学习系统,具备优先调度和可扩展交互功能

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

双向Voronoi偏向的强化学习探索课程

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

超越熵:视频推理中的自我诊断多角色令牌优化

OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation

OuroReward:文本到三维生成中强化学习的顺序奖励调度

Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition

少测量,了解更多:自我监督测试时间特性获取

Hierarchical Control via MPC-RL for Multi-Timescale Battery Systems

通过MPC-RL实现多时间尺度电池系统的分层控制

Mastering Atari 2600 Games with Discovered Options

通过发现选项掌握Atari 2600游戏

UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning

UniIntervene++:一种高效现实世界强化学习的自适应干预代理

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

NeutronGym:面向大型语言模型代理的物理级中子仪器设计

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

关键处署名:终端代理依赖感知策略优化

On the Convergence of Success Conditioning for Policy Optimization

关于成功条件与政策优化的趋同

Planning to Learn

计划学习

Keyword: diffusion policy

CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

CriticHack:在机器人政策优化下评估视觉奖励

AdaTempo: Learning Shared Relative Tempo from Demonstrations for Faster Robot Manipulation

AdaTempo:通过演示学习共享相对节奏以实现更快的机器人操作

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

等变视觉-触觉扩散策略用于接触丰富操作