生成时间: 2026-09-29 22:49:42 (UTC+8); Arxiv 发布时间: 2026-09-29 20:00 EDT (2026-09-30 08:00 UTC+8)

今天共有 144 篇相关文章

Keyword: reinforcement learning

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

MaD-RL:校准LLM与强化学习的匹配分布

Humanoid Badminton: Learning Dynamic Racket Skills from Limited Human Motion Data

类人羽毛球:从有限的人体运动数据中学习动态球拍技能

PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning

PHIRL:将学习奖励与任务进展对齐,实现逆向强化学习

Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining

股票交易的深度强化学习:基于前向再训练的基准分析者-批评者方法

FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization

FARE:深度强化学习,实现公平暴露、受限不确定性意识的金融内容个性化

CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense

CyberWorld:样本高效自主网络防御的世界模型

Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training

理解 LLM 后期培训中 SFT、RLVR 和 OPD 之间的协同效应

Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning

利用强化学习增强胸部X光报告生成中的视觉推理能力

On-Policy Attention Linearization

政策上的注意力线性化

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

CaptchaArena:一个大规模、细粒度的数据集,用于在交互式验证码上训练计算机代理

Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation

交互式分布稳健多智能体学习,采用一般函数近似

Graph Forward Distribution Matching for Molecular Inverse Design

分子逆设计中的图前向分布匹配

Grasp2Twist: Learning Bimanual Dexterous Jar Opening by Reinforcement Learning

Grasp2Twist:通过强化学习学习双手灵巧罐开

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

找到你做不到的事:针对自我改进VLA模型的代理现实强化学习

When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk

何时人类应重新掌控?在动荡的人工智能风险下的最佳委派

Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions

发挥到标准:可证明的最优四边形块分解的强化学习

Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles

不确定性感知在线算法的模拟集群选择

Noisy Test-Time Reinforcement Learning for Code LLMs

针对代码大型语言模型的噪声测试时强化学习

Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments

见证:互动谜题环境中的发现、解读与顿悟

RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process

RoboFFT:通过在线强化学习和前向过程精细化生成机器人政策

Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging

Train4Merge:一项针对强化学习与SFT教师在基于门诊教学(OPD)模式合并中的受控单一教师研究

All On-Board: Fully On-Chip Neuromorphic Q-Learning with Embedded CartPole Simulation

全程上载:全片上神经形态Q学习,内置CartPole模拟

RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning

RLHarness:与强化学习共同演进的过程性技能,用于长视野多模态推理

SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO

SGA-Flow-GRPO:流动GRPO的空间梯度引导学分分配

OpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRI

OpenMASC:一个用于加速MRI中交叉轨迹金属感知采样与校正的开源流水线

Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization

超越脚本化搜索:通过代理黑箱优化实现样本高效奖励发现

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

重新思考大型语言模型强化学习中的训练-推理不匹配:其出现之处及纠正方法

On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models

关于口头置信先验在校准大型推理模型时的陷阱

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

AdaTutoRank:通过自适应辅导优化学习如何重新排序文档集,用于RAG和深度研究

From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning

从结果到策略:学习策略在数学推理中的实用性

What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor

ProcGen推广差距衡量什么?动作规则、收敛性与缺失的随机底层

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Code Agent RL 的分组代理分级与优势再分配

SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents

SWE-MILE:长视野软件工程代理的异步电位诱导里程碑学分分配

PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models

PF-RL:通过目标条件值几何学实现视觉-语言-行动模型的进步场强化学习

Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

时间步权权:有效基于ELBO的流量匹配强化学习的隐藏钥匙

Scaling Properties of Same-Family On-Policy Distillation

同族策略蒸馏的缩放性质

Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication

分布稳健平均奖励强化学习:弱交流下的有限样本保证

Retrospective Distillation Attribution via Normalized Response Similarity

通过归一化反应相似度进行回溯性蒸馏归因

CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning

CUA沙盒:计算机使用代理强化学习的高效环境

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

CT-OPD:扩散视觉语言模型中的反事实追踪策略提炼

Understanding and Exploiting Anisotropy in Post-Training

理解和利用后期培训中的各向异性

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

OpenTumorBoard:多学科肿瘤委员会讨论轨迹的现实基准

Improving LLM Collaboration via Multi-Agent Preference Learning

通过多智能体偏好学习提升LLM协作

Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping

重新定时的钟声流:逃离速度自助的不可能三角

Allspark: Weak to Strong Transfer via Alternating Chain of Thought

全火种:通过交替思维链由弱到强转移

Counting on Thinking: Tracing Evidence Integration in Language Models

依靠思维:在语言模型中追踪证据整合

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

结构化受限MDP中在线强化学习的最后迭代保证

Communication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task Allocation

用于去中心化多人多机器人任务分配的通信感知异构图学习

Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View

受限流策略更新:广义薛定谔桥视图

Self-Confirming Superposition Traps in Reinforcement Learning

强化学习中的自我确认叠加陷阱

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

模型还懂另一种方式:策略切换以实现有效RLVR探索

Evolving Dexterous Robots from Scratch

从零开始进化灵巧机器人

MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning

MedRouter:解析基于路由推理的医学大型语言模型知识差异

Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface

一起训练还是以后合并?通过共享动作界面统一VLA专家

Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL

保存你饱和的数据:在基于群体的强化学习中超越奖励饱和

Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation

政策可塑性在离线到线上强化学习中至关重要:为在线适应调整线下政策

Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?

任务的天花板:变压器什么时候能在没有思维链的情况下取得成功?

Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?

测试时间的缩放和训练在个人站姿预测中有哪些不足?

Downstream-Aware Context Selection for Online In-Context Reinforcement Learning

在线上下文强化学习的下游感知上下文选择

When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning

模型何时承认自己错误?在强化学习下,失败披露是不稳定的

RMB: Reward Model Boosting Mitigates Reward Hacking

RMB:奖励模式提升缓解奖励黑客行为

CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents

CodeSkill:长期代码代理的潜在技能抽象

ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents

ActiveMem:面向长视界代理的动态潜在记忆树

Minimax-Optimality of Posterior Sampling for Reinforcement Learning

后置采样的极小极大最优性用于强化学习

GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences

GTRL:扎根、分化并征服时间差异的价值学习

Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space

接下来的思考是分布:潜在空间中的生成自回归推理

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

学习销售:多产品市场中战略性大型语言模型代理的强化学习

Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models

校准,而非答案选择:推理模型中内部信心的提炼

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

安全评分匹配:传播政策与汉密尔顿-雅各比可达性,用于在线安全强化学习

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

超越时间戳:长期战略代理的决策对齐政策提炼

COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement

COEVO:递归自我提升的共演化背景与参数

SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

SciGen-Verifier:科学图像生成中可解释验证的多模态推理器

VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings

VaME:探索多模嵌入的变分潜在推理

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

TTRSD:视觉语言模型的测试时间强化学习与自我蒸馏

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

TeacherGRPO:通过教师对齐缩小推理提炼能力差距

Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning

阐明基于回归的扩散强化学习设计空间

DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

DISCO:带有基础推理分解的分布式长上下文尺度

Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning

利用多智能体强化学习实现的联合多模态人类活动识别

MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

MA-JEPA:多智能体强化学习的联合嵌入世界模型

OpenFC: Learning Verification Policies towards Open-Search Fact Checking

OpenFC:学习验证政策以实现开放搜索事实核查

TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

TGRL:温度分组强化学习,用于高效探索大型语言模型

You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

你只需编辑一次:通过本地演示细化激励LLM的上下文内能力

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

ParaAgent:在开放世界工具环境中强化并行行动

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

攀登高峰:提示注入红队对抗前沿模式与课程强化学习

Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss

基于通信丢失下仓库机器人协调的层级多智能体强化学习

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

无演示成功概率奖励学习适用于通用机器人政策

LieDiscover: Adaptive Symbolic Library Construction for Explicit Open-form Symmetry Discovery

LieDiscover:显式开放形式对称发现的自适应符号库构造

CompoWorld: Compositional Environment Scaling for General Agents

CompoWorld:通用代理的合成环境尺度

Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

共享经验,独立学习:伴随置信度校准用于大型语言模型

DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

DuraS2ST:针对时长对齐的语音转语音翻译的思维链与强化学习

Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies

冻结生成机器人策略在线适配的主要引导子空间

Selecting Diverse SFT Traces Improves Post-RL Generalization

选择多样化的SFT迹可以提升强化后推广能力

Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning

意外成功,反复失败:熵引导学分作业用于探索大型语言模型推理

RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation

RAISE:通过符号评估强化LLM中的访问控制策略综合

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

QwenGyre:用于训练xLong-Horizon代理的弹性强化学习框架

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

R$^2$ 流:通过递归技能进化实现递归自我提升

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

反事实推广重放:可分叉环境作为软件工程代理的免费流程奖励

DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation

DexTaG:灵巧操作强化学习中的触觉指导

No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability

无免费效率:重新审视训练效率与模型脆弱性之间的权衡

ICMAPE: In-Context Multiagent Pure Exploration

ICMAPE:上下文多智能体纯探索

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

奖励解释器:从反事实偏好反馈中学习奖励模型解释

ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim

ZeroBot:通过生成式Real2Sim几分钟内从零开始学习

Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation

视觉——受限强化学习中的语言信号:安全优势而无预期

Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

学习具有稳定优化的LLM代理的扰动稳健策略

PhysFieldBench: Can Multimodal Models Understand Physical Fields?

PhysFieldBench:多模态模型能理解物理场吗?

Evolution of fairness in multi-objective reinforcement learning framework

多目标强化学习框架中公平性的发展

Learning to Optimize through Solver-Grounded Self-Play

通过解算器基础的自我游戏学习优化

GlyphBench: A Playground for Language-Model Reinforcement Learning

GlyphBench:语言模型强化学习的游乐场

MAS-OPD: On-Policy Distillation for Multi-agent Systems

MAS-OPD:多智能体系统的策略提炼

AdaGuard: An Adaptive Guard Model with User-defined Policies

AdaGuard:一个带有用户自定义策略的自适应保护模型

SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents

SALMONN-duo:全双工语音代理的自适应双系统协调

Evolving Support Priorities in Empathetic Reinforcement Learning

同理强化学习中支持优先事项的演变

Rotated Manifold Optimization for Low-Rank Adaptation

低秩适应的旋转流形优化

ZeroCode: On-demand Error-Correcting Code Construction from the Zero Matrix via Reinforcement Learning

ZeroCode:通过强化学习从零矩阵构建按需纠错代码

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

参见、测量与推理:学习病理学中的视觉基础推理

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

博士学分:深度研究代理的评分标准基础过程学分作业

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

知道何时思考还不够:教导小推理模型超越其参数化知识

Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors

学习引导,引导观察:通过可训练向量揭示大型语言模型中RLVR的几何结构

Improving Large Language Models for Code through Runtime Program-State Reasoning

通过运行时程序状态推理改进代码大型语言模型

Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

涌现,而非带宽:物理耦合与多智能体通信学习的局限性

ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control

ABC-Align:预测驱动的对齐与自适应偏倚控制

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

马拉松者:超长视野自主智能

CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

CAR-VLA:自动化的复杂性感知和风险适应性推理

Distilling Visual Reasoning into Text Space

将视觉推理提炼进文本空间

OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue

OSPD:针对人格一致对话的政策自我提炼

Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning

编码代理记忆后训练:通过强化学习释放预训练文件操作的记忆潜力,用于长视野任务

Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

Q-learning 因安全离线强化学习而受罚的变换器

LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

LLMs作为自适应元求解器:多策略强化学习用于工业规模优化

Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning

逃离局部视图:发现可解释多智能体强化学习的潜在概念

Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction

通过逐步预测实现的模型知情安全强化学习

FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL

FlexLoop:深度弹性循环策略用于深度强化学习中自适应测试时间计算

Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets

人工智能能在加密货币领域赚钱吗?衡量从回测到真实市场的差距

Verifying Neural Networks with Reinforcement Learning

通过强化学习验证神经网络

Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

远景离线目标条件强化学习的扩散子目标规划

Reinforcement Learning from Intermediate Renders for Image-to-Code Generation

从中间渲染中强化学习图像到代码生成

Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation

$Q^*$估计的深度加权贝尔曼残差最小化

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

接地前总结:查询引导块凝聚,用于长视频时间接地

The Low-Rank Structure of VLA Reinforcement Learning

VLA强化学习的低秩结构

Nereus: Adaptive Parallelism for LLM Post-Training

Nereus:LLM 后期培训的自适应并行

Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

统一轨迹匹配策略优化:多样化的T2I生成与VLA泛化

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

稳定控制中零阶奖励塑造政策梯度的充分性

Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning

质量决定方向,长度决定形状大小:长度控制用于开放式强化学习

DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations

DexWeave:通过人类演示学习灵巧的人形机车操控

Keyword: diffusion policy

CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

资本支出:通过体验自适应推理,高效将基础模型行为提炼为可部署机器人策略