生成时间: 2026-10-06 22:54:38 (UTC+8); Arxiv 发布时间: 2026-10-06 20:00 EDT (2026-10-07 08:00 UTC+8)

今天共有 97 篇相关文章

Keyword: reinforcement learning

What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study

可验证的奖励教会视频语言模型关于时间的什么?一项受控多模型研究

Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking

晶体发生器离散组成通道上的强化学习:验证的收益与奖励黑客

Barrier-Shaped Recurrent Reinforcement Learning for Autonomous Landing on a Heaving Ship Deck

障碍状反复强化学习,用于在颠簸船甲板上自主着陆

Network Adaptation in IRS-Aided Hybrid RF/VLC Systems Using Cooperative Multi-Agent DRL

利用合作多智能体日程学习(DRL)的IRS辅助混合射频/极低噪声系统中的网络适配

Exploration-Preserving Policy Optimization

勘探保全策略优化

Reinforcement Learning with Comparative Evidence for Social Intelligence

社会智能的强化学习与比较证据

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

IdeaScientist:策划扎实科学构思的代理

PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

PB-GRPO:通过偏好批处理GRPO从人格驱动模拟学习社会适应LLM代理

How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws

强化学习如何重塑大型语言模型推理:可转移性、覆盖率与扩展性规律

Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks

异步几乎可以自由地用于长视野能动任务的演化策略

ROOT: Discovering Rewards for User-Specified Embodied Behaviors

ROOT:发现用户指定具身行为的奖励

Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry

RLVR在融合格罗莫夫-瓦瑟斯坦几何上的层级学分作业

Large Language Models and Augmented Democracy

大型语言模型与增强民主

CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies

CORE-RL:基于信心的黑盒强化学习策略可靠性评估

AgroGround: Multi-Granularity Grounded Recognition in Agriculture

AgroGround:农业中的多粒度基础认可

LocusRL: Diagnosing LLM Reward and Policy Interventions in Competitive Games

LocusRL:诊断竞技游戏中的大型语言模型奖励与政策干预

Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?

LLM代理能否自动化文本转语音的强化学习?

DreamTest: World-Model Surrogates for Search-Based Testing of Deep Reinforcement Learning Agents

DreamTest:基于搜索测试的深度强化学习代理的世界模型替代品

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate:针对政策提炼的困难门槛教师指导

Anticipating the Consequences of Curriculum Decisions with Large Language Models

用大型语言模型预判课程决策的后果

Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning

用于从非规范密度抽样的评分校准流,应用于生成在线强化学习

Learning to Clarify Underspecified Intents Under Limited Interaction

在有限互动下学习澄清未明确的意图

PatternDex: Learning Interaction Patterns to Guide Reinforcement Learning of Bimanual Dexterous Manipulation of Articulated Objects

PatternDex:学习互动模式以指导双手灵活操作关节物体的强化学习

Multi-Agent Spectrum Sharing

多智能体频谱共享

CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery

CURIO:以好奇心驱动的测试时间学习,实现开放式发现

Rewrite What Matters: Adaptive Multilingual Query Rewriting for Reasoning via Agentic Reinforcement Learning

重写重要内容:通过智能强化学习实现自适应多语言查询重写推理

Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning

残余视觉学分优化:多模态强化学习的保守证据路由

PWM: Personalized World Models with Online Reinforcement Learning

PWM:带有在线强化学习的个性化世界模型

A Unified Dynamics Framework for Reinforcement Learning and Classical Control of a Six-DOF Pipeline-Tracking ROV in NVIDIA Isaac Sim

用于增强学习和经典控制的统一动力学框架,用于NVIDIA Isaac Sim中六自由度流水线跟踪ROV

How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation

教师应如何准备?关于学生诱导的政策提炼状态的现实学习

AlphaPADI: Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion

AlphaPADI:通过池感知层级离散扩散实现公式化Alpha发现

Outcome-Guided On-Policy Self-Distillation

以结果为导向的政策自我提炼

Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning

在线目标条件强化学习的方向条件政策

Small Agents with Semantic Search: Efficient Multilingual Code Localization

具语义搜索的小型代理:高效的多语言代码本地化

Arithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning

算术演员头与非分布强化学习的训练稳定

The Law of DeepSeek

深搜法则

MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation

MGPO:任务感知数据集蒸馏中的流形引导扩散比对

Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes

有证据的答案:一致性意识的接地视觉问答,适用于路边交通场景

Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

Red-TTT:自动化越狱大型语言模型的测试时训练

Erased, Rerouted, or Rescaled? Post-Training and the Causal Quotient of a Language Model's Belief State

是被抹除、重新定向还是重新调整比例?训练后与语言模型信念状态的因果商

RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation

RubricArmor:对抗进化提升基于LLM的评分标准生成

On Semi-Markov Suboptimality in Hierarchical Reinforcement Learning

关于分层强化学习中的半马尔可夫次优性

Optimal Control with Learned Critics under Unmodeled State Dependencies

在未建模状态依赖下,学习批评者的最优控制

AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

AIProver:通过证书驱动的进化工具实现数学研究的能动自形式化

Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks

Sibyl:一个高效的小-大模型协作框架,用于长期任务

Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking

未提及的检查清单发现改变了强化学习改善胸部X光报告检查的方式

Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning

层级时间感知自助,用于非策略子目标价值学习

Groupwise Distortion Guarantees for Preference-Based Alignment

基于偏好的群组失真保证

AI Safety via Debate is Compromised by Cognitive Biases

通过辩论进行的人工智能安全受到认知偏见的影响

Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

语言模型代理的层级强化学习与稳定时间抽象

An LLM-in-the-loop RL Framework for Bioinformatics Feature Selection

一个用于生物信息学特征选择的LLM在环中强化学习框架

SCOUT: Supply-Aware Cold-Start Proactive Query Suggestion for Travel Search

SCOUT:供应意识冷启动的主动查询建议,适合旅行搜索

Bellman-Centric Learning: Near-Optimal Regret for Linear Bandits with Memory

贝尔曼中心学习:记忆力强的线性强盗近乎最佳后悔

End-to-End Safe Social Navigation via Multi-Task Reinforcement Learning and Probabilistic Perception

通过多任务强化学习和概率感知实现端到端安全社交导航

Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning

通过多目标强化学习,用四足机器人运输未加固的堆叠有效载荷

Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped

针对欠驱动双足行走的无碰撞运动的层级强化学习

Generative-AI for XR Content Transmission in the Metaverse: Potential Approaches, Challenges, and a Generation-Driven Transmission Framework

元宇宙中XR内容传输的生成式人工智能:潜在方法、挑战与代际驱动传输框架

Safe Image Generation via Reinforcement Learning

通过强化学习实现安全图像生成

MEND: RL For Flow Models via Proximal Velocity Matching

MEND:通过近距离速度匹配实现流模型的强化学习

Strategic Multi-Agent Learning for Interpretable Action Valuation of All Players in Football

战略多智能体学习,用于对所有球员进行可解释的行动评估

Reachability-Aware Diffusion Policy Optimization

可达性感知扩散策略优化

Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers

强化学习改进推理教师的转移分层政策提炼

How (and How Not) to Use Data Augmentation in VLA Post-Training

如何在VLA训练后使用数据增强(以及如何不使用)

Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning

自适应专家指导,用于高效策略上的强化学习

Scalable Minimal-Change Learning for Controllable Image Editing

可扩展的最小变化学习用于可控图像编辑

Grounded Joint-Attention Other-Play for Zero-Shot Coordination

接地、联合注意力、其他对抗以实现零射击协调

Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

增强对深度强化学习的可转移对抗性攻击

I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning

I-BFM:通过无监督强化学习实现的奖励条件强健类人互动

Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks

基于强化学习的工作负载感知电力传输网络优化

Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

小型语言模型会学会谈判吗?一项针对强化学习培训卖家的受控规模研究

Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning

受限目标导向平面图生成,结合基于语法的强化学习

Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments

通过混合状态深度强化学习在部分可观察的联网车辆环境中实现匝道计量控制

GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories

GAMBIT:学习规划连续多机器人轨迹

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

VepAgent:通过工具增强强化学习桥接因果过渡,用于视频事件预测

Graph Neural Network-Driven Deep Reinforcement Learning for Scalable RIS Allocation

图神经网络驱动深度强化学习,实现可扩展RIS分配

RollPlace: Improving Macro Placement via Monte Carlo Rollout Search

RollPlace:通过蒙特卡洛推送搜索提升宏置位置

CRAFTER: Causality-based Self-adaptation for Autonomous IoT Systems

CRAFTER:基于因果律的自主物联网系统的自我适应

Dynamic Minimax Regret Optimization for Robust LLM Post-Training

动态极大后悔优化,用于稳健的大型语言模型后训练

MeSD: Multi-Evidence Self-Distillation for VideoLLM

MeSD:视频LLM的多证据自我提炼

Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering

本体概念重叠作为训练信号:基于知识的强化学习用于临床问答

Visual Swarm Navigation via Deep Reinforcement Learning and Evolutionary Hybrid Design

通过深度强化学习和进化混合设计实现可视化群体导航

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

辅助困境:通过多回合强化学习学习教学

GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement

GPlaceRL:一个用于详细定位的开源图强化学习框架

MIRT: Transformers for Truthful Generative Auctions with Whole-feed Permutation Externalities

MIRT:适用于真实生成拍卖的全给置换外部性转换器

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

LoGRA:利用低阶梯度草图扩展大型语言模型强化学习

Considering Context: When World Models Need Context Encoders

考虑上下文:当世界模型需要上下文编码器时

Reward Stealing Attack on Large Language Models

对大型语言模型的奖励窃取攻击

MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks

MedPrune:医疗VQA任务中的拓扑高效多模态多代理通信演进

Improving Diversity in LLM Short Story Generation

提升LLM短篇故事生成的多样性

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

CLIFT:Web代理培训和测试时间缩放的共形自我验证

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

MemPilot:为LLM代理编排按需多模态内存管理

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

迈向正确完成的循环模型,第二部分:固定点的重新思考

Base Models Can Reason By Taking a Cue From Training Data

基础模型可以通过从训练数据中获取提示来推理

Keyword: diffusion policy

The Unexpired Plan: A Free Monitor for Accelerated Diffusion Policies

未过期计划:加速扩散政策的免费监测

Demonstration-Calibrated Port-Hamiltonian Retuning for Manipulation Policies

演示校准的Port-Hamiltonian操作策略重调

Reachability-Aware Diffusion Policy Optimization

可达性感知扩散策略优化

Robotizing Human Videos with Physically Consistent Interactions

机器人化具有物理一致性交互的人类视频