生成时间: 2026-07-21 18:28:21 (UTC+8); Arxiv 发布时间: 2026-07-21 20:00 EDT (2026-07-22 08:00 UTC+8)

今天共有 57 篇相关文章

Keyword: reinforcement learning

Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization

增强学习引导的NSGA-II加灰色关系系数增强多目标优化:在纳斯达克投资组合优化中的应用

Rater State Bias in RLHF Preference Data: An Audit Framework

RLHF偏好数据中的评级者状态偏差:审计框架

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

蒙面扩散语言模型是针对代理强化学习的强大且可引导的文本世界模型

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

它需要8个令牌:弱到强的策略外强强化学习,通过辅助分支

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization

PPO-HSC:基于广域保单覆盖优化的探索性强化学习框架

A Survey on the Verification of Reinforcement Learning Policies

关于强化学习政策验证的调查

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents

CIGPO:多回合证据读取LLM代理的上下文信息获取策略优化

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

从结果到行动:利用事后诸葛亮进行长期语言代理培训

Interactive Task Alignment as a POMDP

作为POMDP的交互式任务对齐

When to Plan: Learning to Select Between Reactive Control and Deliberative Planning

何时制定计划:学会在反应控制与深思熟虑规划之间做出选择

Certifiable Safe Model-Based Reinforcement Learning with Control-Affine Dynamics Approximation

可认证的安全基于模型的强化学习,含控制-仿射动力学近似

Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models

用于视觉-语言-动作模型的远程视野机器人操作的前瞻性残留强化学习

Differentiable Reinforcement Learning for Path Tracking by an Agile Fish-Like Robot

敏捷鱼类机器人用于路径跟踪的可微强化学习

Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning

Building2Building:一个面向可推广现实世界强化学习的大规模基准

From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology

从最优政策到个体差异:重新思考生物学中的强化学习

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

PAVXploreRL:物理-动作-视觉世界模型强化学习结合行动探索

FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images

FUSAR-R1:用于智能解读SAR图像的大规模推理模型

Group Entropy-Controlled Policy Optimization

群熵控制策略优化

Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators

通过无模型的认知自由能估计量实现有原则的无方向内在动机

Trace-Based On-Policy Distillation for Masked Diffusion Language Models

基于跟踪的策略上蒸馏,用于掩蔽扩散语言模型

Enhancing Personalized Bladder Cancer Treatment Through Reinforcement Learning: A Recurrent Patient State Transition Decision Support Framework

通过强化学习提升个性化膀胱癌治疗:反复患者状态转换决策支持框架

Counterfactual Shapley Credit Assignment

反事实沙普利信用分配

Scalable Causal Imitation Learning

可扩展因果模仿学习

Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making

以奖励为驱动的大型语言模型代理工作流:综合POMDP路由与自我纠正以实现自主决策

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

STBridge:在UMM中连接理解与生成的共享目标对齐

EdgeCoInfer: Hierarchical Collaborative Inference for On-Device Multimodal Large Models

EdgeCoInfer:用于设备内多模态大型模型的层级协作推理

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

医患对话的有力总结:TalTech为超越转录挑战而打造的系统

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

LenGuard-GPC:带引导提示一致性的长度防护空间推理强化学习

Distilled Reinforcement Learning for LLM Post-training

LLM后期培训的精炼强化学习

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning

WAR:同步代理强化学习的工作负载感知推广

Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies

玻尔兹曼理性理论的合理化:熵正则化政策的公理化刻画

TAPAS: Throughput-adaptive Perception for Autonomous Systems

TAPAS:自主系统的吞吐量自适应感知

Rethinking the Suitability of Reinforcement Learning Algorithms Under Practical Transfer Constraints

在实际转移约束下重新思考强化学习算法的适用性

CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

CORAL:学习淀粉样纤维配体结合并获得合作结合奖励

TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning

TraversRL:可穿越的步行路径生成与强化学习

Reinforcement Learning: From Algorithms To Foundation Models

强化学习:从算法到基础模型

AGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

AGG:雅可比聚合群梯度用于扩散模型高效GRPO训练

Concentration and Mean-Square Bounds for Contractive Stochastic Approximation: A Unified Elementary Approach

收缩随机近似的集中度与均方界限:统一初等方法

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

ConsiSpace:学习几何一致性对视频空间推理至关重要

On Optimal Event-Triggered Distributed Control for Stochastic Multi-Agent Systems via Reinforcement Learning

关于通过强化学习实现随机多智能体系统最优事件触发分布式控制

Mobile Network Control with a World Model

用世界模型进行移动网络控制

Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning

推广与指导:分解奖励以实现少数样本逆强化学习

Theoretical Foundations of $\max$@$k$ Reinforcement Learning

$\max$@$k 强化学习的理论基础

Distributional Soft Bellman Operator under the Cramér Geometry

Cramér 几何下的分布软贝尔曼算子

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

在通信丢失下实现多智能体协调的价值感知预测

PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks

PRIME:无人机辅助紧急通信网络的多智能体环境中可塑性恢复

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

优势中汇总,而非比率:合作多智能体策略优化的典型形式分析

A Geometric Perspective on Stabilizing Value Conflict Resolution

稳定价值冲突解决的几何视角

Information-Based Exploration via Random Features for Reinforcement Learning

基于信息的随机特征探索强化学习

PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

PAMD:视觉强化学习中双模拟表示的结构化自适应距离

Generalised Bellman recurrence and three dualities in sequential decision-making

广义贝尔曼递归与顺序决策中的三对偶性

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

稀疏证据可以满足:代理证据寻求多模态视频错误信息检测

Importance Sampling and PCA for Finding Failures in Commercial Autonomous Vehicles

重要性抽样和PCA用于发现商用自动驾驶车辆故障

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

LLM作为教练:针对不可验证任务的体验式学习

Isaac Sim-to-Real: Reinforcement Learning based Locomotion for Quadrupeds

Isaac模拟到现实:基于强化学习的四足动物运动

Keyword: diffusion policy

HyperDCM: Dynamic Cluster Memory Replay in Hyperbolic Space for Continual Robotic Navigation Across Scenes

HyperDCM:双曲空间中的动态集群记忆回放,实现场景间持续机器人导航

Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion

通过延迟感知引导融合实现异步多模态扩散策略组合