生成时间: 2026-07-22 18:27:13 (UTC+8); Arxiv 发布时间: 2026-07-22 20:00 EDT (2026-07-23 08:00 UTC+8)

今天共有 38 篇相关文章

Keyword: reinforcement learning

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

S2T-RLHF:稳定优先RLHF的分层学分分配

Multi-Timescale Latent-Action DRL for Joint Optimization in Edge-Cloud Networks

多时间尺度的潜在作用持续DLL用于边缘云网络中的联合优化

Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal

深度强化学习以掌握巴格查尔的非对称策略

Decentralized Multi-agent Reinforcement Learning for Resilient Critical Infrastructures

针对韧性关键基础设施的去中心化多智能体强化学习

FARO: Feasibility-Aware Robot Motion Optimization

FARO:可行性感知机器人运动优化

Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion

迈向四足行走的扭矩驱动强化学习

Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability

部分可观察性下时间知识图记忆的神经符号元政策

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

RRPO:参考相对策略优化与分层条件推广

Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning

Search-on-Graph-R1:用强化学习训练大型语言模型搜索知识图谱

The Open Ant: A Robot Platform for Reinforcement Learning Research

开放蚂蚁:强化学习研究机器人平台

Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling

自动化数据工程与特征选择,用于熔融沉积建模中翘曲检测案例研究

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

连续状态-动作空间的网络多智能体强化学习可扩展策略优化

Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

在强化学习中以关系隐性状态为规划作为涌现行为

A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space

具有连续动作空间的协作任务的自我演化默认动作

Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach

ITNTN中的智能多无人机导航:一种层级大型语言模型方法

Exposure-Based Reinforcement Learning to Rank

基于暴露的强化学习以排名

Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents

策略跟随多智能体深度强化学习,考虑提供给其他智能体的控制策略

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

陈旧但稳定:稳定异步强化学习的陈旧自适应信任区域

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

从轨迹到指令:语言条件元强化学习

Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

动态环境中基于无人机的参与感测传递强化学习

Circuit Claims Depend on What Is Extracted and How It Is Compared

电路权利要求取决于提取的材料及其比较方式

H$^2$SD: Hybrid Hindsight Self-Distillation

H$^2$SD:混合后见之明自我蒸馏

Measuring Reward-Seeking via Contrastive Belief Updates

通过对比性信念更新测量追求奖励

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

捞出搭便车者:基于Shapley的奖励归因,通过强化学习实现平行推理

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio

雅典娜-大脑技术报告:一款用于通用智能和具身互动的高效机器人大脑

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

DobicVLM:通过群体相对政策优化,使胸部X光报告生成与临床基础的项目化奖励保持一致

Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation

采用具有可验证奖励的强化学习以促进分子生成

Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning

参数化动作强化学习中多智能体演员-批评算法的比较研究

Coherence in Control: Bridging Many-Core Mapping and Routing through Cost Unification

控制中的一致性:通过成本统一桥接多核映射与路由

Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning

翻译前的推理:用结构化推理增强法律机器翻译

Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

在不确定性估计下离线强化学习的保守查询与自适应正则化

Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

超越评分预测:基于LLM的论文评分与通过强化学习与评分标准奖励的反馈生成

The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation

推理的代价:神经机器翻译强化学习中的成本与质量权衡

S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning

S3:通过限制粗动态学的不确定性,在层级强化学习中实现稳定子目标选择

A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors

一种增强学习增强型液体燃料反应堆网络模型,用于预测燃气轮机的稀薄爆出

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

脱离上下文的GRPO:利用特权信息学习推理困难问题

ISO: An RLVR-Native Optimization Stack

ISO:RLVR 原生优化栈

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

OmniReasoner:通过原生工具使用长音视频思考

Keyword: diffusion policy

There is no result