生成时间: 2026-07-27 19:23:53 (UTC+8); Arxiv 发布时间: 2026-07-27 20:00 EDT (2026-07-28 08:00 UTC+8)

今天共有 21 篇相关文章

Keyword: reinforcement learning

Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems

模块会保持其自身的定位吗?复合式大型语言模型系统中的角色漂移

Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning

元强化学习的准蒙特卡洛初始化

Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

调整速度作为非固定强化学习的安全约束

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Molt:一个可扩展的PyTorch原生智能强化学习训练框架

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

Oxygen-TryOn:时尚原生基础模型,适用于任何物品的虚拟试穿

Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

保持一致性!增强具有一致性约束的LVLM中强健的视觉推理

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

QLPO:象限加权抽样用于长度感知策略优化

Variance-Reduced Q-Learning over Static and Time-Varying Networks

静态和时间变化网络上的方差减少Q学习

Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

通过深度强化学习实现蛇状机器人在动态粘性环境中的自适应波动运动

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

教授大型语言模型自我进化:通过强化学习培养核心元技能

Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning

推理展开中的学习:高效强化学习的渐进推广分配

On the Runtime Analysis of Reinforcement Learning Hyper-Heuristics

关于强化学习超启发式的运行时分析

Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs

够了,就像盛宴一样:强化学习如何缓解大型语言模型任务冲突的综合分析

Predictive Lightweight MARL for Resilient Coverage in Sparse-Signaling Aerial Networks

预测轻量级MARL用于稀疏信号天线网络的韧性覆盖

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

拆解非策略比率:异步强化学习中的熵尺度信任区域

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR

LayoutLite:用于高效文档OCR的令牌级隐式布局分析

Active few-shot segmentation by reinforcing data selection

通过强化数据选择实现的主动少数样本分割

Conformal Constraint Tightening for Chance-Constrained Motion Planning with Unknown Dynamics

在未知动力学条件下的概率约束运动规划中的共形约束紧缩

Explainable Reinforcement Learning for assisting Air Traffic Controllers

可解释强化学习,协助空中交通管制员

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

技能自玩:通过共同进化技能推动LLM能力的前沿

Keyword: diffusion policy

GRACE: Gradient-Free Robot Action Generation via Combined Diffusion-MPPI Posterior Mean Estimation

GRACE:通过联合扩散-MPPI后验平均估计生成无梯度机器人动作