生成时间: 2026-07-30 18:26:43 (UTC+8); Arxiv 发布时间: 2026-07-30 20:00 EDT (2026-07-31 08:00 UTC+8)

今天共有 24 篇相关文章

Keyword: reinforcement learning

Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning

深度强化学习中的冻结随机CNN特征提取器中的涌现稀疏性

Guess Where You Go: Generative Next Point-of-Interest Recommendation in Amap

猜猜你去哪里:Amap 生成性下一个兴趣点推荐

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

元学习奖励塑造以实现人类反馈的强化学习

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

探究推理表现的起源:强化学习数学问题解决与SFT精细调优模型中的表征质量

ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter Networks

ProFlow:基于强化学习(RL)驱动且性能感知的主动流量部署,应用于数据中心网络

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

早期结论,更优预算:序列自适应部署分配以实现高效计算的RLVR

Learning Implicit Causal World Models from Multi-Agent Demonstrations

从多智能体演示中学习隐性因果世界模型

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

检测边缘的后训练:一种博弈论方法进行微调

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

CaM-Wolf:社会演绎游戏中的因果感知多模态代理

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

SCOUT:稀疏奖励强化学习的按情境重置课程

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

成本受限的四足硬件上的强化学习

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

DHRCL:训练代码LLMs,具有密集的层级奖励和课程学习

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

RLMM-Flow:基于流动的移动操作框架,具备潜在空间强化学习

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

与悲观批评者合作加权,以缓解非策略强化学习中的高估

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

视觉-语言-行动模型的显式运动指导分析概念

Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection

图是验证者:过程间漏洞检测中的代理强化学习

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

风险感知型自动实时学习的高效异方差贝叶斯优化

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

SkillRise:跨任务技能演进的能动强化学习

ReCo: Reweighting GRPO Against Distributional Concentration

ReCo:根据分布集中度重新加权GRPO

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

SERPO:自我演进的评分规策略优化,用于开放式测试时间强化学习

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

RL$^2$-VLA:视觉-语言-动作模型的自适应强化学习潜在组合引导,具备测试时间缩放

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

全息覆盖决策过程中通过稳定商实现的最小马尔可夫化

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

你真的需要预训练Q函数来进行在线强化学习微调吗?

Keyword: diffusion policy

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

以时间为中心的SIGReg提升多任务LeWorld模型学习:从分析到方法的改进