生成时间: 2026-09-30 22:47:33 (UTC+8); Arxiv 发布时间: 2026-09-30 20:00 EDT (2026-10-01 08:00 UTC+8)

今天共有 100 篇相关文章

Keyword: reinforcement learning

Learning from the Gap Between Pass@K and Pass@1

从Pass@K与Pass@1之间的鸿沟中学习

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

OpenAI-HuggingFace:复刻版及对齐测试经验教训

IMPACT: Intent-driven Multi-agent Policy with Attention for SLO-guaranteed Microservice Migration in Cloud-edge Systems

影响:意图驱动的多智能体策略,关注云边缘系统中SLO保证的微服务迁移

From Static Policies to Adaptive Priors in Offline Reinforcement Learning

从静态策略到离线强化学习中的自适应先验

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

在经验时代自我发现强化学习:学习历史是优势还是负担?

Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion

受被动动力步行启发的动力学指导,用于节能类人机动

Question-Specific Knowledge Graphs for Efficient Visual Reasoning

针对问题的知识图谱,促进高效的视觉推理

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

ROSS:通过选择性监督重新学习自发的推广

PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning

PowerZooJax:基于JAX的强化学习动力系统基准测试

ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards

ABC:基于优势的控制变量用于强化学习,提供可验证的奖励

FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design

FLOORA:一个面向人类的领域特定语言模型,用于建筑设计

Dyad: Extending Large Language Models with Native Typed Decision-Making

Dyad:用原生类型决策扩展大型语言模型

Xiaomi-OCR-0 Technical Report

小米OCR-0技术报告

Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes

大-辅弱耦合马尔可夫决策过程中的公平策略优化

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

在代理强化学习中针对学分分配的关键决策

LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning

LeRF:学习参考坐标系以进行视角推理

Providing Rapid Design Feedback for 3D Obstacle Course Games Using Constrained Solvability Queries

利用受限可解性查询为3D障碍赛游戏提供快速设计反馈

ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning

ChronoSRL:自我监督强化学习的时间几何

Action Chunking Proximal Policy Optimization with Feedback Correction

带反馈修正的动作分块近端策略优化

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

在大型推理模型中缓解欺骗性安全对齐

Understanding LLM Parameter Update Sparsity through the Lens of Fisher

通过Fisher的视角理解LLM参数更新稀缺性

Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks

完全去中心化且安全意识的多智能体强化学习,用于网络控制

CheatBench: Measuring Reward Gaming in AI Agents

CheatBench:衡量AI代理中的奖励游戏

Massively Parallel Reinforcement Learning with a Chaotic Reconfigurable Clockless Chip

采用混沌可重构无时钟芯片的大规模并行强化学习

StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

StructRL:面向远景视觉-语言-行动任务的在线结构化强化学习

Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

高效机器学习工程代理的奖励率策略梯度

Sample Complexity of Equivariant Reinforcement Learning

等变强化学习的示例复杂度

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

教师是一个方向,而非终点:在策略上提炼中推算强化学习诱导的表征残差

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

桥梁:双级检索-学分感知能动强化学习

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

SERA:量表均衡推广分配以实现最大似然强化学习

Learning to Explore Hidden Kinematics for Articulated Object Manipulation

学习探索隐性运动学以实现关节物体操作

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

视觉敏感性不是主张可撤回性:多模态强化学习中的持久感知学分赋值

Learned Reporting Preferences in RLVR Can Conflict with the Current Request

RLVR中的学习报告偏好可能与当前请求冲突

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

通过强化精细调优实现的协作多智能体视觉-语言-行动模型

OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport

OTRetarget:通过最优传输实现联合机器人与物体运动的重新定向

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

CrossTimeEdit:一个跨十年的交叉视图数据集和基于奖励的历史街景生成编辑

Inducing Process Supervision from Outcome-Only Reinforcement Learning

从仅结果强化学习中诱导过程监督

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

PR-OPD:特权代表政策自我提炼用于能动强化学习

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

RankBuffer:高效的基于排名的奖励,用于开放式生成

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

SIPO:将强化学习与政策自提纯相结合

EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents

EASE:为自我进化智能体提供行为适应技能策划

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

群体边缘化的自我奖励强化学习驱动零标签自我进化

DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning

DSPO:多元感知的主观政策优化,以实现强健的情感推理

TaRL: Learning General and Physical Rewards from Tactile Demonstrations

TaRL:从触觉演示中学习一般性和身体上的奖励

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

GitHarness:git init 你的束带工作内存,支持永久用户需求

Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs

了解应听什么:诊断和修复全模态大型语言模型中的跨模态捷径

EasyPPO: Stabilizing the Critic Is Key

EasyPPO:稳定批评者是关键

VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction

VAA-CSEC:中文语义错误纠正的投票引导优势分配

Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows

利用提案条件精炼流程改进扩散政策

RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning

救援:通过强化学习修复语言模型错误至稀疏电路

Towards Better Training Signal: Advantage Clipped Policy Optimization

迈向更优的训练信号:优势剪裁策略优化

Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training

陈旧在哪里积累?池感知:LLM后培训中异步强化学习的有效陈旧控制

RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts

RoXDrive:通过忠实行动推广实现端到端自动驾驶的闭环强化学习

Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

用于强化学习的马尔可夫非凸ADMM:Bellman-解式稳定性在光滑块之外