生成时间: 2026-09-09 20:50:15 (UTC+8); Arxiv 发布时间: 2026-09-09 20:00 EDT (2026-09-10 08:00 UTC+8)

今天共有 72 篇相关文章

Keyword: reinforcement learning

Benchmarking Storage Systems for Machine Learning Workloads Using NIO Bench

利用NIO Bench进行机器学习工作负载的存储系统基准测试

Compiling VGDL into Causal Models

将VGDL编译成因果模型

Information-Guided Safe Reinforcement Learning for Autonomous Gas Source Localization using sUAS

利用sUAS实现自主气体源定位的信息引导安全强化学习

EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

EnvCraft:在Agentic RL中为爪状代理综合可执行环境

Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

内生探索基于内在好奇心的强化学习

Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling

牛顿匹配用于生成建模:一个用于微调和抽样的统一框架

Inference-Time Graph Engineering for Multi-Agent LLM Workflows

多智能体LLM工作流的推理时间图工程

Generalizing HVAC Control With Domain Randomized Reinforcement Learning

将HVAC控制推广到领域随机强化学习

SLA-Safe Energy Control for AI-Native NG-RAN Using Stability-Aware Constrained PPO

基于AI原生网络的SLA-Safe能源控制,采用稳定性感知受限PPO

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

UniRRM:跨语言和评估范式的统一推理奖励模型

Test-Time Weak-to-Strong Alignment: Transferring Implicit Rewards from Weak to Strong Flow Models

测试时间弱到强对齐:从弱流向强流模型转移隐性奖励

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

如何从人类会避免的事情中学习?干预意识世界模型与现实现实强化学习,实现灵巧操作

FALCON-S: Fixed-wing ground-effect Aerodynamics Simulator and Flight Control Learning Suite

FALCON-S:固定翼地面效应空气动力学模拟器和飞行控制学习套件

SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking

SCoCaT:成功条件条件约束强化学习以助航天器对接

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

DataFlex-RL:RLVR数据政策评估平台

Fixed-Time Integral Reinforcement Learning for Saturated Nonlinear Multi-Agent Systems Under FDI Attacks

在外商直接投资攻击下,针对饱和非线性多智能体系统的固定时间积分强化学习

Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning

非平稳强化学习中的谱级优先扫描

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

MobileVLA-R1 2.0:增强学习驱动的移动机器人控制推理

Geometric Distributional Control: Learning Progress with Partial Structural Knowledge

几何分布控制:部分结构知识下的学习进展

Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts

验证器错误在GRPO组内是独立的吗?Qwen2.5推广的证据

Local and Global Stability in Performative Reinforcement Learning

执行强化学习中的局部与全局稳定性

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

一步一导线:通过跨步控制缓解多域强化学习中的高阶干扰

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

用摄像头思考:通过动态视点控制实现监控视频理解的主动视觉推理

Certifying cooperation: a novel approach to cooperative multi-agent task generation

认证合作:合作多智能体任务生成的一种新方法

Inducing Emergent Misalignment from Reward Hacks with Iterative DPO

通过迭代DPO诱发奖励黑客中的涌现错位

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

PARSER:并行阅读,长上下文LLM代理的深度解析

SkillX: Unified Multi-Skill Policy Learning for Humanoid Soccer

SkillX:人形足球统一多技能政策学习

Agentic Visual Generation: From Generative Models to Agentic Control

智能视觉生成:从生成模型到智能控制

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

离线强化学习中扩散策略的噪声空间策略梯度

Distributed Secure Learning Control for Large-scale Multirobots under Stealthy Actuator Attacks

针对大规模多机器人在隐形执行器攻击下的分布式安全学习控制

CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation

CARDEA:基于空间证据的可听性推理,用于端到端冠状动脉造影解读

Mind the Phase: Effective Rank and Representation Health in Legged Locomotion

注意阶段:腿行走中有效的等级与表现健康

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

TrojanWorld:通过想象力引导后门世界模型代理

SupGRPO: Enhancing GRPO with Matching-based Online SFT for Text Spotting

SupGRPO:通过基于匹配的在线SFT增强GRPO文本检测

Revisiting Complete Reasoning Traces for Post-Training

重新审视完整的推理痕迹以备培训后

Beyond Sparse Rewards: A New Benchmark and Structure-Aware Graph Alignment for Micro-Drama Understanding

超越稀疏奖励:微剧理解的新基准与结构感知图对齐

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

大规模、长上下文强化学习后期推测解码的在线草稿共训

From LLM-Generated Specifications to Learned Quadruped Locomotion

从LLM生成的规范到学习的四足行走

Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer

Flow3D-OPD:多教师策略蒸馏,用于使用流量匹配扩散变换器生成三维几何

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

Stable-MM-R1:通过熵引导分层锚定多模推理动力学

CircuitLens: Reasoning Circuits as Data Selection Signals for Reinforcement Learning with Verifiable Rewards

CircuitLens:推理回路作为强化学习数据选择信号,可验证奖励

Phase-and-First-Arrival VLM Feedback for Sparse-Reward Reinforcement Learning in Surgical Manipulation

手术手法中稀疏奖励强化学习的阶段与首次到达VLM反馈

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

通过渐进点匹配实现的长视野语言模型强化学习

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

通过自动机学习进行时间因果推断以实现强化学习

Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining

通过本体感觉锚定的跨模态预训练实现零射点模拟到真实接触丰富组装

HyCO: A Hybrid Neural Solver for Combinatorial Optimization

HyCO:一种用于组合优化的混合神经求解器

Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study

迷你批次风险规避深度Q学习:机器人导航案例研究

mjorbit: A Simulation Framework for Space Robotics

mjorbit:空间机器人模拟框架

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

通过强化学习引发多模推理代理的自我验证

Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning

非平稳强化学习的主动上下文预测安全约束

When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation

当指标奖励最差翻译时:内化文化推理以评估社交媒体翻译

QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement Learning

QoS感知的RACH前言切片,通过配额投影分支深度强化学习

A Better Spur Should Start From Each Objective

每个目标都应该有一个更好的推动

Bridging Language and Physics: Automated Design of Continuum Robots with Large Language Models

连接语言与物理:利用大型语言模型自动设计连续介质机器人

Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTM

利用历史感知离线强化学习(LSTM)路由密集布局

From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment

从一眼到细致:渐进畸变推理用于细粒度图像质量评估

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

环境作为支架:在长期任务中自助反馈自我演化代理

SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs

SRPO:多代理大型语言模型的集合相对策略优化

Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models

哪些形式的照护者反馈支持语法学习?儿童类语言模型的强化学习研究

SUN: Reaching for Novelty in Reinforcement Learning

SUN:在强化学习中追求新颖性

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

难度自适应树结构策略优化以扩展RLVR推理覆盖范围

Learning to build covering structures with continuous adjustments

学习通过持续调整构建覆盖结构

Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology

基于图的多智能体系统安全强化学习,具有时间变化拓扑结构

CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

CAST:基于模型的强化学习中交替的状态值目标与扩展策略梯度

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

AuK 技术报告:语音生成与编辑的开源基础模型

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

PlayTrain:一个高效的强化学习框架,用于LLM生成的可适应JavaScript游戏

ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR

ThinkPrior:RLVR 冷启动提示选择的零启动难度先验

ExecCritic: Learn to Test, Test to Improve for Coding Agents

执行批评:学会测试,测试以改进编码代理

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

熵正则化秩掩盖策略优化,用于代码生成中的测试时强化学习

Keyword: diffusion policy

MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference

MemCorr-DP:以参考为导引的扩散政策的反事实对应条件

CAVEAT: Recurrent Multimodal Diffusion Planning for Mapless Aerial Exploration

注意:无图空中探测的重复多模态扩散规划

M3-Tele: A Unified Multimodal Teleoperational Framework for Compliant Whole-Body Mobile Manipulation

M3-Tele:一个统一的多模态远程操作框架,用于符合合规的全身移动操作