生成时间: 2026-08-29 03:58:07 (UTC+8); Arxiv 发布时间: 2026-08-28 20:00 EDT (2026-08-29 08:00 UTC+8)

今天共有 29 篇相关文章

Keyword: reinforcement learning

The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

人工实验者:通过自我强化学习发现与控制自组织现象

Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

奖励知情的稀疏自编码器与解完全性混淆

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

AdaThinking-e:适应性思维的单代币熵调控

CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models

CARE:医学大型语言模型的因果对齐推理探索

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

用潜在世界模型预测后果并强化导航政策

AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

AffectOmni:现实学习可验证的以人为中心的基于社会和艺术场景的情感推理

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

SpeechGym:一个通过强化学习培训语音代理的音频原生健身房

Active Curriculum Refinement for Reinforcement Learning

强化学习的主动课程精炼

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

共享行为者不必分享批评者:并行强化学习中价值不匹配的影响

Video-FLAIR: Not Whether to Reason, But How

视频标签:不是讲究是否讲道理,而是怎么讲

SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning

SPEAR:通过顺序符号对齐在强化学习中提炼领域自适应推理骨架

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

到达并生存:从一比特故障信号中学习安全目标条件策略的扩展

Simple Actors and Deep Critics for Scalable Reinforcement Learning

简单演员与深度批评者,适用于可扩展强化学习

SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

SIGMA:结构化噪声效应感知分组多智能体聚合

Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

你说的话中的理性:视频大型语言模型中推理提炼的非政策痕迹的意语直译

Residual Deep Reinforcement Learning-Based Computed Torque Control for a Cable-Driven Lower-Limb Rehabilitation Robot under Disturbances and Parametric Uncertainties

基于基于残余深度强化学习的计算扭矩控制,适用于受干扰和参数不确定性条件下的电缆驱动下肢康复机器人

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Behavior2Trip:通过用户行为轨迹实现个性化旅行规划

Towards Safe Reinforcement Learning with Reduced Conservativeness: A Case Study on Drone Flight Control

迈向降低保守性的安全强化学习:关于无人机飞行控制的案例研究

Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic

基于强化学习的CAV排在混合交通中加入机动的控制

AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion

人工智能代理在算法电力市场中的出现:关于默契合谋的出现

Energy-Neutral Coverage Optimization by Joint Deployment and Scheduling in Ambient IoT Devices with Directional Sensing

通过联合部署和调度环境物联网设备实现的能量中和覆盖优化,并具备方向感测

Performance Foundations of Parallel & Distributed Reasoning Language Models

并行与分布式推理语言模型的性能基础

Emotional Preferences as Goal-Priority Regulation

情感偏好作为目标优先级调节

GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

GRAIN:通过不变性奖励的能动强化学习,弥合现实世界图论推理中的名称与叙事转变

A Trans-Domain Digital Twin for Bio-Aware Control of Climate and Energy in Cattle Fattening Barns Using Single-Episode Optimizer Learning

跨域数字孪生,利用单集优化器学习实现牛场中气候和能源的生物感知控制

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

跨领域整合RLVR能力:融合范式深度解析

Boosting LLM Exploration via Weak-Model Guidance in RLVR

通过RLVR中的弱模型指导促进LLM探索

TTPO: Test-Time Policy Optimization

TTPO:测试时策略优化

Keyword: diffusion policy

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

机器人人群导航短期规划的扩散政策