生成时间: 2026-09-04 20:34:54 (UTC+8); Arxiv 发布时间: 2026-09-04 20:00 EDT (2026-09-05 08:00 UTC+8)

今天共有 32 篇相关文章

Keyword: reinforcement learning

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

RL-ADA:针对对抗性强健企业对话代理的世界反馈框架

LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games

多智能体战斗游戏中自适应NPC行为的LLM引导强化学习

Tail-Likelihood Reinforcement Learning

尾似然强化学习

GPU-Accelerated Astrodynamics World Models for Spacecraft Rendezvous and Proximity Operations

用于航天器会合和近距离操作的GPU加速天体动力学世界模型

Adaptive Beam Hopping and Power Control for Dual-Layer Over-the-Air Online Federated Learning in LEO Satellite Networks

近地轨道卫星网络中双层空中在线联邦学习的自适应波束跳跃和功率控制

SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

SWIM:通过熟练条件生成实现的学生写作模拟

Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

长视界一致且交互感知的世界模型,用于多风格端到端驾驶

Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification

基于强化学习和不确定性量化,风险与异常识别以实现配电网络最优运行

TIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code Generation

TIPCODER:强化学习增强测试时间的代码生成指令提案

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

DE-Venus:大型语言模型的数据高效RLVR框架

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

梯度知道结果不了解的事:通过梯度对齐奖励解锁LLM推理的强化学习

StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

StrixAE:在复杂失真耦合下实现音频增强的智能代理

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO:消除扩散强化学习中的冗余重计算

A Semantic-Aware Multiple Access Scheme Leveraging Spatial Redundancy for Uplink-Dominant Network Services

一种利用空间冗余实现上行主导网络服务的语义感知多址方案

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

从先导启发式到可部署代理:加速演示驱动的强化学习以实现截止时间限制的网络控制

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

离线多智能体强化学习中的序列模型的分布外推广

LLM4AIGQ: LLM-based AI Guidance Query Generation Framework for Multi Interest Mining

LLM4AIGQ:基于LLM的AI引导查询生成框架,用于多兴趣挖矿

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

WISE:世界模型引导的想象调度,实现视觉-语言-行动模型的高效后期训练

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

展开世界:在强化空间推理中分解四维属性

SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

SVG-Score:文本到SVG生成的人类对齐评估

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

离线强化学习中的多步近端策略改进

EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments

EF1约束纳什社会福利与相同加法估值:复杂性、保证与实验

Revisiting Topological Graphs for Macro Action based Closed-loop Reinforcement Learning of Vision Language Navigation in Continuous Environment

在连续环境中视觉语言导航的宏观动作闭环强化学习中重新审视拓扑图

Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

代码大型语言模型中声音和对抗性测试生成的两阶段强化学习

FiMI Banking: A Sovereign Model for Indian Retail Banking

FiMI银行:印度零售银行的主权模式

The Dually Flat Geometry of Planning as Inference

规划的对偶平面几何作为推断

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

模型编辑过多时:关于最小代码编辑的忠实度

Spurious Advantage Hidden in GRPO

GRPO中隐藏的虚假优势

Subspace Inference Enables Efficient Active Reward Learning from Preferences

亚空间推理实现了从偏好中高效的主动奖励学习

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO:细粒度学分分配,动态评分标准,用于长视线特工培训

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

顺序节拍联合:论政策提炼与RLVR之间的相互作用

Keyword: diffusion policy

MulDP: Multimodal Diffusion Policy for Autonomous Quadruped Parkour Navigation across Complex Terrains

MulDP:跨复杂地形自主四足跑酷导航的多模态扩散政策