生成时间: 2026-07-24 18:20:42 (UTC+8); Arxiv 发布时间: 2026-07-24 20:00 EDT (2026-07-25 08:00 UTC+8)

今天共有 28 篇相关文章

Keyword: reinforcement learning

AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics

AINTMA:具备生成智能、安全云通信和自适应质量分析的自主测试管理智能人工智能架构

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

基于不一致的人类反馈实现可靠性感知的LLM对齐

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

当RLVR缩小推理边界:诊断反转Pass@k

Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

利用多智能体强化学习在降级监控下解决空中走廊的冲突

Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design

超越SBDD:多药理学与多靶点药物设计中的几何深度学习

CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning

CMI-Mem:通过CMI增强强化学习实现可推广的长期记忆管理

SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

SalesLoop:通过绩效反馈进行强化学习,以获得销售线索排名

Adaptive Multi-Horizon Reinforcement Learning

自适应多视界强化学习

Learning to Detect UI Principle Violations via Reinforcement Learning

通过强化学习学习检测UI原则违规

Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

透视潜在作为主动推理代理因果出现的建筑条件

Robostral Navigate

机器人导航

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

在奖励和状态损坏下的稳健异步Q-学习,通过批处理实现

Offline RL with Hierarchical Action Chunking

带有层级动作分块的离线强化学习

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

多回合强化学习,具结构性和性能感知奖励,用于CUDA内核生成

The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning

沉默的重量:潜在国际象棋推理中权重在草稿板上的因果论证

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

EmoAgent-R1:迈向基于强化学习的动态智能体专化的多模态情感理解

From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python

从评估到优化:Python 中基于层级感知的训练信号用于 CWE 预测

Training Large Language Models for Self-Explanation Faithfulness

训练大型语言模型以实现自我解释的忠实性

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

巡演:离线强化学习的轨迹级去学习基准

Relative Value Learning

相对价值学习

Deep Reinforcement Learning for Adaptive Gain Tuning in Control of Teleoperation Manipulators with Joint Flexibility and Time-Varying Delays

深度强化学习用于控制具关节灵活性和时变延迟的远程操作机械臂中的自适应增益调谐

FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

FORGE-plus:带冻结LLM Supervisor的强制预算恢复,用于接触丰富组装

Expert Behavior Prior Reinforcement Learning

专家行为:事前强化学习

How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning

适配器能写多少位?参数高效微调的容量与记忆度的测量

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

PATS:策略意识培训支架用于代理强化学习

AREX: Towards a Recursively Self-Improving Agent for Deep Research

AREX:迈向一个递归自我改进的深度研究代理

Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections

无信号灯路口自动驾驶车辆的紧凑潜在协调

MIRROR: Learning from the Other View for Multi-Modal Reasoning

镜像:从他者视角学习多模态推理

Keyword: diffusion policy

There is no result