生成时间: 2026-08-18 16:39:46 (UTC+8); Arxiv 发布时间: 2026-08-18 20:00 EDT (2026-08-19 08:00 UTC+8)

今天共有 57 篇相关文章

Keyword: reinforcement learning

When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL

何时沟通:多智能体强化学习中原则门控的信念分布与认知分歧

SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization

SKILL:自纠正知识引导迭代大型语言模型代理,用于逻辑优化

Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach

城市无线网络中的智能基站部署:一种基于地理数据的数字孪生方法

Explaining Reinforcement Learning Decisions in Self-adaptive Systems

解释自我适应系统中的强化学习决策

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

大型语言模型中对抗性政治偏见的推理时间缓解

Belayer: Efficient Fault Tolerance for LLM Agentic RL Training

边路者:LLM代理强化学习训练的高效容错

Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions

在每集分布中培训和评估伦理强化学习代理

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

通过离线强化学习发现高质量的国际象棋谜题

Deep Reinforcement Learning for 6G AI-RAN: A Comprehensive Survey

6G AI-RAN深度强化学习:一项全面调查

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

指令空间反事实解释帕累托条件强化学习

MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems

元理性:通过编辑元信息解决几何问题,精确交错多模态推理

LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning

基于LLM的层级协调控制与持续感知策略学习

Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning

Max-Q 人类在线机器人学习的选择性模仿

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

StructRL:基于流的VLA的结构化动作空间探索

PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search

PureTD:无评估时间搜索的双陆棋奖金游戏强化学习

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

LAPF:基于 UAVScenes 数据集的基于 LLM 代理的路径查找器

Temporal Logic Guided Universal Task Representations for Reinforcement Learning

强化学习的时序逻辑引导通用任务表示

Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation

现在谁来领导?用于图表到代码生成的令牌级模态仲裁

GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning

GLaQ:将潜在查询建立在视觉证据中,用于多模态推理

Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback

为什么摘要变得中立:强化学习中从人类反馈中归因的政策归因

When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction

当熵不足以实现:在LLM输出长度预测中重新夺回失去的语义

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

机器人多巴胺2.0:基于历史条件和值班时间感知的过程奖励建模用于机器人操作

Adaptive Mixing of Policies from Searching and Policies from Learning

从搜索和学习策略的自适应混合

GAINS: Leveraging Inconsistent Human Intervention Signals in Reinforcement Learning

收益:利用不一致的人类干预信号在强化学习中

TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness

TaoLive 数字化身代理技术报告:培训代理与其安全带同步进化

Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning

通过基于重中心的对抗性逆强化学习学习股票交易策略

Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading

股票交易中稳定强化学习的自监督辅助任务发现

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

询问确认:信息丰富的互动,助你自信地推荐多回合大型语言模型

DER Allocation without Load Prediction via Reinforcement Learning

通过强化学习实现无负载预测的DER分配

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

学习剩下的,而非已掌握的:多奖励政策优化中的饱和优势重权重

US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina

US-VLA:一种针对具身腹部的超声视觉-语言-动作模型

Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System

图神经辅助演员-批判者,用于延迟高效的边缘视觉系统

TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

TRCA:长视野LLM代理的过渡性评分标准学分分配

Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics

通过受控自助和受控价值动态理解并稳定深度Q学习

RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing

RoboStriker:自动人形拳击的潜在空间战略游戏

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

BaT:迈向具有阶段评分标准的自我进化医学研究代理

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

Defake-o3:从推测性理由到可验证的AIGI检测证据

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

TransAnyText:通过结构化视觉生成翻译电商图片中的任意文本

PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster

海报文本:迈向电子商务海报统一的视觉文本生成与编辑

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

StreamOPD:带有时空提示门控的培训后配方,用于流媒体视频理解

KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning

KC-BFPRL:通过双级基于前指点的强化学习实现知识引导多无人机草地恢复协作

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

PertMind:通过基于细胞扰动数据的强化学习,激发LLM中涌现的生物推理

Proving the Utility of Large Language Models in Cybersecurity Simulations: A Comprehensive Examination

验证大型语言模型在网络安全模拟中的实用性:全面考察

Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics

通过不确定性引导混合动力学实现的稳定多步推广

Drive, Pack, Fly: The Travelling Thief Problem with Drone

驾驶、打包、飞行:无人机的旅行小偷问题

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

ICU败血症血流动动力学管理的离线强化学习:结合双重离场策略评估的MIMIC-IV研究

FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning

FLEET:基于令牌的特征提取,用于基于事件摄像机的强化学习

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

问、条件还是弃权:缺失前提推理的强化学习

Interactive Whole Slide Images for RL-based Tumour Segmentation

用于基于强化学习的肿瘤切割的交互式全片图像

A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory

基于RFID支持的智能工厂的车间生产排程案例

Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents

Chronocooked:强化学习代理隐式区间时序基准

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

Le Critique:LLM强化学习中的特权价值函数

ClawGym II: Exploring Black-Box RL on Agent Harness

ClawGym II:探索黑盒强化游戏中的特工安全带

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

HAF:通过层级动作流和光谱潜在强化学习,将通用VLA适应为类人生物全身机动操作

Q-based Variational Inverse Reinforcement Learning

基于Q的变分逆强化学习

Keyword: diffusion policy

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

计划者条件扩散用于协调多智能体探索

MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning

MatchingPolicy:对应感知策略支持跨对象上下文学习