生成时间: 2026-08-13 17:14:13 (UTC+8); Arxiv 发布时间: 2026-08-13 20:00 EDT (2026-08-14 08:00 UTC+8)

今天共有 31 篇相关文章

Keyword: reinforcement learning

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

回头交易员-工作台:用自生成的多选题对算法交易的LLM代理进行基准测试

Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization

基于强化学习的数据库管理系统缓冲池自动调优以实现内存利用率

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach

迈向在线教育中的可持续学习:强化学习方法

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

一瞥、审视与思考:将视频异常检测从无培训提升为智能推理

Dynamics Models for Offline Hyperparameter Selection in Real-World RL

现实现实强化学习中离线超参数选择的动力学模型

Self-Evolving Embodied Agents via Skill-Harness Evolution

通过技能驱动进化自我进化的具身代理

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

PAIR:RLVR中自适应推广分配的成对感知包含重权

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

通过反事实临床审计揭露ICU败血症管理中医疗离线强化学习中的有毒模仿

Benchmarking LLM Judges for Mobile Agent Evaluation

对移动代理评估的大型语言模型评判基准测试

Let it Cook: Learning to Wait in Sequential Decision Making

让它烹饪:学会在顺序决策中等待

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

强化大型语言模型中有效自我纠正的步骤级推理

IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework

物联网驱动的智能港口自主海事导航:课程指导共享政策学习框架

Learning from Online User Feedback for Shopping Agents

从在线用户反馈中学习购物代理

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

学会说服揭示了大型语言模型(LLM)多么容易放弃正确的信念

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

主张:领先的开放领域主动澄清大型语言模型的不确定性测量

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

每位代理人的保单构成安全吗?重新思考合作式多智能体强化学习中的继任特征转移

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

评分标准退出:在评分标准作为奖励的强化学习中,缓解奖励被篡改的简单方法

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

GCPO:在LLM的Rollout RL中诊断和约束亚空间几何

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

当API说错语言时:重新审视多语言工具使用的培训后

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

GRPO用于金融建议生成:在CATE评估中优于商业大型语言模型

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

HarmoniDPO:通过偏好优化扩散实现视频引导音频生成

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

LODESTAR:可信熵是被导航的,而不仅仅是测量——强化极化仪防止冻结的大型语言模型被错误的证据自信地误导

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

LoongReflect:通过全球视角提炼提升搜索代理的长远视野反射

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

重试、切换,还是弃权?通过受控错误注入学习战略感知工具使用策略

Token-Level Credit Assignment Optimization for Generative Document Retrieval

生成式文档检索的令牌级信用分配优化

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

从SMPC演示中学习机车操作,且离线到在线的强化学习稀疏

RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning

RoutePack:专家配置与注意力感知数据打包,用于 MoE 强化学习

Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation

基于机器学习的云基础设施网络防御:一种自适应深度Q网络架构,用于智能入侵检测和自动化威胁缓解

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

侦察兵:通过结构化思维链和多目标过程奖励解锁增强的空间推理能力

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

一个冻结的模拟器不够:多智能体强化学习中的模拟器崩溃

SelectLight: Learning to Select Signal Plans Generated by Distributed Model Predictive Control for Urban Traffic Networks

SelectLight:学习选择由分布式模型预测控制生成的城市交通网络信号计划

Keyword: diffusion policy

There is no result