生成时间: 2026-09-07 22:03:46 (UTC+8); Arxiv 发布时间: 2026-09-07 20:00 EDT (2026-09-08 08:00 UTC+8)

今天共有 26 篇相关文章

Keyword: reinforcement learning

Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

退一步迈向前进:视觉生成的反射感知偏好优化

Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

通过样本引导分布匹配实现视频生成的联合比对与蒸馏

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

VLA精度:非对称共自助,实现视觉-语言-行动模型的高效现实在线强化学习

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

精炼:基于预算文本属性图表的LLM细化,实现个性化医疗概念表示

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

HarvestBench:衡量大型语言模型代理是否愿意支付避免杀害动物的费用

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

多线束强化学习内容?编码代理中的学分分配与可移植性

Deep Reinforcement Learning for Optimization of STAR-RIS Phase and Energy Splitting Coefficients in OTFS-NOMA Framework

OTFS-NOMA 框架中 STAR-RIS 相位和能量分割系数优化深度强化学习

Extremely Sparse Supervision Incentivizes Reasoning Ability

极度稀疏的监督激励推理能力

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

ConsensusBench:通过结果奖励密度化实现LLM推理的共识节点基准

LookThere! Sparse Vision by Reinforced Selection

看那里!《稀疏视野》由强化选择制作

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

利用合成任务缩放训练大语言模型以实现小分子设计

Inventory-Grounded Policy-Level Optimization for Training-Free AI Search

基于库存的策略级优化,实现无训练AI搜索

Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

强化学习以提升大型语言模型的加泰罗尼亚语文本简化能力

Generating Constructive Feedback on Stories via Reinforcement Learning

通过强化学习生成对故事的建设性反馈

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

CoSkill:层级技能进化的推理与元技能代理的联合强化学习

Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

在不确定性下顺序太阳能光伏政策设计的强化学习:基于主体的方法

Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing

股票和加密市场中的人工智能:进展、盈利证据及自动化投资的局限性

Compositional Reward Models for Conditional Medical Image Generation

条件医学图像生成的合成奖励模型

Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG

超越未压缩的压缩:RAG软上下文压缩的两阶段训练配方

Morphology and actuation as inductive biases in robotic hand manipulation

形态学与驱动作为机器人手操作中的归纳偏置

First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

首先:教基于LLM的代理优先考虑必备功能,而不是“可拥有”的

GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

GUT:通过图复杂度量化和优化大型语言模型的推理不确定性

Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning

协作多智能体强化学习的在线变点检测

Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds

人机交互变压器(H2INT),用于在密集且不确定的人群中进行机器人导航

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

建筑能源系统中暖通空调操作的大型语言模型:方法、应用及部署准备度的批判性综述

Keyword: diffusion policy

Dressing in Motion: A Human Motion-Aware Diffusion Policy for Robot-Assisted Dressing

移动穿衣:机器人辅助穿衣的人机动作感知扩散政策