生成时间: 2026-07-28 18:39:20 (UTC+8); Arxiv 发布时间: 2026-07-28 20:00 EDT (2026-07-29 08:00 UTC+8)

今天共有 58 篇相关文章

Keyword: reinforcement learning

CRAFT: Learn the Schema, Execute the Plan

技巧:学习模式,执行计划

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

DocHRL:一种用于成本优化文档分类的分层强化学习框架

STAIF: A Stage-wise Optimization for Complex Instruction Following

STAIF:复杂指令跟随的分阶段优化

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

LLM任务适应如何重塑对齐:行为与表征漂移的多维研究

LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory

LazyMem:广泛检索,选择性构造,以实现高效的长期代理记忆

HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation

HiLLTS:零射分层LLM引导交通信号控制,用于可持续交通

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

Cortex:带冻结视觉特征的Quake紧凑行为克隆

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

令人沮丧的简单语言模型黑箱式通过Logit偏置

A Replay-Constrained Simulation Framework for Personalization of Powered Knee--Ankle Prosthesis Controllers

一个用于个性化动力膝踝假肢控制器的回放约束模拟框架

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

空间智商:通过层级能力测试解构空间智能

Not All LLM Reasoning is Visible in the Chain-of-Thought

并非所有LLM推理都能在思维链中显现

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

有限视界马尔可夫决策过程中自然政策梯度的有限时间分析

Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics

通过对抗性逆强化学习实现无标签工业故障检测:一个运行到故障预报系统

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

Real2Sim2Real 用于视觉-语言-行动操作操作:基于 AMD ROCm 的流水线

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

针对具有动态UBSR措施的MDP在线政策评估

DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

DishSeg24k:基于随机专家解码的大规模食品分段基准

Traceable LLM Reasoning for Fake-Order Fraud Detection

可追溯的大型语言模型推理用于假订单欺诈检测

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

AgentOmnia:全场景应用的智能模型尺度化

Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation

带有噪音学生政策自我提炼的自我提升视觉语言模型

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

SeekJudge:计算机智能体强化学习的实用奖励框架

On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

关于无偏且长度不变的策略优化与结果奖励的不可能性

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

LA-RL:信息提取中强化学习的标签感知自我反思

Learning to Optimize: Joint Routing and Flow Allocation on Sparse Non-Euclidean Networks

学习优化:稀疏非欧几里得网络上的联合路由与流量分配

PRISM: Polynomial Representations for Interaction-Structured Motor Control

棱镜:交互结构运动控制的多项式表示

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

通过哈达玛德超参数化实现稀疏高斯混合模型Q函数的在线强化学习

Learning Sampling Parameters for Diffusion Models

学习扩散模型的采样参数

LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks

LEACL:LLM增强型自动课程学习,用于长期视野操作任务中的强化学习

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

ObsDriveBench:在恶劣天气下以可观测性意识进行多模态理解的基准测试

Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter

前瞻性风险引导强化学习,助力在动态杂波中安全飞行

Optimal Reward Shaping: Autonomous Car Parking Case Study

最佳奖励塑造:自动停车案例研究

Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

突破总方差障碍:带有固定动作集的线性异方差盗垒的锐利样本复杂度

Offline-Online Curriculum RL for Multimodal Reasoning

多模态推理的离线在线课程强化学习

Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

稀疏奖励长视界强化学习的层级软性演员-批评者

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

TRUAV:分布式多智能体强化学习,用于无人机辅助物联网VANET中的轨迹规划与路由增强

Zing: Social Mind for LLMs

Zing:面向大型语言模型的社交思维

Training Language Models to Cooperate with Inference-Time Controllers

训练语言模型以配合推理时间控制器

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

$N_0$-VTLA:带潜在触觉令牌的视野-触觉语言-动作模型缩放

GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks

基于GNN的多代理控制稀疏车辆自组网络中的流量冲击波

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

PathScale-R1:病理图像分析的跨尺度推理

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

从RLVR到RLSVR:任务转换为开放式LLM自我提升带来自我验证的奖励

Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

通过学习与搜索理解组合优化中的类人解

Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping

通过可行的动作映射桥接强化学习与最优控制

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

EviBack:通过证据受限教师退让进行搜索-代理强化学习

WARL: Wrench-Augmented Reinforcement Learning for Task-Agnostic Learning in Legged Robots

WARL:用于腿部机器人任务无关学习的扳手增强强化学习

Constrained Reinforcement Learning Using Successor Representations

使用继任表示的约束强化学习

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning

ACRL:稳定强化学习中训练-推断差异的自适应控制

Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation

学习可重复使用的混合运动先验,用于模拟运动中的类人移动

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

MemChain:为内存增强的LLM代理学习可解释的记忆痕迹

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

监管推理:交通规则理解的思路链

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

从专有到开源:通过多智能体协议提炼弥合代理搜索中的分发差距

Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks

通过超网络学习自适应多任务引导、导航和控制

CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models

CONSISTRE:一个统一一致性感知框架,用于大型语言模型的文档级关系提取

Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

Continual-RL 用于 RoboRacer 平台上的自动竞速推广

VulnGym: Evaluating Vulnerability Management Strategies against Advanced Persistent Threats

VulnGym:评估针对高级持续威胁的漏洞管理策略

Evaluating Fuzz Testing for Reinforcement Learning Agents

评估强化学习代理的模糊测试

Kimi K3: Open Frontier Intelligence

Kimi K3:开放边境情报

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

通过物理感知策略蒸馏的可解释强化学习

Keyword: diffusion policy

PRISM: Polynomial Representations for Interaction-Structured Motor Control

棱镜:交互结构运动控制的多项式表示