生成时间: 2026-09-15 21:16:34 (UTC+8); Arxiv 发布时间: 2026-09-15 20:00 EDT (2026-09-16 08:00 UTC+8)

今天共有 66 篇相关文章

Keyword: reinforcement learning

Harnessing human expertise for high-precision robotic assembly in industrialized construction: A sample-efficient installer-in-the-loop interactive reinforcement learning framework

利用人类专业知识实现工业化建筑中高精度机器人组装:一个高效的安装在环互动强化学习框架

Self-Evolving AI for Humanoids: Mechanisms, Safety, and Evaluation of Post-Deployment Self-Improvement

类人生物自我进化的人工智能:机制、安全性及部署后自我提升的评估

GzDRL: Reproducible and Scalable Deep Reinforcement Learning with Gazebo

GzDRL:带有凉亭的可重复且可扩展深度强化学习

An Evolutionary Computation Framework for Multi-Agent Q-Learning with Mean-Field Environmental Feedback

一个多智能体Q学习的进化计算框架,结合平均场环境反馈

Extending the Speed Limit of Quadrupedal Locomotion via Refined Actuator Modeling and Adaptive Command Scheduling

通过精细的执行器建模和自适应指令调度扩展四足行走的速度限制

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

利用图像问题依赖性以实现更好的VLM测试时间强化学习

Task-Based CT Protocol Optimization Using Reinforcement Learning and Virtual Imaging Trials

基于任务的CT协议优化,利用强化学习和虚拟成像试验

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

迈向自适应物理人工智能:LLM代理能否管理长期的物理任务?

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

TimeThink:在时间序列大型语言模型中激发组合推理

Constraint-Grounded Reinforcement Learning for Variable Impedance Control in Contact-Rich Robotic Insertion

约束接地强化学习用于接触富裕机器人插入中可变阻抗控制

Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control

基于强化学习的自适应控制运行时增量变换器

Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

规划或学习:多资产维护中的可靠性与成本

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

漂移约束优化:在微调指令模型中,只有方向才重要

Learning In-Hand Object Reaching to General 6D Poses

学习手持物体达到通用的6D姿势

Force-Aware Reinforcement Learning with Hybrid Sensorless Force Estimation for Wheeled-Legged Loco-Manipulation

基于混合无感应力估计的原力感知强化学习,用于轮式腿式机车操控

LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems

大型语言模型增强型多智能体强化学习,用于公共充电系统中统一电动汽车-充电站-电网优化

Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

绕过观察:一种非侵入式层级语义提取架构的概念设计

ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading

ViperQ:通过拍卖市场理论进行订单流模式识别,用于强化学习交易

UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

UniCAR-RL:在视觉数学中先看得更好再深入思考

North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)

North Small Translate:高级经济翻译(Cohere CAT+)

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

解锁无解之谜:教师引导课程,实现数据高效RLVR

SafePG: Safe and Globally Optimal Reinforcement Learning with Hard Constraints

SafePG:安全且全局最优的硬约束强化学习

DQN-Scheduler: A Multi-Objective Optimization Framework for Scheduling Microservices in Cloud Computing

DQN调度器:一个用于云计算微服务调度的多目标优化框架

VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching

VGFM:通过密集值指导实现表达式机器人策略,用于流量匹配

Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning

非参数方差惩罚的行为者-批评者:风险敏感强化学习的统计推断

Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors

仅使用三个测距传感器的无人机基于学习的动态障碍避让

EMoG: Emotion-Modulated Gait Generation for Expressive Humanoid Locomotion

EMoG:情感调制步态生成,用于表达性类人运动

Grounded in Sound: Reinforcement Learning with a Frozen Acoustic Judge to Curb ASR Insertion Hallucinations

扎根于声音:用冷冻声学法官进行强化学习以抑制ASR插入幻觉

Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots

工厂中学习多智能体任务分配与导航:从模拟到真实机器人

A note on goal-based hierarchical RL

关于基于目标的层级强化学习的说明

Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

通过价值引导偏好提炼,通过密集的行为信号优化稀疏结果

Comparative Evaluation of MILP, MPC, and Reinforcement Learning for Commercial Battery Dispatch Under Time-of-Use Tariffs

在分时使用费率下,商业电池调度中MILP、MPC和强化学习的比较评估

Multi-Agent Reinforcement Learning in Markets with Congestion

拥塞市场中的多智能体强化学习

Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation

利用MPC脚手架进行灵活操作的真实世界强化学习

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

《四十度蓝色:通过模态条件强化学习实现质量-多样性对齐》

Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement Learning

基于图的注意力驱动层级强化学习的云工作流调度

HiGFRL: Hierarchical Graph Fusion-Driven Reinforcement Learning for Dependency-Aware Task Scheduling in Heterogeneous Cloud

HiGFRL:用于异构云依赖感知任务调度的层级图融合驱动强化学习

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

学习如何求解未知漂移和运行奖励的随机控制:理论、算法与收敛

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

并非所有提示都相同:探索引导提示支架用于多模态培训后强化

What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

LLM从强化学习中学到什么?机械可解释性视角与固定SAE轨道

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Salesforce Koa:用于代理工具的企业语言模型

Refinement-based Flow Policy Optimization

基于精细化的流程策略优化

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip:自适应视觉令牌跨层跳跃,实现高效的MLLM推断

GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving

GRAVA:自动驾驶的基于推理到行动的表征与学习

Admissable: Training Reinforcement Learning Agents against Adversarial Missingness

可采纳:培训强化学习代理以防对抗性缺失

Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding

从可靠负面中学习:基于信心的测试时间适配以实现GUI接地

Evaluation Metrics for Safe Reinforcement Learning

安全强化学习的评估指标

Robust and Efficient Communication for Multi-Agent Learning

多智能体学习的稳健高效通信

CodeTS: Verifiable Text-to-Time Series Generation via Executable Code

CodeTS:通过可执行代码生成可验证的文本到时间序列

Dynamics-Informed Reinforcement Learning for Agile and Energy-Efficient Locomotion of a Monopedal Hopping Quadcopter

动力学知情强化学习,实现单足跳跃四旋翼的敏捷且节能的行走

HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments

HISPO:层级重要性-抽样策略优化,含熵衍生片段

Strong and Compact Policies for Submodular Markov Decision Processes via LP-Based Submodular Orienteering

通过基于LP的次模块定向,为亚模块化马尔可夫决策过程制定强而紧凑的策略

Specifying Reward Functions for RL Without Environment Sampling

在不使用环境抽样的情况下指定强化学习的奖励函数

VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding

VideoScout:学习能动主动探索,采用自适应推理节奏,实现长视频理解

AI-Native Open RAN: A Roadmap from xApps and rApps to Autonomous Network Agents

AI原生开放RAN:从xApp和rApp到自主网络代理的路线图

Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation

组装CREW:一个用于自动化相关工作生成的协作多智能体强化学习框架

JEPLO: Joint-Embedding Predictive Learning for LiDAR-Based Legged Locomotion

JEPLO:基于激光雷达的关节嵌入预测学习

Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation

应对稀疏证据:通过显式上下文选择与巩固实现能动视觉RAG

Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

通过价值加权最优运输学习多模态单步流策略

Inoculation Midtraining with Learned Neologisms

学习新词的培训中接种

Safe Meta-Reinforcement Learning via Information Space Reachability

通过信息空间可达性实现安全的元强化学习

Bellman Policy Optimization

贝尔曼策略优化

ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids

ResSafe:人形生物的残余强化学习安全过滤

Keyword: diffusion policy

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Attention-DP3:通过几何对齐注意力条件的空间感知对象3D扩散策略

LieSpline-DP: Lie-Group B-Spline Diffusion Policy for Smooth Robot Manipulation

LieSpline-DP:用于平滑机器人操作的李群B样条扩散策略

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Bench2Dex:双手灵活操作的Visuo-触觉基调