生成时间: 2026-09-10 20:46:34 (UTC+8); Arxiv 发布时间: 2026-09-10 20:00 EDT (2026-09-11 08:00 UTC+8)

今天共有 25 篇相关文章

Keyword: reinforcement learning

SMCC-Empowered Digital Twins for Sensorless Monitoring in Large-Scale AI-Driven IoT Systems

SMCC赋能的数字孪生,用于大规模AI驱动的物联网系统中的无传感器监测

Learning to Fly: Stable Vision-Guided UAV Servoing with Compact Target-Centric Cues and Reinforcement Learning

学习飞行:稳定视觉引导无人机服务,配备紧凑目标中心提示和强化学习

Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

视频-MOPD:多教师政策提炼以促进视频理解

From Learning to Control: Data-Driven Multi-Agent Reinforcement Learning for Multivariable Control in a Microalgae Bioprocess

从学习到控制:数据驱动的多智能体强化学习,用于微藻生物工艺中的多变量控制

Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management

跨连续体的智能自适应计算:物联网-边缘-云资源管理中的大型语言模型

A Decade of Bayesian Optimization for Controller Tuning and Robot Learning: Tutorial, Review, and Future Prospects

十年来的贝叶斯优化与控制器调校与机器人学习:教程、评测与未来展望

Actuator Dynamics Curricula for Narrow-Viability Tasks in Legged Robot Learning

执行器动力学课程,适用于腿式机器人学习中窄可行性任务

CityPlanner: A Sandbox Agent for Executable Urban Planning

城市规划师:可执行城市规划的沙盒代理

A Risk-Sensitive and Uncertainty-Aware Decision-Making and Control Framework for Safe and Robust Autonomous Driving

一个风险敏感且具不确定性的决策与控制框架,实现安全且稳健的自动驾驶

HiRAD: A Flexible Large-Scale AGV Routing System

HiRAD:一种灵活的大规模全地形车路由系统

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

SocialRL:通过多轮强化学习和奖励设计优化大型语言模型的社会智能

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

带证据的认知:用现实定型奖励弥合验证差距

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

BRACE:异步强化学习中为陈旧批评者锚定的贝尔曼残余修正

Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications

无人机安装的RIS辅助动态D2D通信决策变换器

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

HaWMPO:幻觉感知世界模型的通用机器人政策优化

VLX-VR: An Agentic-Aware Video Reasoning Model

VLX-VR:一种代理感知视频推理模型

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

以数据为中心的金融推理后期培训:挖掘、蒸馏与可验证学习

Assembling Two Parts in One Hand

一手组装两个部件

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

为什么要采样你能列举的?基因组工具选择的精确策略优化

Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search

通过策略引导嵌入搜索实现层级和置换不变特征转换学习

Learning Terrain-Adaptive Humanoid Locomotion on Granular Terrain

在颗粒地形上学习地形自适应类人运动

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

TRACE:训练推理代理进行因果探索,并以合成奖励进行

Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response

多智能体强化学习用于野火响应中的自主无人机探索

ConvMem: Convolutional Memory for Long-Context Reasoning

ConvMem:用于长上下文推理的卷积记忆

Keyword: diffusion policy

JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction

JEPA政策:通过配对动作和未来表示预测实现无扩散模仿学习