生成时间: 2026-07-20 18:54:07 (UTC+8); Arxiv 发布时间: 2026-07-20 20:00 EDT (2026-07-21 08:00 UTC+8)

今天共有 22 篇相关文章

Keyword: reinforcement learning

Co-Design of Aeroelastic Systems with Deep Reinforcement Learning

与深度强化学习的气动弹性系统共同设计

Robust Peak-cost Constrained Reinforcement Learning

稳健的峰值-成本约束强化学习

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

从黑匣子到可执行逻辑:通过Prolog专家系统实现的可解释强化学习

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

流程奖励知情树推广以实现有效的多回合强化学习

Difference-Based Relational Learning for Zero-Shot Object-Goal Visual Navigation With Direct Sim-to-Real Transfer

基于差分的关系学习用于零射程目标视觉导航,实现直接模拟到现实的传输

PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval

PCTD:代理工具检索中的偏好引导反事实任务分解

GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking

GoStop:事件型特征追踪中自适应时间聚合的强化学习

RAVEN: Reinforcement-Adaptive Visibility-Graph Planning for Robust Humanoid Navigation with Collision-Free MPC

RAVEN:强化-自适应可视化-图规划,实现无碰撞MPC的稳健类人形导航

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

通过隐性文化对齐奖励建模去偏见文本到图像评估

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

模块化动态粒度视频LLM用于多事件长视频理解

QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

QUAD:通过跨双侧的质量化-错误对齐稳定NVFP4强化学习,实现MoE中的NVFP4强化学习

Scalable Supervisory HVAC Control for Linear Objectives

线性物镜的可扩展监督暖通空调控制

DSWorld: A Data Science World Model for Efficient Autonomous Agents

DSWorld:高效自主智能体的数据科学世界模型

Learning Reach-Avoid Task with Reinforcement Learning: Vectorized Simulation and Benchmark

通过强化学习学习实现距离-避免任务:矢量化模拟与基准测试

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

少用更多:一个用简单配方打造的大型遥感VLM

Data and Learning Where it Matters for Contact-Rich Manipulation

数据与学习对丰富接触操作至关重要

CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach

CLaC@FinMMEval 2026 任务3:情感增强深度强化学习用于主动交易——一种阿尔法奖励方法

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

JoyNexus:面向服务的多租户VLA模型后培训

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

DADiff:扩散驱动的跨域策略适配用于强化学习

Understanding Reasoning from Pretraining to Post-Training

从预培训到培训后理解推理

Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

物理增强强化学习用于动态系统实时最优控制

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

以树状搜索视频:基于长视频质量保证的自我纠正代理

Keyword: diffusion policy

There is no result