生成时间: 2026-09-21 22:55:50 (UTC+8); Arxiv 发布时间: 2026-09-21 20:00 EDT (2026-09-22 08:00 UTC+8)

今天共有 36 篇相关文章

Keyword: reinforcement learning

COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training

COAL-SQL:覆盖引导增强与失败驱动的文本转SQL后培训学习

Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces

通过自发的深度研究推广,提升深度研究和长上下文的能力

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

BI-Agent与BI-Bench:迈向自动化端到端商业智能

Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

连续延迟记忆随机梯度下降与连续时间强化学习,源自天体物理时间序列研究的历史

Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

高效贝叶斯自适应强化学习与时间逻辑规范

ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

阿斯加德:通过强化学习实现无人机韧性行动空间守卫

DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real

DEXTERA:从单一图像到可部署的灵活操作,通过真实到模拟再到真实

Design of Adaptive PID Controller Based On Asynchronous Advantage Actor Critic Learning Method for QuadCopter Control

基于异步优势演员批评者学习方法设计的自适应PID控制器,用于四旋翼控制

Dynamics-Induced Commitment in Learning-Based Robotic Penalty Kicks

基于学习的机器人点球中的动力学诱导承诺

REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

REFINEPPO:通过迭代动作精炼学习持续控制策略

MetaPusher: Meta Learning and Planning for Nonprehensile Manipulation of Unseen Objects with Rapid Online Adaption

MetaPusher:基于在线快速适应,实现非抓握式操作的元学习与规划

SAGE: Safety-Aligned Gradient Enforcement for Human--Robot Collaboration

SAGE:人机协作的安全对齐梯度执行

Stability-aware Residual Reinforcement Learning Framework for Robotic Manipulator Disturbance Compensation

机器人机械臂干扰补偿的稳定性感知残差强化学习框架

LEMCA: LLM-Guided Synthesis of Efficient Mode-Switching Control Architectures

LEMCA:高效模式切换控制架构的LLM引导综合

Deep Reinforcement Learning with Buffered Quantile Objectives

带缓冲分位数目标的深度强化学习

Co-Evolving Zero-Day Jamming: Adaptive Attack Synthesis and Graph Attention-Based Online Detection

共同演进的零日干扰:自适应攻击综合与基于注意力的图在线检测

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

ArenaFlow:从轨迹排名到开放式智能体强化学习的层级信用传播

DPed-VLN: A Benchmark for Socially Compliant Vision-and-Language Navigation in Dynamic Pedestrian Environments

DPed-VLN:动态步行环境中社会合规的视觉与语言导航基准

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

IncentRL:偏好指导与任务绩效之间的权衡

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

OneBid:针对多种oCPX广告场景的统一自动竞价基础模型

Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

在接触富丰富操作中强化学习中的势场动作表示

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

SynthDemo-RL:通过LLM引导的合成演示突破VLA适应中的零奖励障碍

DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning

DRT:密集推理迹,用于高效且扎根的多模态推理

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC:通过专家指导规划平衡勘探与开发

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

从预训练到熟练:以最小人工干预实现长期操作的现实子任务强化学习

Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning

超越运动学:肌肉驱动模仿学习的模拟忠真度基准测试

What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

接下来我们应该问什么?部分证据下的检索感知问题学习

Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks

学习移动城市:城市网络校准与控制的深度元模型与强化策略

Automata-Theoretic Verification of Interval Markov Decision Processes

自动机理论验证区间马尔可夫决策过程

$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

$λ$-受控GRPO:将流量匹配比率不稳定性转化为预算资源

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

CodeMidas:从代码本身扩展代理编码强化学习环境

MintAct: A Unified Visual Agent for Digital Environments

MintAct:数字环境统一视觉代理

Keyword: diffusion policy

Demonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning

通过高斯喷涂单次扫描演示合成,用于视觉运动策略学习

Robotic Multiphase Interaction: Manipulating Coupled Liquid and Solid Dynamics with a World Model

机器人多相相互作用:利用世界模型操控耦合的液体和固体动力学

Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning

迈向细粒度物体操作:SAM3引导的身体运动策略结合持续记忆学习和聚焦视觉条件反射

AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance

声学扩散:用于搜救协助的语义条件音频引导扩散政策