生成时间: 2026-09-25 21:21:26 (UTC+8); Arxiv 发布时间: 2026-09-25 20:00 EDT (2026-09-26 08:00 UTC+8)

今天共有 40 篇相关文章

Keyword: reinforcement learning

Pistis Technical Report

Pistis技术报告

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

DEEPO:双熵增强策略优化,用于MLLM中的幻觉

RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory

RLVR的迭代乘法景观可能是良性的:自旋玻璃理论的见解

Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

形态计量模仿:从形态学和接触感知手部重定向到模拟到真实的体力运动策略

Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

超越表面风格:将多回合用户模拟器与行为一致性相结合

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

强化学习中的策略复杂性、反应时间与有界理性

Reinforcement Learning with Verifiable Rewards for Small Search Agents

小型搜索代理的可验证奖励强化学习

An Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics

机器人学中流式深度强化学习在自适应持续学习中的分析

FlyCNS: Connectome-Grounded Information Organization for Communication-Constrained Embodied Control

FlyCNS:连接组接地信息组织,用于通信受限的内在控制

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

不确定性门控探索噪声抑制在线强化学习中任务崩溃,对流程匹配视觉-语言-行动策略进行微调

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

返回定义:通过轨迹图估算代理强化学习的步骤级优势

Deep-learning-aided dismantling of interdependent networks

深度学习辅助拆解相互依赖网络

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

从混合质量部署经验中学习机器人操作

Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching

结果敏感运动搜索,实现冲击感知的灵巧接球

Simple Torque-Observation Alignment for Zero-Shot Sim-to-Real Grasping with a Direct-Drive Gripper

简单的扭矩观察对准,用于零次模拟到实抓,使用直接驱动夹持器

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

SLCA-GRPO:解决工具调用强化学习中的跨部门信用错误归因

A Human-Like Pedestrian Model for Automated Driving Simulations

一种类人步行模型用于自动驾驶模拟

Right Choice of Classification Algorithms Based on Reinforcement Learning for Prediction of Non-Alcoholic Fatty Liver

基于强化学习的分类算法正确选择,用于预测非酒精性脂肪肝

EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards

EAGER:通过可验证奖励的强化学习提升生成事件提取

Safe Learning-Based Adaptive Augmentation Control for Fixed-Wing UAV under Uncertainty

固定翼无人机在不确定性下的安全基于学习的自适应增强控制

Temperament Engineering: Designing Strategic Behavioural Diversity in Robot Swarms

气质工程:设计机器人群体中的战略行为多样性

Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators

混合刚性-气动机械臂的耦合状态空间建模、控制与策略蒸馏

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

问问Jev:作为AI对齐失败零射值检测器的校准决策强化学习

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

IterSynth:通过角色解耦迭代合成重新思考深度搜索代理

AdaPilot: Towards Scene-Adaptive Policy Learning for Cross-Generator Text-to-Image Quality Optimization

AdaPilot:迈向场景自适应策略学习,实现跨生成器文本转图像质量优化

CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning

CataOPD:用于大型语言模型推理的催化策略上提纯

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

认证的预测价值建议门禁,用于强化学习中成本感知的语言模型指导

iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

iCoder-27B:递归AI驱动的前沿工业编码模型开发

To Think or Not to Think: Allocating Reasoning Where It Helps

思考还是不思考:在有帮助的地方分配推理

Learning to Ideate for Scientific Impact

学习如何为科学产生影响而构思

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Qwen-Planner-Agent:一个面向现实世界移动规划代理的闭环AI框架

Learning Better Reasoning for Generative Recommendation with Semantic IDs

学习更好地推理语义ID生成式推荐

Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation

Res-HIL:人主导残余强化学习,用于样本高效灵巧操作

TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs

TEMA:基于证据的时间问答,采用多回合多音频对话

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

SciWalker:综合科学编码问题与算子图和执行反馈

Self-Play Pretraining with Zero Data

零数据自玩预训练

Graph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management

基于图的推理与拓扑感知的多智能体强化学习,适用于大规模铁路网络管理

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

用于 Roblox 游戏搜索中多元查询理解的搜索感知强化学习

PoEM: Predicting RL Outcomes from Existing Policies

PoEM:从现有政策预测强化学习结果

Keyword: diffusion policy

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen:基于类别级策略泛化的无监督关键点对象中心表示