Data-Efficient Safe Policy Improvement Using Parametric Structure

cs.AI updates on arXiv.org 07月22日 12:34

Data-Efficient Safe Policy Improvement Using Parametric Structure

本文提出一种基于参数依赖的SPI算法，通过利用已知分布间的相关性，提高数据效率，并采用预处理技术减少冗余动作，显著提升安全策略改进的数据效率。

arXiv:2507.15532v1 Announce Type: new Abstract: Safe policy improvement (SPI) is an offline reinforcement learning problem in which a new policy that reliably outperforms the behavior policy with high confidence needs to be computed using only a dataset and the behavior policy. Markov decision processes (MDPs) are the standard formalism for modeling environments in SPI. In many applications, additional information in the form of parametric dependencies between distributions in the transition dynamics is available. We make SPI more data-efficient by leveraging these dependencies through three contributions: (1) a parametric SPI algorithm that exploits known correlations between distributions to more accurately estimate the transition dynamics using the same amount of data; (2) a preprocessing technique that prunes redundant actions from the environment through a game-based abstraction; and (3) a more advanced preprocessing technique, based on satisfiability modulo theory (SMT) solving, that can identify more actions to prune. Empirical results and an ablation study show that our techniques increase the data efficiency of SPI by multiple orders of magnitude while maintaining the same reliability guarantees.

Fish AI Reader

AI辅助创作，多种专业模板，深度分析，高质量内容生成。从观点提取到深度思考，FishAI为您提供全方位的创作支持。新版本引入自定义参数，让您的创作更加个性化和精准。

FishAI

鱼阅，AI 时代的下一个智能信息助手，助你摆脱信息焦虑

联系邮箱 441953276@qq.com

相关标签

安全策略改进数据效率预处理技术

相关文章

Unifying Vision and Language Models with Mohit Bansal - #636

Use of Pretrained BERT to Predict the Rating of Reviews

Google DeepMind Researchers Introduce Diffusion Augmented Agents: A Machine Learning Framework for Efficient Exploration and Transfer Learning

Researchers from Princeton University Introduce Metadata Conditioning then Cooldown (MeCo) to Simplify and Optimize Language Model Pre-training

少用33％数据，模型性能不变，陈丹琦团队用元数据来做降本增效

少用33％数据，模型性能不变，陈丹琦团队用元数据来做降本增效

This AI Paper from UC Berkeley Introduces a Data-Efficient Approach to Long Chain-of-Thought Reasoning for Large Language Models

机器人泛化能力大幅提升：HAMSTER层次化方法和VLA尺度轨迹预测，显著提升开放世界任务成功率

OpenAI：现在只需 5~10 人即可从头重建 GPT-4，发展瓶颈已经从“算力”转变为“数据效率”

OpenAI 揭秘 GPT-4.5 训练：10 万块 GPU，几乎全员上阵，出现“灾难性问题”