二、最简单的 MC-based RL algorithm 【蒙特卡洛方法思想核心,但数据利用效率太低】蒙特卡洛方法思想核心,但效率太低,在实际中无法使用。1、将策略迭代转换为无模型方法理解该算法的关键是理解how to convert the policy iteration algorithm to be model-free.首先,策略迭代算法在每次迭代中分为两步:{ Policy