本章统一比较两条强化学习算法演进路线:REINFORCE→Actor-Critic→A2C→TRPO→PPOREINFORCE\rightarrow Actor\text{-}Critic\rightarrow A2C\rightarrow TRPO\rightarrow PPOREINFO