二、The algorithm of deterministic actor-critic基于policy gradient,the gradient-ascent algorithm就可以最大化J ( θ ) J(\theta)