Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Which approach focuses on finding the optimal policy that maps a given state to an action?

Policy-based approach directly optimizes the policy that maps each state to an action. It treats the policy as the object to be learned and uses methods like policy gradients to adjust the policy parameters so as to maximize the expected cumulative reward. This means the algorithm searches for the best mapping from states to actions rather than first estimating a value function and deriving actions from it. Value-based methods, by contrast, focus on learning value functions (the value of states or state-action pairs) and derive the action choice from those values (for example, choosing actions with the highest Q-value). Q-Learning is a classic example of a value-based method. Monte Carlo methods estimate returns and can be used in either framework, but they are not inherently about directly optimizing the state-to-action mapping. Thus, the approach that centers on finding the optimal state-to-action mapping is the policy-based approach.

Policy-based approach directly optimizes the policy that maps each state to an action. It treats the policy as the object to be learned and uses methods like policy gradients to adjust the policy parameters so as to maximize the expected cumulative reward. This means the algorithm searches for the best mapping from states to actions rather than first estimating a value function and deriving actions from it. Value-based methods, by contrast, focus on learning value functions (the value of states or state-action pairs) and derive the action choice from those values (for example, choosing actions with the highest Q-value). Q-Learning is a classic example of a value-based method. Monte Carlo methods estimate returns and can be used in either framework, but they are not inherently about directly optimizing the state-to-action mapping. Thus, the approach that centers on finding the optimal state-to-action mapping is the policy-based approach.