Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Reinforcement learning aims to maximize which of the following?

In reinforcement learning the goal is to maximize the total return collected over time. The agent learns a policy that, when followed, yields the highest expected sum of rewards across the horizon (often written as a discounted sum of future rewards). This emphasis on the long term means actions are valued not just for the immediate payoff but for how they influence future opportunities and outcomes. A high discount factor makes the agent consider far-reaching effects, while a low discount factor makes it focus more on near-term rewards. If you aimed only for short-term rewards or immediate correctness, you’d risk missing bigger gains that come from planning ahead. Prediction accuracy isn’t the objective here; the feedback signal is the rewards, and the policy is evaluated by the cumulative return it produces.

In reinforcement learning the goal is to maximize the total return collected over time. The agent learns a policy that, when followed, yields the highest expected sum of rewards across the horizon (often written as a discounted sum of future rewards). This emphasis on the long term means actions are valued not just for the immediate payoff but for how they influence future opportunities and outcomes. A high discount factor makes the agent consider far-reaching effects, while a low discount factor makes it focus more on near-term rewards. If you aimed only for short-term rewards or immediate correctness, you’d risk missing bigger gains that come from planning ahead. Prediction accuracy isn’t the objective here; the feedback signal is the rewards, and the policy is evaluated by the cumulative return it produces.