Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Which term describes the architecture that enables attention across the entire input and parallel processing?

Transformers are built to attend to every position in the input at once and to do so in parallel. They use self-attention to calculate, for each token, how much every other token should influence its representation, effectively weighting and summing contributions from all parts of the sequence. This global, token-to-token interaction happens in a single pass, so the entire sequence can be processed at the same time rather than step by step. Positional information is added separately so the model still understands the order of tokens. This combination lets the model capture long-range dependencies efficiently and with high parallel throughput. Other options either describe components (attention mechanisms) or rely on architectures that process data sequentially in time (recurrent neural networks) or focus on local patterns (convolutional neural networks), which don’t inherently provide the same global, parallel attention across the whole input.

Transformers are built to attend to every position in the input at once and to do so in parallel. They use self-attention to calculate, for each token, how much every other token should influence its representation, effectively weighting and summing contributions from all parts of the sequence. This global, token-to-token interaction happens in a single pass, so the entire sequence can be processed at the same time rather than step by step. Positional information is added separately so the model still understands the order of tokens. This combination lets the model capture long-range dependencies efficiently and with high parallel throughput. Other options either describe components (attention mechanisms) or rely on architectures that process data sequentially in time (recurrent neural networks) or focus on local patterns (convolutional neural networks), which don’t inherently provide the same global, parallel attention across the whole input.