Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Which approach in word embedding training predicts surrounding context words from a center word?

Predicting surrounding context words from a center word is the essence of the Skip-Gram approach. In this setup, you take a single center word and the model learns to maximize the probability of each word that appears within a chosen window around it. As training proceeds, the word vectors adjust so that words that tend to appear in similar contexts end up close in the vector space. This is why Skip-Gram is especially good at capturing semantic relationships and works well with rare words, since it treats each center word as the predictor for its neighbors. In contrast, another common model—Continuous Bag-of-Words—does the opposite: it uses the surrounding context words to predict the center word. Doc2Vec extends the idea to document-level embeddings, incorporating a vector for the whole document, while Latent Semantic Analysis relies on matrix factorization rather than predictive training. So the approach that aligns with predicting context from a center word is Skip-Gram.

Predicting surrounding context words from a center word is the essence of the Skip-Gram approach. In this setup, you take a single center word and the model learns to maximize the probability of each word that appears within a chosen window around it. As training proceeds, the word vectors adjust so that words that tend to appear in similar contexts end up close in the vector space. This is why Skip-Gram is especially good at capturing semantic relationships and works well with rare words, since it treats each center word as the predictor for its neighbors.

In contrast, another common model—Continuous Bag-of-Words—does the opposite: it uses the surrounding context words to predict the center word. Doc2Vec extends the idea to document-level embeddings, incorporating a vector for the whole document, while Latent Semantic Analysis relies on matrix factorization rather than predictive training. So the approach that aligns with predicting context from a center word is Skip-Gram.