Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Which term describes the measure used to determine similarity between word vectors?

Cosine similarity is the measure used to determine similarity between word vectors. It compares the angle between two vectors, not their lengths, so words that are conceptually related point in similar directions even if their overall magnitudes differ. The calculation is the dot product of the vectors divided by the product of their magnitudes, yielding a value from -1 to 1 where values near 1 indicate strong similarity. This is particularly useful for word embeddings because training can produce vectors with varying magnitudes due to word frequency or other factors, while the semantic relationship is captured by direction. Cosine similarity therefore focuses on orientation in the embedding space, which aligns with how we interpret word similarity. The other terms describe how word vectors are created or weighted, not how to measure similarity. Word embeddings are the vector representations themselves. Inverse Document Frequency downweights common terms in document representations. Continuous Bag of Words is a model architecture for learning word vectors.

Cosine similarity is the measure used to determine similarity between word vectors. It compares the angle between two vectors, not their lengths, so words that are conceptually related point in similar directions even if their overall magnitudes differ. The calculation is the dot product of the vectors divided by the product of their magnitudes, yielding a value from -1 to 1 where values near 1 indicate strong similarity.

This is particularly useful for word embeddings because training can produce vectors with varying magnitudes due to word frequency or other factors, while the semantic relationship is captured by direction. Cosine similarity therefore focuses on orientation in the embedding space, which aligns with how we interpret word similarity.

The other terms describe how word vectors are created or weighted, not how to measure similarity. Word embeddings are the vector representations themselves. Inverse Document Frequency downweights common terms in document representations. Continuous Bag of Words is a model architecture for learning word vectors.