Get ready for the GARP Risk and AI Exam with flashcards and multiple choice questions. Each question comes with hints and explanations. Prepare for success!

Multiple Choice

Which measure downweights common words to reflect rarity across a corpus?

Inverse document frequency is the measure that downweights words based on how common they are across the whole collection. It looks at how many documents contain a term; if a word appears in many documents, its IDF is small, so it contributes less to the document’s weight. Rare terms, appearing in few documents, get a higher weight and thus stand out in matching queries or distinguishing documents. This idea is central to TF-IDF: you multiply a term’s frequency in a document by its IDF to balance local importance with global rarity. BM25 uses IDF as part of its scoring but the specific downweighting of across-corpus rarity is the IDF itself. Latent Semantic Analysis doesn’t perform this kind of term-by-term rarity weighting; it uses matrix factorization to capture latent structure in the data.

Inverse document frequency is the measure that downweights words based on how common they are across the whole collection. It looks at how many documents contain a term; if a word appears in many documents, its IDF is small, so it contributes less to the document’s weight. Rare terms, appearing in few documents, get a higher weight and thus stand out in matching queries or distinguishing documents. This idea is central to TF-IDF: you multiply a term’s frequency in a document by its IDF to balance local importance with global rarity. BM25 uses IDF as part of its scoring but the specific downweighting of across-corpus rarity is the IDF itself. Latent Semantic Analysis doesn’t perform this kind of term-by-term rarity weighting; it uses matrix factorization to capture latent structure in the data.