ISSN :2582-9793

Comparative Analysis of Classical Information Retrieval Models for Relevance Judgment List Expansion

Original Research (Published On: 24-Jun-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.63319

Fadi Ibrahim Yamout

Adv. Artif. Intell. Mach. Learn., 6 (3):5769-5784

1. Fadi Ibrahim Yamout: Lebanese International University

Download PDF Here

DOI: 10.54364/AAIML.2026.63319

Article History: Received on: 07-Jan-26, Accepted on: 16-Jun-26, Published on: 24-Jun-26

Corresponding Author: Fadi Ibrahim Yamout

Email: fadi.yamout@liu.edu.lb

Citation: Fadi Yamout. Comparative Analysis of Classical Information Retrieval Models for Relevance Judgment List Expansion. Advances in Artificial Intelligence and Machine Learning. 2026;6(3):319. https://dx.doi.org/10.54364/AAIML.2026.63319


Abstract

    

Document representations commonly used in classical information retrieval (IR), combined with machine learning classifiers, can be used to automatically expand upon relevance judgment lists with high fidelity. Automatically expanded judgments can be useful when existing relevance assessments are incomplete, which is frequently the case given how labor-intensive and expensive relevance assessment can be. We conduct experiments on multiple IR test collections comparing various IR document representations combined with two classifiers. The IR document representations used include TF-IDF, cosine-normalized TF-IDF, BM25, variants of DFR (PL2, INL2, and IFB2), Dirichlet-smoothed language modeling (LMDir), and Binary weighting. Separate Naïve Bayes and Support Vector Machine (SVM) classifiers were trained on each document representation given existing relevance judgments and then used to predict the relevance of previously-unjudged documents. The experimental results, supported by paired randomization significance testing, demonstrate that both document representation and classifier selection significantly affect relevance judgment expansion performance. Empirically, using paired randomization significance testing, we find that SVM consistently outperforms Naïve Bayes, and TF-IDF and BM25 are typically among the top-performing representations. Additional experiments designed to measure rank-preserving characteristics of the expanded judgments using Kendall's Tau show that the expanded judgments preserve retrieval-system rankings better than the incomplete judgments do, and that rankings generated using expanded judgments are statistically significantly closer to the rankings generated using the originally complete set of relevance assessments.

Statistics

   Article View: 441
   PDF Downloaded: 9