Perbandingan Kinerja Metode TF-IDF Dan BM25 Pada Pencarian Artikel Berita Bahasa Indonesia

Authors

  • Muhammad Al Adib Universitas Potensi Utama
  • Yiska Dayanti Zagoto
  • Bualazatulo Laia
  • Roslina

Keywords:

BM25, document ranking, Indonesian news retrieval, information retrieval, TF-IDF

Abstract

The rapid growth of online media has significantly increased the number of Indonesian news articles, creating the need for an information retrieval system capable of retrieving relevant documents efficiently and accurately. This study aims to compare the performance of the Term Frequency–Inverse Document Frequency (TF-IDF) and Best Matching 25 (BM25) methods for Indonesian news article retrieval. The dataset used was the Indonesia News Dataset (2025), consisting of 80,472 news articles. The research process included data exploration, preprocessing through case folding, cleaning, stopword removal, and tokenization, followed by the development of TF-IDF and BM25 retrieval models. Performance evaluation was conducted using the query "harga cabai" by comparing document relevance, ranking scores, and retrieval execution time. The results show that both methods successfully retrieved relevant documents and produced the same top-ranked article. However, BM25 demonstrated superior efficiency with a retrieval time of 0.097587 seconds, while TF-IDF required 0.417316 seconds. These findings indicate that BM25 is more suitable for Indonesian news article retrieval, as it maintains retrieval effectiveness while providing faster search performance.

Keywords: BM25; Information Retrieval; news article retrieval; TF-IDF.

Downloads

Download data is not yet available.

References

The growth of online media has increased the number of Indonesian-language news articles and created a need for effective and efficient retrieval systems. This study compares Term Frequency–Inverse Document Frequency (TF-IDF) and Best Matching 25 (BM25) using the Indonesia News Dataset (2025). Of 80,472 articles, 80,463 documents were retained after removing nine articles with empty content. The evaluation used 10 multi-category queries. Ground truth was constructed by pooling the top 20 results from each method, producing 304 query–document pairs manually assessed by one evaluator using relevance labels from 0 to 2. Retrieval effectiveness was measured using Precision@k, Recall@k, Mean Average Precision (MAP), Mean Reciprocal Rank (MRR), and normalized Discounted Cumulative Gain (nDCG). Retrieval time was measured over 30 repetitions after three warm-up runs. At Top-10, BM25 achieved a Precision@10 of 0.79, Recall@10 of 0.351630, MAP@10 of 0.728635, and nDCG@10 of 0.630976, which were higher than the corresponding TF-IDF scores. TF-IDF obtained a slightly higher MRR of 0.90 than BM25 at 0.85. In the evaluated implementation, BM25 required 0.010191 seconds on average, whereas TF-IDF required 0.066014 seconds. Descriptively, BM25 performed better on most Top-10 metrics and retrieval time, but the results depended on query characteristics and experimental configuration

Downloads

Published

2026-10-07