arXiv · 1004.4489
MIREX: MapReduce Information Retrieval Experiments
Abstract
We propose to use MapReduce to quickly test new retrieval approaches on a cluster of machines by sequentially scanning all documents. We present a small case study in which we use a cluster of 15 low cost ma- chines to search a web crawl of 0.5 billion pages showing that sequential scanning is a viable approach to running large-scale information retrieval experiments with little effort. The code is available to other researchers at: http://mirex.sourceforge.net
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Djoerd Hiemstra, Claudia Hauff. 2010-04-26. MIREX: MapReduce Information Retrieval Experiments. https://arxiv.org/abs/1004.4489
Cite the original work for its findings. Save a collection to share your selection of sources.