MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair.
Code (1)
Tasks
Information RetrievalRe-RankingRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
G\'en\'erer une grammaire d'arbres adjoints pour l'arabe \`a partir d'une m\'eta-grammaire (Generate a tree adjoining grammar for arabic from a meta-grammar)
La raret{\'e} des ressources num{\'e}riques pour la langue arabe, telles que les grammaires et corpus, rend son traitement plus difficile que les autres langues naturelles. A ce jour il n{'}existe pas une grammaire forme…
UMAIR-FPS: User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style
The rapid advancement of high-quality image generation models based on AI has generated a deluge of anime illustrations. Recommending illustrations to users within massive data has become a challenging and popular task. …
Image GenerationMulti-modal RecommendationRecommendation SystemsSentenceGrammaires phrastiques et discursives fond\'ees sur les TAG : une approche de D-STAG avec les ACG
Nous pr{\'e}sentons une m{\'e}thode pour articuler grammaire de phrase et grammaire de discours qui {\'e}vite de recourir {\`a} une {\'e}tape de traitement interm{\'e}diaire. Cette m{\'e}thode est suffisamment g{\'e}n{\'…
TAGEdge Federated Learning Via Unit-Modulus Over-The-Air Computation
Edge federated learning (FL) is an emerging paradigm that trains a global parametric model from distributed datasets based on wireless communications. This paper proposes a unit-modulus over-the-air computation (UMAirCom…
Autonomous DrivingFederated LearningMAIR++: Improving Multi-view Attention Inverse Rendering with Implicit Lighting Representation
In this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely us…
Inverse Rendering