Multilingual Reranking & Retrieval Leaderboard

Open rerankers (cross-encoders, LLM rerankers) and embedding models on the same multilingual reranking task: each query's candidate list from public MTEB reranking sets is re-sorted by the model, scored with nDCG@10. Embedding models rank candidates by cosine similarity. Same script and same queries for every model (code: rerank/eval_rerank.py, emb/eval_emb.py).

Sets: MIRACL dev (18 languages, 60 queries x 100 candidates each), WikipediaRerankingMultilingual (16 languages, 60 x 9), ESCI (es/jp/us product search), RuBQ (ru), T2Reranking (zh), mMARCO (ja), AskUbuntu (en). "mean" = mean of the MIRACL, Wikipedia and "other" averages. Caveats: these are subsets (not full MTEB scores); bge-reranker-v2-m3 was trained on MIRACL's training split; "→ bge-m3 index" = queries embedded by the small model, candidates by bge-m3. Click a column to sort.