Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking

Domain:

natural language processing

Record type:

papermodel
Creator:
DanNguPha
Host:avatar
This paper presents ViRanker, a cross-encoder reranking model tailored to the Vietnamese language. Built on the BGE-M3 encoder and enhanced with the Blockwise Parallel Transformer, ViRanker addresses the lack of competitive rerankers for Vietnamese, a low-resource language with complex syntax and diacritics. The model was trained on an 8 GB curated corpus and fine-tuned with hybrid hard-negative sampling to strengthen robustness. Evaluated on the MMARCO-VI benchmark, ViRanker achieves strong early-rank accuracy, surpassing multilingual baselines and competing closely with PhoRanker. By releasing the model openly on Hugging Face, we aim to support reproducibility and encourage wider adoption in real-world retrieval systems. Beyond Vietnamese, this study illustrates how careful architectural adaptation and data curation can advance reranking in other underrepresented languages. 9 pages

Visit

arxiv.org

Tasks

information retrieval

Tags

Computation and LanguageArtificial Intelligence

Similar

mradermacher/BGE-m3-swahili-GGUFllama-lang-adapt/BGE-m3-swahiliAmazigh Speech Recognition via Parallel CNN Transformer-Encoder Modelychafiqui/bge-m3-law-morocco-ar-embllama-lang-adapt/BGE-m3-swahili-33Exploring Graph-based Transformer Encoder for Low-Resource Neural Machine Translation

mradermacher/BGE-m3-swahili-GGUF

llama-lang-adapt/BGE-m3-swahili

Amazigh Speech Recognition via Parallel CNN Transformer-Encoder Model

ychafiqui/bge-m3-law-morocco-ar-emb

llama-lang-adapt/BGE-m3-swahili-33

Exploring Graph-based Transformer Encoder for Low-Resource Neural Machine Translation

The Transformer is commonly used in Neural Machine Translation (NMT), but it faces issues with over-