跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 2026-08-05AI 评分44

教 Nemotron 学希腊语:挖掘语料、适配检索并实现生成落地

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

AI 导读

一项研究将 NVIDIA Nemotron 检索栈端到端适配至现代希腊语,覆盖语料挖掘、检索训练、重排器适配与阅读器微调。微调后 Nemotron 1B 嵌入器 nDCG@10 从 0.362 提升至 0.835,LoRA 微调的 Nemotron 30B-A3B 阅读器将答案正确率从 29.4% 提升至 66.9%。

正文

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org