Nepali Neural Lemmatizer
This model is a part of the study "Evaluating Multilingual Transformer Models for Lemmatization in Nepali: A Low-Resource Case Study" presented at LREC-COLING 2026. It addresses the challenges of lemmatization in the morphologically rich and low-resource Nepali language by leveraging pre-trained multilingual Transformer models.
π Model Description
This project frames Nepali lemmatization as a text-to-text generation problem. We fine-tuned several state-of-the-art architectures to map inflected Nepali words (e.g., ΰ€ΰ€Ύΰ€¨ΰ₯ΰ€ΰ₯) to their canonical base forms (e.g., ΰ€ΰ€Ύΰ€¨ΰ₯).
Key Features:
- Architecture: Based on
mT5-base,mT5-small, andmBART-large-50. - Script: Optimized for the Devanagari script.
- Robustness: Handles Out-of-Vocabulary (OOV) words and spelling inconsistencies better than traditional rule-based or TRIE-based systems.
π Performance
Our experiments show that the mT5-base model achieved the highest precision, while mBART-large-50 demonstrated superior morphological coverage.
| Model | CER β | Accuracy β | BLEU (Char) β | Morph. Coverage β |
|---|---|---|---|---|
| mT5-base | 1.1% | 96.1% | 0.980 | 0.964 |
| mBART-large-50 | 1.6% | 96.0% | 0.986 | 0.970 |
| mT5-small | 1.7% | 95.2% | 0.983 | 0.955 |
Downstream Impact:
The use of this lemmatizer significantly improves performance in other NLP tasks:
- Cross-lingual Alignment: Improved accuracy from 12.86% to 41.61% for Hindi-Nepali word alignment.
- Information Retrieval: Increased Mean Average Precision (MAP) from 0.71 to 0.90.
π Usage
You can use this model through the Hugging Face transformers library:
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_name = "sunilregmi/nepali-lemmatizerV1-mt5-base" # Example HF path
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
def lemmatize(word):
# Prompt format used during training
text = f"lemmatize: {word}"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(**inputs, max_length=32)
return tokenizer.decode(outputs[0], skip_special_tokens=True)
print(lemmatize("ΰ€ΰ₯ΰ€°ΰ€Ώΰ€―ΰ€€ΰ€Έΰ€Ήΰ€Ώΰ€€ΰ€ΰ€Ύ")) # Output: ΰ€ΰ₯ΰ€°ΰ€Ώΰ€―ΰ€€
π Training Data
The model was trained on a curated corpus of 8,000 unique word-lemma pairs:
- Gold Standard: 5,000 manually verified pairs.
- Augmented Data: 3,000 samples generated based on common Nepali inflectional patterns and morphological rules.
π Citation
If you use this model or the research findings, please cite:
@inproceedings{regmi2026evaluating,
title={Evaluating Multilingual Transformer Models for Lemmatization in Nepali: A Low-Resource Case Study},
author={Regmi, Sunil and Dawadi, Sundeep and Bal, Bal Krishna},
booktitle={Proceedings of the 31st International Conference on Computational Linguistics (LREC-COLING 2026)},
year={2026}
}
π Resources
- Downloads last month
- 12