Nepali Neural Lemmatizer

This model is a part of the study "Evaluating Multilingual Transformer Models for Lemmatization in Nepali: A Low-Resource Case Study" presented at LREC-COLING 2026. It addresses the challenges of lemmatization in the morphologically rich and low-resource Nepali language by leveraging pre-trained multilingual Transformer models.

πŸš€ Model Description

This project frames Nepali lemmatization as a text-to-text generation problem. We fine-tuned several state-of-the-art architectures to map inflected Nepali words (e.g., ΰ€–ΰ€Ύΰ€¨ΰ₯ΰ€›ΰ₯) to their canonical base forms (e.g., ΰ€–ΰ€Ύΰ€¨ΰ₯).

Key Features:

  • Architecture: Based on mT5-base, mT5-small, and mBART-large-50.
  • Script: Optimized for the Devanagari script.
  • Robustness: Handles Out-of-Vocabulary (OOV) words and spelling inconsistencies better than traditional rule-based or TRIE-based systems.

πŸ“Š Performance

Our experiments show that the mT5-base model achieved the highest precision, while mBART-large-50 demonstrated superior morphological coverage.

Model CER ↓ Accuracy ↑ BLEU (Char) ↑ Morph. Coverage ↑
mT5-base 1.1% 96.1% 0.980 0.964
mBART-large-50 1.6% 96.0% 0.986 0.970
mT5-small 1.7% 95.2% 0.983 0.955

Downstream Impact:

The use of this lemmatizer significantly improves performance in other NLP tasks:

  • Cross-lingual Alignment: Improved accuracy from 12.86% to 41.61% for Hindi-Nepali word alignment.
  • Information Retrieval: Increased Mean Average Precision (MAP) from 0.71 to 0.90.

πŸ›  Usage

You can use this model through the Hugging Face transformers library:

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_name = "sunilregmi/nepali-lemmatizerV1-mt5-base" # Example HF path
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)

def lemmatize(word):
    # Prompt format used during training
    text = f"lemmatize: {word}"
    inputs = tokenizer(text, return_tensors="pt")
    outputs = model.generate(**inputs, max_length=32)
    return tokenizer.decode(outputs[0], skip_special_tokens=True)

print(lemmatize("ΰ€–ΰ₯ˆΰ€°ΰ€Ώΰ€―ΰ€€ΰ€Έΰ€Ήΰ€Ώΰ€€ΰ€•ΰ€Ύ")) # Output: ΰ€–ΰ₯ˆΰ€°ΰ€Ώΰ€―ΰ€€

πŸ“‚ Training Data

The model was trained on a curated corpus of 8,000 unique word-lemma pairs:

  • Gold Standard: 5,000 manually verified pairs.
  • Augmented Data: 3,000 samples generated based on common Nepali inflectional patterns and morphological rules.

πŸ“ Citation

If you use this model or the research findings, please cite:

@inproceedings{regmi2026evaluating,
  title={Evaluating Multilingual Transformer Models for Lemmatization in Nepali: A Low-Resource Case Study},
  author={Regmi, Sunil and Dawadi, Sundeep and Bal, Bal Krishna},
  booktitle={Proceedings of the 31st International Conference on Computational Linguistics (LREC-COLING 2026)},
  year={2026}
}

πŸ”— Resources

Downloads last month
12
Safetensors
Model size
0.6B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support