You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

mT5 Base Fine-tuned for Sinhala Spell Correction

An mT5-base model fine-tuned for Sinhala spell correction as part of the LMSpell project.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_name = "lm-spell/mt5-base-ft-ssc"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
ZWJ_token = '<ZWJ>'
ZWJ_unicode = '\u200D'

text = "යින් හැඟෙන්නේ පවත්නා අධ්‍යාපන ක්‍රමයේ පවතින" # max tokens in mt5 = 128
text = text.replace(ZWJ_unicode, ZWJ_token) 

inputs = tokenizer(text, return_tensors="pt")
output = model.generate(**inputs)

zwj_id = tokenizer.convert_tokens_to_ids("<ZWJ>")
special_ids = set(tokenizer.all_special_ids) - {zwj_id}

tokens = [token for token in output[0].tolist() if token not in special_ids]

decoded = tokenizer.decode(tokens, skip_special_tokens=False)
decoded = decoded.replace(f'{ZWJ_token} ',ZWJ_unicode)
print(decoded)

Zero Width Joiner Handling

mT5 family of models do not preserve the Unicode Zero Width Joiner (\u200D) during tokenization and generation. To address this, LMSpell represents the Zero Width Joiner using a custom special token, <ZWJ>, during inference.

Before tokenization, occurrences of \u200D are replaced with <ZWJ>. After generation and decoding, <ZWJ> is converted back to the original Unicode Zero Width Joiner (\u200D).

This ensures that Zero Width Joiners are preserved throughout the spell-correction process.

Limitations

  • The model is specifically fine-tuned for Sinhala spell correction and may not generalize to other languages.
  • Performance may vary for informal text, social media content, code-mixed text, and text containing uncommon or domain-specific vocabulary.
  • The model may occasionally modify text that is already correctly spelled.

Research

For details on the training methodology, dataset preparation, evaluation, results, limitations, and comparisons with other models, please refer to the LMSpell paper:

LMSpell: Spell Correction with Pre-Trained Language Models

A. Gunathilake, N. Karunarathna, T. Bandaranayake, S. Ranathunga, N. de Silva and N. Jayatilleke, "LMSpell: Spell Correction with Pre-Trained Language Models," 2026 Moratuwa Engineering Research Conference (MERCon), Moratuwa, Sri Lanka, 2026, pp. 503-508.

Citation

@INPROCEEDINGS{11691371,
  author={Gunathilake, Akesh and Karunarathna, Nadil and Bandaranayake, Tharusha and Ranathunga, Surangika and de Silva, Nisansa and Jayatilleke, Nevidu},
  booktitle={2026 Moratuwa Engineering Research Conference (MERCon)},
  title={LMSpell: Spell Correction with Pre-Trained Language Models},
  year={2026},
  pages={503-508},
  doi={10.1109/MERCon71835.2026.11691371}
}
Downloads last month
29
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train lm-spell/mt5-base-ft-ssc

Space using lm-spell/mt5-base-ft-ssc 1