Instructions to use lm-spell/mt5-base-ft-ssc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lm-spell/mt5-base-ft-ssc with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("lm-spell/mt5-base-ft-ssc") model = AutoModelForSeq2SeqLM.from_pretrained("lm-spell/mt5-base-ft-ssc", device_map="auto") - Notebooks
- Google Colab
- Kaggle
mT5 Base Fine-tuned for Sinhala Spell Correction
An mT5-base model fine-tuned for Sinhala spell correction as part of the LMSpell project.
- Model: mT5-base
- Base model:
google/mt5-base - Language: Sinhala (
si) - Task: Sinhala Spell Correction (SSC)
- Dataset: Sinhala Spell Correction Dataset
- License: CC BY 4.0
- Project: LMSpell
Usage
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_name = "lm-spell/mt5-base-ft-ssc"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
ZWJ_token = '<ZWJ>'
ZWJ_unicode = '\u200D'
text = "යින් හැඟෙන්නේ පවත්නා අධ්යාපන ක්රමයේ පවතින" # max tokens in mt5 = 128
text = text.replace(ZWJ_unicode, ZWJ_token)
inputs = tokenizer(text, return_tensors="pt")
output = model.generate(**inputs)
zwj_id = tokenizer.convert_tokens_to_ids("<ZWJ>")
special_ids = set(tokenizer.all_special_ids) - {zwj_id}
tokens = [token for token in output[0].tolist() if token not in special_ids]
decoded = tokenizer.decode(tokens, skip_special_tokens=False)
decoded = decoded.replace(f'{ZWJ_token} ',ZWJ_unicode)
print(decoded)
Zero Width Joiner Handling
mT5 family of models do not preserve the Unicode Zero Width Joiner (\u200D) during tokenization and generation. To address this, LMSpell represents the Zero Width Joiner using a custom special token, <ZWJ>, during inference.
Before tokenization, occurrences of \u200D are replaced with <ZWJ>. After generation and decoding, <ZWJ> is converted back to the original Unicode Zero Width Joiner (\u200D).
This ensures that Zero Width Joiners are preserved throughout the spell-correction process.
Limitations
- The model is specifically fine-tuned for Sinhala spell correction and may not generalize to other languages.
- Performance may vary for informal text, social media content, code-mixed text, and text containing uncommon or domain-specific vocabulary.
- The model may occasionally modify text that is already correctly spelled.
Research
For details on the training methodology, dataset preparation, evaluation, results, limitations, and comparisons with other models, please refer to the LMSpell paper:
LMSpell: Spell Correction with Pre-Trained Language Models
A. Gunathilake, N. Karunarathna, T. Bandaranayake, S. Ranathunga, N. de Silva and N. Jayatilleke, "LMSpell: Spell Correction with Pre-Trained Language Models," 2026 Moratuwa Engineering Research Conference (MERCon), Moratuwa, Sri Lanka, 2026, pp. 503-508.
- Paper: https://doi.org/10.1109/MERCon71835.2026.11691371
- Extended Preprint: https://arxiv.org/html/2512.05414v3
Citation
@INPROCEEDINGS{11691371,
author={Gunathilake, Akesh and Karunarathna, Nadil and Bandaranayake, Tharusha and Ranathunga, Surangika and de Silva, Nisansa and Jayatilleke, Nevidu},
booktitle={2026 Moratuwa Engineering Research Conference (MERCon)},
title={LMSpell: Spell Correction with Pre-Trained Language Models},
year={2026},
pages={503-508},
doi={10.1109/MERCon71835.2026.11691371}
}
- Downloads last month
- 29