Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated
Viewer • Updated • 417k • 405 • 1
How to use Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std with NeMo:
import nemo.collections.asr as nemo_asr
asr_model = nemo_asr.models.ASRModel.from_pretrained("Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std")
transcriptions = asr_model.transcribe(["file.wav"])Research checkpoint from the study Lithuanian speech recognition: effects of dialect training and
transcript spelling. Fine-tuned from nvidia/parakeet-tdt-0.6b-v3 on LIEPA-3 telephone speech (420.17 h) + LIEPA-3 dialect speech (82.30 h), dialect clips in standard spelling.
Output is lowercase spoken form without punctuation. This is a study model, not a product model.
| Test | WER |
|---|---|
| Dialect test, dialect-spelling reference | 37.52 |
| Dialect test, standard-spelling reference | 27.06 |
| Dialect test, either spelling accepted | 26.06 |
| LIEPA-3 telephone test | 10.65 |
| FLEURS lt test | 26.87 |
| Common Voice 19 lt test | 27.75 |
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.from_pretrained("Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std")
print(model.transcribe(["audio.wav"])[0].text)
| Run | ltd26/S+D80.std/s1/a1 |
| Recipe | S+D80.std (recipes/S+D80.std.json in the code repo) |
| Splits | splits-v1 |
| Training | 10,000 updates, 192 clips per update, last checkpoint; full spec in config.json |
| Scorer | see registry/scores.csv in the code repo |
Base model
nvidia/parakeet-tdt-0.6b-v3