AI & ML interests

None defined yet.

Recent Activity

lianghsun 
posted an update about 2 months ago
view post
Post
328
🇹🇼 Releasing https://huggingface.co/lianghsun/tw-tokenizer-v1 — a tokenizer trained from scratch for Traditional Chinese (Taiwan).

**46% better Chinese compression than Qwen3.8-27B with 81% of its vocab (201K vs 248K), and English essentially untouched (4.657 vs 4.674 chars/token).**

The gain isn't from the regex — it's the corpus. Qwen carries **27,364 Simplified-only multi-char tokens**, 11% of its vocab, dead weight for Traditional Chinese. Train on pure Traditional and that waste never appears.

Recent work is skeptical that compression predicts quality (Lotz et al. 2025 measured ρ = −0.59), so we validated two levels deeper:

**Segmentation** — boundary hit rate against jieba: **85.6%** vs Qwen's 77.8%. Single-character tokens: **17.6%** vs 41.7%.

專業素養、特質或經公告審查優勝
  ours: ['專業素養', '、', '特質', '或經', '公告', '審查', '優勝']
  Qwen: ['專業', '素', '養', ...]     ← 「素養」split mid-word


**Downstream** — trained a 270M model from scratch with each tokenizer, compared bits-per-character (the only metric fair across tokenizers). At equal compute: **4.434 vs 4.591**, a 3.4% win — with 13% fewer parameters. Same token budget means our model saw 440M characters vs 308M: **43% more data for the same compute**.

Also: 6-char cap on pure-CJK tokens (long tokens obscure orthographic info — Haslett, CL 2025), NFC not NFKC, 1,024 reserved tokens.

Known limits (weak Tâi-lô support, small-scale downstream validation, vocab sweep hadn't flattened) are in the card.

👉 https://huggingface.co/lianghsun/tw-tokenizer-v1
Felguk 
posted an update over 1 year ago
view post
Post
2316
Where gone streamlit in huggingface?
  • 3 replies
·
lianghsun 
posted an update over 1 year ago
view post
Post
4757

With the arrival of Twinkle April — Twinkle AI’s annual open-source celebration held every April — our community is excited to unveil its very first project:

📊 Twinkle Eval (https://github.com/ai-twinkle/Eval), a next-generation evaluation tool led by our contributor @tedslin .

Unlike traditional evaluation tools like iKala’s ievals (https://github.com/ikala-ai/ievals), which can only evaluate language models (LMs) one sample at a time, Twinkle Eval is designed with Large Reasoning Models (LRMs) in mind. As reasoning time increases with more complex models, traditional tools become increasingly inefficient 😲 — for example, evaluating LRMs on the ikala/tmmluplus benchmark could take *
half a day without finishing.

One question we were especially curious about:
Does shuffling multiple-choice answer order impact model accuracy? 🤔
→ See: "Change Answer Order Can Decrease MMLU Accuracy" – arXiv:2406.19470v1

To address these challenges, Twinkle Eval brings three key innovations to the table:

1️⃣ Parallelized evaluation of samples
2️⃣ Multi-round testing for stability
3️⃣ Randomized answer order to test robustness

After running experiments, we observed that Twinkle Eval can speed up evaluation by up to 15× 🚀🚀. Interestingly, most models scored slightly lower under the 2️⃣3️⃣ test settings compared to their claimed performance — suggesting further benchmarking is needed.

This framework also comes with additional tunable parameters and detailed logging of LM behavior per question — perfect for those who want to dive deeper. 😆

If you find Twinkle Eval useful, please ⭐ the project and help spread the word 🤗
  • 5 replies
·
Baicai003 
updated a Space over 1 year ago