Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ποΈ
Building on HF
83262.9
TFLOPS
Huang Liang Hsun
PRO
lianghsun
102
13
261
Follow
cslin's profile picture
fastyangmh's profile picture
albertkingdom's profile picture
143 followers
Β·
29 following
https://www.lianghsun.dev
lianghsun
lianghsunhuang
AI & ML interests
Founder of π§ππΆπ»πΈπΉπ² ππ. Focused on applying deep learning in legal and scientific domains, with expertise in NLP and model fine-tuning.
Recent Activity
posted
an
update
about 9 hours ago
πΉπΌ Releasing https://huggingface.co/lianghsun/tw-tokenizer-v1 β a tokenizer trained from scratch for Traditional Chinese (Taiwan). **46% better Chinese compression than Qwen3.8-27B with 81% of its vocab (201K vs 248K), and English essentially untouched (4.657 vs 4.674 chars/token).** The gain isn't from the regex β it's the corpus. Qwen carries **27,364 Simplified-only multi-char tokens**, 11% of its vocab, dead weight for Traditional Chinese. Train on pure Traditional and that waste never appears. Recent work is skeptical that compression predicts quality (Lotz et al. 2025 measured Ο = β0.59), so we validated two levels deeper: **Segmentation** β boundary hit rate against jieba: **85.6%** vs Qwen's 77.8%. Single-character tokens: **17.6%** vs 41.7%. ``` ε°ζ₯η΄ ι€γηΉθ³ͺζηΆε ¬εε―©ζ₯εͺε ours: ['ε°ζ₯η΄ ι€', 'γ', 'ηΉθ³ͺ', 'ζηΆ', 'ε ¬ε', 'ε―©ζ₯', 'εͺε'] Qwen: ['ε°ζ₯', 'η΄ ', 'ι€', ...] β γη΄ ι€γsplit mid-word ``` **Downstream** β trained a 270M model from scratch with each tokenizer, compared bits-per-character (the only metric fair across tokenizers). At equal compute: **4.434 vs 4.591**, a 3.4% win β with 13% fewer parameters. Same token budget means our model saw 440M characters vs 308M: **43% more data for the same compute**. Also: 6-char cap on pure-CJK tokens (long tokens obscure orthographic info β Haslett, CL 2025), NFC not NFKC, 1,024 reserved tokens. Known limits (weak TΓ’i-lΓ΄ support, small-scale downstream validation, vocab sweep hadn't flattened) are in the card. π https://huggingface.co/lianghsun/tw-tokenizer-v1
updated
a dataset
about 10 hours ago
lianghsun/fineweb-zhtw
published
a dataset
about 10 hours ago
lianghsun/fineweb-zhtw
View all activity
Organizations
lianghsun
's models
30
Sort:Β Recently updated
lianghsun/tw-tokenizer-v1
Updated
about 13 hours ago
lianghsun/fineweb-2-edu-zhtw-classifier
Text Classification
β’
Updated
15 days ago
lianghsun/peptide
Updated
16 days ago
lianghsun/pii-tool-bundle
Token Classification
β’
Updated
29 days ago
lianghsun/Llama-3.2-Taiwan-3B-Instruct
Text Generation
β’
4B
β’
Updated
May 4
β’
71
β’
28
lianghsun/Llama-3.2-Taiwan-Legal-3B-Instruct
Text Generation
β’
3B
β’
Updated
May 4
β’
13
lianghsun/Llama-3.2-Taiwan-Legal-1B-Instruct
Text Generation
β’
1B
β’
Updated
May 4
β’
5
lianghsun/win98-gemma-270m-onnx
Text Generation
β’
Updated
May 4
β’
12
lianghsun/TangYin
Text Generation
β’
0.4B
β’
Updated
May 4
lianghsun/privacy-filter-tw
Token Classification
β’
1B
β’
Updated
May 4
β’
23
lianghsun/Marble-3B-Instruct
Text Generation
β’
3B
β’
Updated
May 4
β’
2
lianghsun/Marble-3B
Text Generation
β’
3B
β’
Updated
May 4
β’
5
lianghsun/Llama-3.2-Taiwan-3B-Instruct-GGUF
Text Generation
β’
4B
β’
Updated
May 4
β’
670
β’
11
lianghsun/Llama-3.2-Taiwan-3B
Text Generation
β’
4B
β’
Updated
May 4
β’
1
β’
30
lianghsun/Llama-3.2-Taiwan-1B-Instruct
Text Generation
β’
1B
β’
Updated
May 4
β’
3
lianghsun/Llama-3.2-Taiwan-1B
Text Generation
β’
1B
β’
Updated
May 4
β’
6
lianghsun/Llama-3.2-3B-F1-Reasoning-Instruct
Text Generation
β’
4B
β’
Updated
May 4
lianghsun/Llama-3.2-3B-F1-Instruct
Text Generation
β’
4B
β’
Updated
May 4
lianghsun/Llama-3.2-3B-F1-Base
Text Generation
β’
4B
β’
Updated
May 4
β’
1
lianghsun/Llama-3.1-DeepFox-70B-Instruct
Text Generation
β’
Updated
May 4
lianghsun/keyboard-warrior
Text Generation
β’
0.4B
β’
Updated
May 4
β’
15
β’
2
lianghsun/gemma-3-tw-270m-thinking
Text Generation
β’
0.4B
β’
Updated
May 4
lianghsun/gemma-3-tw-270m-it
Text Generation
β’
0.4B
β’
Updated
May 4
β’
6
lianghsun/gemma-3-tw-270m
Text Generation
β’
0.4B
β’
Updated
May 4
β’
2
β’
1
lianghsun/fineweb-edu-zhtw-classifier
Text Classification
β’
0.3B
β’
Updated
May 4
β’
1
lianghsun/F1-24B-Reasoner
Text Generation
β’
24B
β’
Updated
May 4
lianghsun/F1-24B-Instruct-Cybersecurity
Text Generation
β’
24B
β’
Updated
May 4
β’
1
lianghsun/F1-24B-Instruct
Text Generation
β’
24B
β’
Updated
May 4
β’
1
lianghsun/F1-24B-Base
Text Generation
β’
24B
β’
Updated
May 4
lianghsun/embeddinggemma-300m-cot-router
Feature Extraction
β’
Updated
May 4