- π― Hanok LLM 1.0
- β¦ What is Hanok?
- π§ Architecture in depth
- β‘ Training stack
- π₯οΈ Recommended workstation
- π§ͺ Validation & roadmap status
- π§© Technology stack
- π¦ Project structure
- π οΈ Installation (Windows 11)
- π― Quick start
- π¬ Inference
- ποΈ Data sources & pipeline
- π€ Export & commercial release roadmap
- π¬ Model Card & Data Card
- π§ Design principles
- π License
- β¦ What is Hanok?
π― Hanok LLM 1.0
Korean-first modular language model stack
π Languages: English Β· νκ΅μ΄
A clean, native PyTorch Transformer implementation for Korean LLM experimentation, training, inference, checkpointing and data workflows β built for commercialization.
20 Layers Β· 1,920 Hidden Β· 10 Heads Β· SwiGLU 4,800 Β· ~1.09B Parameters
Repository: https://github.com/seoan1024/Hanok-LLM.git
Research base: https://github.com/seoan1024/Korean-llm.git
β¦ What is Hanok?
Hanok LLM is built on the foundation of continuous research from the Korean-llm project.
It is a modular, production-oriented Korean language model stack designed with full commercialization as the clear and explicit goal.
The entire system lives under a maintainable Python package (src/hanok/) with strict separation of concerns. Every major subsystem β model, data, training, inference, evaluation, export and monitoring β can be inspected, extended or replaced independently.
| Layer | Responsibility |
|---|---|
| Model | Native decoder-only Transformer (KoreanLLM) with RMSNorm, RoPE, SwiGLU, tied embeddings and KV cache |
| Data | Hugging Face dataset loading, Parquet caching, Korean-oriented field parsing and response-only SFT masking |
| Training | Pretrain β SFT orchestration, gradient accumulation, cosine schedule, checkpoint resume, optional GUI |
| Inference | Native .pth loading + temperature / top-k / top-p / repetition-penalty generation |
| Evaluation | Perplexity utility |
| Export | Hugging Face / Ollama helpers (full conversion + weights + benchmarks coming with commercial release) |
| UI | Optional real-time training monitor GUI |
Build a Korean LLM stack you can actually inspect β and ship.
Native PyTorch. Explicit defaults. Local checkpoints. Practical Windows + CUDA workflow.
Trained model weights, official benchmark scores, and complete Hugging Face + Ollama (GGUF) conversion support will be released together when the commercial-ready checkpoint is published.
π§ Architecture in depth
Core configuration
| Parameter | Value | Notes |
|---|---|---|
| Transformer blocks | 20 | n_layers |
| Hidden dimension | 1,920 | dim |
| Attention heads | 10 | n_heads |
| Head dimension | 192 | dim // n_heads |
| SwiGLU intermediate | 4,800 | dim Γ 2.5 |
| Normalization | RMSNorm | eps=1e-6 |
| Positional encoding | RoPE | ΞΈ = 10,000 |
| Residual style | Pre-norm | Attention & FFN |
| Embedding / output | Tied weights | output.weight = embed.weight |
| KV cache | Supported | Per-layer, detached |
| Default training seq length | 512 | Configurable |
| Max sequence length | 2,048 | RoPE buffer Γ 2 |
| Vocab size (default) | 128,256 | Follows loaded tokenizer |
| Instantiated parameters | ~1.09B | Depends on final vocab size |
Building blocks
RMSNorm
Root-mean-square normalization without mean centering. Lightweight and stable for deep residual stacks.
RoPE (Rotary Positional Embeddings)
Applied to both queries and keys. Precomputed cosine/sine tables are registered as buffers and sliced by absolute position (supports KV-cache incremental decoding).
Attention
- Multi-head self-attention with
F.scaled_dot_product_attention - Causal masking by default (
is_causal=Truewhen no external mask is given) - Optional external attention mask
- KV cache concatenation along the sequence dimension for efficient generation
SwiGLU Feed-Forward
Three linear projections (w1, w2, w3) with no bias:w2(SiLU(w1(x)) * w3(x)). Intermediate size is dim Γ 2.5 = 4,800.
TransformerBlock
Pre-norm residual:x = x + Attention(RMSNorm(x))x = x + SwiGLU(RMSNorm(x))
KoreanLLM
- Embedding β 20 Γ TransformerBlock β final RMSNorm β tied output projection
- Weight initialization: normal(std=0.02) for Linear and Embedding
- Training path uses
torch.utils.checkpoint(non-reentrant) for memory efficiency - Loss: token-level cross-entropy with
ignore_index=-100, reduced by valid target count (handles empty-target batches safely)
Design philosophy
The architecture deliberately stays close to modern decoder-only practices (RMSNorm + RoPE + SwiGLU + tied embeddings) while remaining fully transparent. There are no opaque custom CUDA kernels or closed-source components β every tensor operation is visible in pure PyTorch.
β‘ Training stack
Default TrainingConfig
stage auto | pretrain | sft
batch_size 2
accumulation_steps 8 # effective batch = 16
learning_rate 5e-5
warmup_steps 200
scheduler Cosine
weight_decay 0.1
max_steps 50,000 # general
pretrain_steps 100,000
sft_steps 10,000
eval_interval 5,000
max_seq_len 512
num_workers 4
use_bfloat16 True
use_8bit_optimizer True # AdamW8bit β AdamW fallback
validation_split 0.02
validation_batches 32
enable_gui True
seed 42
checkpoint_dir checkpoints
force_redownload False
stream_datasets False
samples_per_dataset None
resume_from_checkpoint None
Training stages
ββββββββββββββββββββ
β DATA SOURCES β
ββββββββββ¬ββββββββββ
β
ββββββββββΌββββββββββ
β CACHE / PARQUET β
ββββββββββ¬ββββββββββ
β
βββββββββββββΌββββββββββββ
β PRETRAIN β
β KorMix / custom data β
β (100k steps default) β
βββββββββββββ¬ββββββββββββ
β checkpoint
βββββββββββββΌββββββββββββ
β SFT β
β KuLLM + KoAlpaca β
β (10k steps default) β
β response-only labels β
βββββββββββββ¬ββββββββββββ
β
ββββββββββΌβββββββββ
β NATIVE .PTH β
βββββββββββββββββββ
automode: runs full pretraining first, then starts SFT from the latest pretrain checkpoint (fresh optimizer & schedule for SFT).- Checkpoint resume: supported for both stages.
- SFT label masking: only response tokens contribute to the loss; prompt tokens are set to
-100. - Validation: 2% random split, evaluated every
eval_intervalsteps (up tovalidation_batches). - GUI: optional real-time loss monitor (
enable_gui=Trueby default; disable with--no-gui).
Optimizer & scheduler
- Primary: AdamW8bit (bitsandbytes) when available
- Fallback: standard AdamW
- Learning-rate schedule: Cosine annealing with linear warmup (200 steps)
Memory & precision
- BF16 autocast on CUDA when
use_bfloat16=True - Gradient checkpointing active during training
- Pin memory for DataLoader on CUDA devices
π₯οΈ Recommended workstation
| Component | Recommended profile |
|---|---|
| OS | Windows 11 |
| GPU | NVIDIA GeForce RTX 5090 Laptop GPU (or comparable) |
| VRAM | 12 GB+ |
| System RAM | 64 GB+ |
| CPU | High-performance multi-core |
| Python | 3.10+ (documented on 3.11) |
| Precision | BF16 on supported CUDA hardware |
| Storage | Fast SSD (datasets + checkpoints grow quickly) |
Memory note
12 GB VRAM is a practical minimum for the documented configuration.
Actual usage depends on batch size, sequence length, optimizer state, whether gradient checkpointing is active, and whether you are training or only running inference.
Longer sequences or larger effective batches will require more VRAM or reduced accumulation.
π§ͺ Validation & roadmap status
| Claim | Status |
|---|---|
| Local Windows 11 / RTX 5090 Laptop GPU testing | β |
| Model construction & native inference | β |
| Training / checkpoint / resume workflow | β |
| Architecture / data / move-parity tests | β |
| Model Card & Data Card included | β |
| Commercialization as the explicit goal | β |
| Trained weights + official benchmark scores | π Release |
| Full Hugging Face conversion & publishing | π Release |
| Ollama / GGUF conversion | π Release |
π§© Technology stack
| Category | Libraries / tools |
|---|---|
| Core | Python 3.10+, PyTorch 2.x |
| Tokenization | Hugging Face Transformers (beomi/Llama-3-Open-Ko-8B) |
| Data | Hugging Face Datasets, PyArrow, Pandas |
| Optimization | bitsandbytes (optional 8-bit AdamW) |
| Monitoring | matplotlib + optional GUI monitor |
| Packaging | setuptools, pyproject.toml, CLI entrypoint hanok |
π¦ Project structure
Hanok-LLM/
βββ src/hanok/
β βββ model/
β β βββ architecture.py # KoreanLLM (RMSNorm Β· RoPE Β· SwiGLU Β· KV cache)
β βββ data/
β β βββ legacy.py # DatasetManager, LocalKoreanDataset, collate, Parquet cache
β βββ training/
β β βββ config.py # TrainingConfig dataclass (all defaults)
β β βββ trainer.py # Main orchestration (auto / pretrain / sft)
β β βββ optimizer.py # AdamW / AdamW8bit selection
β β βββ scheduler.py # Cosine + warmup
β β βββ checkpoint.py # Save / load / find_latest
β β βββ reproducibility.py # Seed & distributed helpers
β βββ inference/
β β βββ loader.py # Checkpoint + tokenizer loading
β β βββ generation.py # Autoregressive generation with KV cache
β βββ evaluation/
β β βββ perplexity.py
β βββ export/
β β βββ huggingface.py # HF bundle helper
β β βββ ollama.py # Ollama Modelfile helper (full support in release)
β βββ ui/
β β βββ monitor.py # Optional training GUI
β βββ cli.py # `hanok` entrypoint (train / infer)
βββ assets/svg/ # README visual kit
βββ MODEL_CARD.md
βββ DATA_CARD.md
βββ CONTRIBUTING.md
βββ SECURITY.md
βββ pyproject.toml
βββ requirements.txt
π οΈ Installation (Windows 11)
# 1. Create / activate a Python 3.10+ environment
python --version
# 2. Install PyTorch with the CUDA build that matches your driver
# β https://pytorch.org (select CUDA version carefully)
# 3. Install the project in editable mode
python -m pip install -e .
# 4. Optional: 8-bit optimizer support
python -m pip install -e ".[8bit]"
# 5. Verify the CLI is available
hanok --help
Dependencies (from requirements.txt / pyproject.toml)
torch
transformers
datasets
accelerate
bitsandbytes # optional, for 8-bit AdamW
pyarrow
pandas
requests
tqdm
matplotlib
setuptools
wheel
π― Quick start
SFT smoke test (fast sanity check)
hanok train --stage sft --no-gui --stream-datasets --samples-per-dataset 20 --max-steps 10
Full SFT with defaults
hanok train --stage sft --no-gui
Pretraining
hanok train --stage pretrain --no-gui
Auto pipeline (pretrain β SFT)
hanok train --no-gui
The automatic path first completes pretraining (or resumes the latest pretrain checkpoint), then starts SFT from that checkpoint with a fresh optimizer and learning-rate schedule.
Useful training flags
| Flag | Description |
|---|---|
--stage {auto,pretrain,sft} |
Training stage |
--no-gui |
Disable the optional training monitor |
--stream-datasets |
Stream instead of full download |
--samples-per-dataset N |
Limit examples per dataset (smoke / debug) |
--max-steps N |
Override max training steps |
--batch-size N |
Override batch size |
--resume-from-checkpoint |
Path or latest |
--force-redownload |
Re-download datasets |
--checkpoint-dir PATH |
Where to write checkpoints |
π¬ Inference
hanok infer `
--checkpoint checkpoints/korean_llm_00010.pth `
--prompt "μλ
νμΈμ. λλ λꡬλ?" `
--max-tokens 128
Generation controls
| Flag | Default | Description |
|---|---|---|
--temperature |
0.6 | Softmax temperature (> 0 required) |
--top-k |
40 | Keep only top-k logits |
--top-p |
0.95 | Nucleus sampling |
--repetition-penalty |
1.3 | Penalize already-generated tokens |
--max-seq-len |
512 | Context window used at load time |
--device |
auto | cuda / cpu |
How generation works
- Prompt is wrapped in the instruction format:
### μ§λ¬Έ: {prompt}\n### μλ΅: - Tokens are generated autoregressively with a KV cache.
- Temperature scaling β optional top-k filtering β softmax β optional top-p nucleus filtering β multinomial sampling.
- Repetition penalty is applied to previously seen tokens.
- Stops on EOS or when the sequence length safety limit is reached.
- Only the text after
### μλ΅:is returned.
ποΈ Data sources & pipeline
Default datasets
| Stage | Dataset | Config | Split | Notes |
|---|---|---|---|---|
| Pretraining | AdaMLLab/KorMix |
minhash_deduped |
train | Text corpus |
| SFT | nlpai-lab/kullm-v2 |
default | train | Instruction |
| SFT | beomi/KoAlpaca-v1.1a |
default | train | Instruction |
Tokenizer: beomi/Llama-3-Open-Ko-8B
A dedicated pad token <|pad|> is added when the tokenizer does not already provide a distinct pad token.
Data handling details
- Hugging Face downloads are cached under
./datasets/cache. - Downloaded rows are stored as Parquet and recorded in a manifest.
- Field parsing supports multiple common schemas:
text/content/document/body,instruction/input/output,question/response,question/answer,prompt/response. - SFT mode masks the prompt portion; only response tokens contribute to the loss.
- EOS is appended; PAD positions are ignored via
ignore_index=-100. - Optional streaming and
samples_per_datasetlimits are available for rapid iteration.
Always verify the original dataset licenses and terms of use before training.
This repository does not redistribute or re-license third-party data.
π€ Export & commercial release roadmap
Full Hugging Face conversion, Ollama (GGUF) conversion,
trained model checkpoint weights, and official benchmark scores
will be provided together in a future release as part of the commercialization roadmap.
The current helpers in the codebase are preparatory scaffolding.
They will be completed and documented when the official weight package is published.
π¬ Model Card & Data Card
| Document | Purpose |
|---|---|
MODEL_CARD.md |
Architecture summary, intended use, limitations, tokenizer notes |
DATA_CARD.md |
Dataset sources, preprocessing behavior, reproducibility caveats |
CONTRIBUTING.md |
How to contribute |
SECURITY.md |
Security reporting process |
π§ Design principles
- Inspectability β every layer is ordinary PyTorch; no black-box kernels.
- Explicit defaults β training hyperparameters are declared in one place (
TrainingConfig). - Local-first β checkpoints are native
.pthfiles you can load without a remote registry. - Windows + CUDA practicality β the documented path is a real workstation workflow, not a theoretical Linux-only lab.
- Commercial trajectory β the architecture and tooling are built so that a production release (weights + benchmarks + HF/Ollama) can land cleanly on top of this codebase.
π License
See LICENSE for the repository license.
Third-party datasets, tokenizers and dependencies carry their own terms. Always review them before any commercial use.
π― HANOK Β· KOREAN LLM ENGINEERING
Inspectable architecture Β· Practical training Β· Native inference Β· Windows/CUDA workflow Β· Built for commercialization
Repo: Hanok-LLM Β· Research base: Korean-llm