Korean

🏯 Hanok LLM 1.0

Korean-first modular language model stack

🌐 Languages: English Β· ν•œκ΅­μ–΄

A clean, native PyTorch Transformer implementation for Korean LLM experimentation, training, inference, checkpointing and data workflows β€” built for commercialization.

Python 3.10+ PyTorch Transformers Datasets Windows 11 CUDA License

Hanok LLM overview

20 Layers Β· 1,920 Hidden Β· 10 Heads Β· SwiGLU 4,800 Β· ~1.09B Parameters

Repository: https://github.com/seoan1024/Hanok-LLM.git
Research base: https://github.com/seoan1024/Korean-llm.git


✦ What is Hanok?

Hanok LLM is built on the foundation of continuous research from the Korean-llm project.
It is a modular, production-oriented Korean language model stack designed with full commercialization as the clear and explicit goal.

The entire system lives under a maintainable Python package (src/hanok/) with strict separation of concerns. Every major subsystem β€” model, data, training, inference, evaluation, export and monitoring β€” can be inspected, extended or replaced independently.

Layer Responsibility
Model Native decoder-only Transformer (KoreanLLM) with RMSNorm, RoPE, SwiGLU, tied embeddings and KV cache
Data Hugging Face dataset loading, Parquet caching, Korean-oriented field parsing and response-only SFT masking
Training Pretrain β†’ SFT orchestration, gradient accumulation, cosine schedule, checkpoint resume, optional GUI
Inference Native .pth loading + temperature / top-k / top-p / repetition-penalty generation
Evaluation Perplexity utility
Export Hugging Face / Ollama helpers (full conversion + weights + benchmarks coming with commercial release)
UI Optional real-time training monitor GUI

Build a Korean LLM stack you can actually inspect β€” and ship.
Native PyTorch. Explicit defaults. Local checkpoints. Practical Windows + CUDA workflow.

Trained model weights, official benchmark scores, and complete Hugging Face + Ollama (GGUF) conversion support will be released together when the commercial-ready checkpoint is published.


🧠 Architecture in depth

Hanok architecture

Core configuration

Parameter Value Notes
Transformer blocks 20 n_layers
Hidden dimension 1,920 dim
Attention heads 10 n_heads
Head dimension 192 dim // n_heads
SwiGLU intermediate 4,800 dim Γ— 2.5
Normalization RMSNorm eps=1e-6
Positional encoding RoPE ΞΈ = 10,000
Residual style Pre-norm Attention & FFN
Embedding / output Tied weights output.weight = embed.weight
KV cache Supported Per-layer, detached
Default training seq length 512 Configurable
Max sequence length 2,048 RoPE buffer Γ— 2
Vocab size (default) 128,256 Follows loaded tokenizer
Instantiated parameters ~1.09B Depends on final vocab size

Building blocks

RMSNorm
Root-mean-square normalization without mean centering. Lightweight and stable for deep residual stacks.

RoPE (Rotary Positional Embeddings)
Applied to both queries and keys. Precomputed cosine/sine tables are registered as buffers and sliced by absolute position (supports KV-cache incremental decoding).

Attention

  • Multi-head self-attention with F.scaled_dot_product_attention
  • Causal masking by default (is_causal=True when no external mask is given)
  • Optional external attention mask
  • KV cache concatenation along the sequence dimension for efficient generation

SwiGLU Feed-Forward
Three linear projections (w1, w2, w3) with no bias:
w2(SiLU(w1(x)) * w3(x)). Intermediate size is dim Γ— 2.5 = 4,800.

TransformerBlock
Pre-norm residual:
x = x + Attention(RMSNorm(x))
x = x + SwiGLU(RMSNorm(x))

KoreanLLM

  • Embedding β†’ 20 Γ— TransformerBlock β†’ final RMSNorm β†’ tied output projection
  • Weight initialization: normal(std=0.02) for Linear and Embedding
  • Training path uses torch.utils.checkpoint (non-reentrant) for memory efficiency
  • Loss: token-level cross-entropy with ignore_index=-100, reduced by valid target count (handles empty-target batches safely)

Design philosophy

The architecture deliberately stays close to modern decoder-only practices (RMSNorm + RoPE + SwiGLU + tied embeddings) while remaining fully transparent. There are no opaque custom CUDA kernels or closed-source components β€” every tensor operation is visible in pure PyTorch.


⚑ Training stack

Training pipeline

Default TrainingConfig

stage                      auto | pretrain | sft
batch_size                 2
accumulation_steps         8          # effective batch = 16
learning_rate              5e-5
warmup_steps               200
scheduler                  Cosine
weight_decay               0.1
max_steps                  50,000     # general
pretrain_steps             100,000
sft_steps                  10,000
eval_interval              5,000
max_seq_len                512
num_workers                4
use_bfloat16               True
use_8bit_optimizer         True       # AdamW8bit β†’ AdamW fallback
validation_split           0.02
validation_batches         32
enable_gui                 True
seed                       42
checkpoint_dir             checkpoints
force_redownload           False
stream_datasets            False
samples_per_dataset        None
resume_from_checkpoint     None

Training stages

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚   DATA SOURCES   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ CACHE / PARQUET  β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚       PRETRAIN        β”‚
                 β”‚  KorMix / custom data β”‚
                 β”‚  (100k steps default) β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚ checkpoint
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚          SFT          β”‚
                 β”‚ KuLLM + KoAlpaca      β”‚
                 β”‚ (10k steps default)   β”‚
                 β”‚ response-only labels  β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  NATIVE .PTH    β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • auto mode: runs full pretraining first, then starts SFT from the latest pretrain checkpoint (fresh optimizer & schedule for SFT).
  • Checkpoint resume: supported for both stages.
  • SFT label masking: only response tokens contribute to the loss; prompt tokens are set to -100.
  • Validation: 2% random split, evaluated every eval_interval steps (up to validation_batches).
  • GUI: optional real-time loss monitor (enable_gui=True by default; disable with --no-gui).

Optimizer & scheduler

  • Primary: AdamW8bit (bitsandbytes) when available
  • Fallback: standard AdamW
  • Learning-rate schedule: Cosine annealing with linear warmup (200 steps)

Memory & precision

  • BF16 autocast on CUDA when use_bfloat16=True
  • Gradient checkpointing active during training
  • Pin memory for DataLoader on CUDA devices

πŸ–₯️ Recommended workstation

Windows 11 + RTX 5090 profile
Component Recommended profile
OS Windows 11
GPU NVIDIA GeForce RTX 5090 Laptop GPU (or comparable)
VRAM 12 GB+
System RAM 64 GB+
CPU High-performance multi-core
Python 3.10+ (documented on 3.11)
Precision BF16 on supported CUDA hardware
Storage Fast SSD (datasets + checkpoints grow quickly)

Memory note
12 GB VRAM is a practical minimum for the documented configuration.
Actual usage depends on batch size, sequence length, optimizer state, whether gradient checkpointing is active, and whether you are training or only running inference.
Longer sequences or larger effective batches will require more VRAM or reduced accumulation.


πŸ§ͺ Validation & roadmap status

Validation status
Claim Status
Local Windows 11 / RTX 5090 Laptop GPU testing βœ…
Model construction & native inference βœ…
Training / checkpoint / resume workflow βœ…
Architecture / data / move-parity tests βœ…
Model Card & Data Card included βœ…
Commercialization as the explicit goal βœ…
Trained weights + official benchmark scores πŸ”œ Release
Full Hugging Face conversion & publishing πŸ”œ Release
Ollama / GGUF conversion πŸ”œ Release

🧩 Technology stack

Technology stack
Category Libraries / tools
Core Python 3.10+, PyTorch 2.x
Tokenization Hugging Face Transformers (beomi/Llama-3-Open-Ko-8B)
Data Hugging Face Datasets, PyArrow, Pandas
Optimization bitsandbytes (optional 8-bit AdamW)
Monitoring matplotlib + optional GUI monitor
Packaging setuptools, pyproject.toml, CLI entrypoint hanok

πŸ“¦ Project structure

Hanok-LLM/
β”œβ”€β”€ src/hanok/
β”‚   β”œβ”€β”€ model/
β”‚   β”‚   └── architecture.py          # KoreanLLM (RMSNorm Β· RoPE Β· SwiGLU Β· KV cache)
β”‚   β”œβ”€β”€ data/
β”‚   β”‚   └── legacy.py                # DatasetManager, LocalKoreanDataset, collate, Parquet cache
β”‚   β”œβ”€β”€ training/
β”‚   β”‚   β”œβ”€β”€ config.py                # TrainingConfig dataclass (all defaults)
β”‚   β”‚   β”œβ”€β”€ trainer.py               # Main orchestration (auto / pretrain / sft)
β”‚   β”‚   β”œβ”€β”€ optimizer.py             # AdamW / AdamW8bit selection
β”‚   β”‚   β”œβ”€β”€ scheduler.py             # Cosine + warmup
β”‚   β”‚   β”œβ”€β”€ checkpoint.py            # Save / load / find_latest
β”‚   β”‚   └── reproducibility.py       # Seed & distributed helpers
β”‚   β”œβ”€β”€ inference/
β”‚   β”‚   β”œβ”€β”€ loader.py                # Checkpoint + tokenizer loading
β”‚   β”‚   └── generation.py            # Autoregressive generation with KV cache
β”‚   β”œβ”€β”€ evaluation/
β”‚   β”‚   └── perplexity.py
β”‚   β”œβ”€β”€ export/
β”‚   β”‚   β”œβ”€β”€ huggingface.py           # HF bundle helper
β”‚   β”‚   └── ollama.py                # Ollama Modelfile helper (full support in release)
β”‚   β”œβ”€β”€ ui/
β”‚   β”‚   └── monitor.py               # Optional training GUI
β”‚   └── cli.py                       # `hanok` entrypoint (train / infer)
β”œβ”€β”€ assets/svg/                      # README visual kit
β”œβ”€β”€ MODEL_CARD.md
β”œβ”€β”€ DATA_CARD.md
β”œβ”€β”€ CONTRIBUTING.md
β”œβ”€β”€ SECURITY.md
β”œβ”€β”€ pyproject.toml
└── requirements.txt

πŸ› οΈ Installation (Windows 11)

# 1. Create / activate a Python 3.10+ environment
python --version

# 2. Install PyTorch with the CUDA build that matches your driver
#    β†’ https://pytorch.org (select CUDA version carefully)

# 3. Install the project in editable mode
python -m pip install -e .

# 4. Optional: 8-bit optimizer support
python -m pip install -e ".[8bit]"

# 5. Verify the CLI is available
hanok --help

Dependencies (from requirements.txt / pyproject.toml)

torch
transformers
datasets
accelerate
bitsandbytes          # optional, for 8-bit AdamW
pyarrow
pandas
requests
tqdm
matplotlib
setuptools
wheel

🎯 Quick start

SFT smoke test (fast sanity check)

hanok train --stage sft --no-gui --stream-datasets --samples-per-dataset 20 --max-steps 10

Full SFT with defaults

hanok train --stage sft --no-gui

Pretraining

hanok train --stage pretrain --no-gui

Auto pipeline (pretrain β†’ SFT)

hanok train --no-gui

The automatic path first completes pretraining (or resumes the latest pretrain checkpoint), then starts SFT from that checkpoint with a fresh optimizer and learning-rate schedule.

Useful training flags

Flag Description
--stage {auto,pretrain,sft} Training stage
--no-gui Disable the optional training monitor
--stream-datasets Stream instead of full download
--samples-per-dataset N Limit examples per dataset (smoke / debug)
--max-steps N Override max training steps
--batch-size N Override batch size
--resume-from-checkpoint Path or latest
--force-redownload Re-download datasets
--checkpoint-dir PATH Where to write checkpoints

πŸ’¬ Inference

hanok infer `
  --checkpoint checkpoints/korean_llm_00010.pth `
  --prompt "μ•ˆλ…•ν•˜μ„Έμš”. λ„ˆλŠ” λˆ„κ΅¬λ‹ˆ?" `
  --max-tokens 128

Generation controls

Flag Default Description
--temperature 0.6 Softmax temperature (> 0 required)
--top-k 40 Keep only top-k logits
--top-p 0.95 Nucleus sampling
--repetition-penalty 1.3 Penalize already-generated tokens
--max-seq-len 512 Context window used at load time
--device auto cuda / cpu

How generation works

  1. Prompt is wrapped in the instruction format:
    ### 질문: {prompt}\n### 응닡:
  2. Tokens are generated autoregressively with a KV cache.
  3. Temperature scaling β†’ optional top-k filtering β†’ softmax β†’ optional top-p nucleus filtering β†’ multinomial sampling.
  4. Repetition penalty is applied to previously seen tokens.
  5. Stops on EOS or when the sequence length safety limit is reached.
  6. Only the text after ### 응닡: is returned.

πŸ—‚οΈ Data sources & pipeline

Default datasets

Stage Dataset Config Split Notes
Pretraining AdaMLLab/KorMix minhash_deduped train Text corpus
SFT nlpai-lab/kullm-v2 default train Instruction
SFT beomi/KoAlpaca-v1.1a default train Instruction

Tokenizer: beomi/Llama-3-Open-Ko-8B
A dedicated pad token <|pad|> is added when the tokenizer does not already provide a distinct pad token.

Data handling details

  • Hugging Face downloads are cached under ./datasets/cache.
  • Downloaded rows are stored as Parquet and recorded in a manifest.
  • Field parsing supports multiple common schemas:
    text / content / document / body,
    instruction / input / output,
    question / response, question / answer, prompt / response.
  • SFT mode masks the prompt portion; only response tokens contribute to the loss.
  • EOS is appended; PAD positions are ignored via ignore_index=-100.
  • Optional streaming and samples_per_dataset limits are available for rapid iteration.

Always verify the original dataset licenses and terms of use before training.
This repository does not redistribute or re-license third-party data.


πŸ“€ Export & commercial release roadmap

Full Hugging Face conversion, Ollama (GGUF) conversion,
trained model checkpoint weights, and official benchmark scores
will be provided together in a future release as part of the commercialization roadmap.

The current helpers in the codebase are preparatory scaffolding.
They will be completed and documented when the official weight package is published.


πŸ”¬ Model Card & Data Card

Document Purpose
MODEL_CARD.md Architecture summary, intended use, limitations, tokenizer notes
DATA_CARD.md Dataset sources, preprocessing behavior, reproducibility caveats
CONTRIBUTING.md How to contribute
SECURITY.md Security reporting process

🧭 Design principles

  1. Inspectability β€” every layer is ordinary PyTorch; no black-box kernels.
  2. Explicit defaults β€” training hyperparameters are declared in one place (TrainingConfig).
  3. Local-first β€” checkpoints are native .pth files you can load without a remote registry.
  4. Windows + CUDA practicality β€” the documented path is a real workstation workflow, not a theoretical Linux-only lab.
  5. Commercial trajectory β€” the architecture and tooling are built so that a production release (weights + benchmarks + HF/Ollama) can land cleanly on top of this codebase.

πŸ“œ License

See LICENSE for the repository license.
Third-party datasets, tokenizers and dependencies carry their own terms. Always review them before any commercial use.


🏯 HANOK · KOREAN LLM ENGINEERING

Inspectable architecture Β· Practical training Β· Native inference Β· Windows/CUDA workflow Β· Built for commercialization

Repo: Hanok-LLM Β· Research base: Korean-llm

Hanok closing panel
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train seoan1024/Hanok-LLM