YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

๐Ÿ‡ฐ๐Ÿ‡ท Korean LLM Advanced v3

License: GPL-3.0 Python 3.9+ PyTorch CUDA Model Size VRAM

ํ•œ๊ตญ์–ด ํŠนํ™” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด๋ชจ๋ธ - ํ’€์Šคํฌ๋ž˜์น˜ ๊ตฌํ˜„ ๋ฐ ์–‘์žํ™” ์ ์šฉ

๐Ÿš€ ํ”„๋กœ์ ํŠธ ๊ฐœ๋ฐœ๊ธฐ

์ฒ˜์Œ์—๋Š” ๊ธฐ์—…๋“ค์˜ ๋ฌด๋ฃŒ ํ•œ๋„๊ฐ€ ๋นก์„ธ์ง€๊ณ  '๋ฐ”์ด๋ธŒ ์ฝ”๋”ฉ'์„ ํ•˜๊ธฐ์—๋Š” ํ•œ๋„์˜ ํ•œ๊ณ„๊ฐ€ ์ฐพ์•„์™”์Šต๋‹ˆ๋‹ค. Ollama๋ฅผ ํ™œ์šฉํ•ด ๋กœ์ปฌ๋กœ ๋Œ๋ ค๋ณด๊ธฐ๋„ ํ–ˆ์ง€๋งŒ, ๋ชจ๋ธ์ด ๋„ˆ๋ฌด ๋ฌด๊ฑฐ์›Œ ์ปดํ“จํ„ฐ๊ฐ€ ๋ฒ„๊ฑฐ์›Œํ–ˆ์ฃ . ๊ทธ๋•Œ ๋ฌธ๋“ '์ด๋Ÿด ๋ฐ”์—” ๋‚ด๊ฐ€ ์ง์ ‘ ๋งŒ๋“ค์–ด๋ณผ๊นŒ?'๋ผ๋Š” ์ƒ๊ฐ์ด ๋“ค์—ˆ์Šต๋‹ˆ๋‹ค. ํ•˜์ง€๋งŒ ์ €๋Š” ์ค‘ํ•™๊ต 2ํ•™๋…„์ด์—ˆ๊ณ  ์ธ๊ณต์ง€๋Šฅ ๋ชจ๋ธ์— ๋Œ€ํ•ด ์•„๋Š” ๊ฒƒ์ด๋ผ๊ณค ๋ชจ๋ธ ํฌ๊ธฐ๋ฅผ ๋‚˜ํƒ€๋‚ด๋Š” 'B(Billion)'๋ผ๋Š” ๊ฐœ๋…์ด ์ „๋ถ€์˜€๊ณ , ๊ธฐ๋ณธ์ด ์•„๋‹ˆ๋ผ ๊ณ ๊ธ‰๊ฐœ๋…์„ ์งค ์ˆ˜ ์—†์—ˆ์ฃ . ๊ฒฐ๊ตญ ํ‰์†Œ์ฒ˜๋Ÿผ ChatGPT์— ๋“ค์–ด๊ฐ€ โ€œ๋‚˜ ๋…์ž์ ์ธ ํ•œ๊ตญ์–ด LLM ๋ชจ๋ธ ๋งŒ๋“ค๋ž˜!โ€๋ผ๋Š” ํ•œ๋งˆ๋””๋ฅผ ๋˜์ง€๋ฉฐ ๋ฌด๋ชจํ•œ ๋„์ „์„ ์‹œ์ž‘ํ–ˆ์Šต๋‹ˆ๋‹ค. GPT์˜ ๋„์›€์„ ๋ฐ›์œผ๋ฉด์„œ๋„ ํ•œ๊ณ„๋Š” ๊ณ„์† ์ฐพ์•„์™”์Šต๋‹ˆ๋‹ค. ์ฒ˜์Œ์—๋Š” ๊ทธ์ € GPT๊ฐ€ ์ค€ ์ฝ”๋“œ ์กฐ๊ฐ๋“ค์„ ๋ชจ์•„ ์ฐจ์› ์˜ค๋ฅ˜(Dimension Error)๊ฐ€ ๋‚˜์ง€ ์•Š๊ธฐ๋งŒ์„ ๋ฐ”๋ผ๋ฉฐ ์—ด์‹ฌํžˆ ๋Œ๋ ค๋ณผ ๋ฟ์ด์—ˆ์Šต๋‹ˆ๋‹ค. ๋ฐ์ดํ„ฐ๋ฅผ ์ˆ˜์ง‘ํ•˜๊ณ  ์ •์ œํ•˜๋ฉฐ ์˜จ์ข…์ผ ์ปดํ“จํ„ฐ ๋ชจ๋‹ˆํ„ฐ๋งŒ ๋ฐ”๋ผ๋ณด์•˜์Šต๋‹ˆ๋‹ค. ๊ทธ๋ ‡๊ฒŒ ํƒœ์–ด๋‚œ ์ œ ์ฒซ ์ž‘ํ’ˆ(v1 ์ด์ „ ๋ฒ„์ „)์€ ์œ„ํ‚คํ”ผ๋””์•„ ๋ฐ์ดํ„ฐ๋กœ ํ•™์Šตํ•œ, ๊ณ ์ž‘ 50M(5์ฒœ๋งŒ ํŒŒ๋ผ๋ฏธํ„ฐ) ํฌ๊ธฐ์˜ ๊ทน์†Œํ˜• ๋ชจ๋ธ์ด์—ˆ์Šต๋‹ˆ๋‹ค. ์ง€๊ธˆ์€ ๋‚จ์•„์žˆ์ง€ ์•Š์ง€๋งŒ์š”. ๋น„๋ก ๋Œ€ํ™”๋Š” ๋ถˆ๊ฐ€๋Šฅํ–ˆ์ง€๋งŒ, ์–ด๋А ์ •๋„ ๋ฌธ๋ฒ•์— ๋งž๋Š” ๋ฌธ์žฅ์„ ๊ตฌ์ƒํ•˜๋Š” ๋ชจ์Šต์„ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ๊ทธ ์ž‘์€ ์„ฑ๊ณต์ด ๋„ˆ๋ฌด ๊ธฐ๋ป์„œ ์ด๋•Œ๋ถ€ํ„ฐ ๋ณธ๊ฒฉ์ ์œผ๋กœ '์ฑ„ํŒ…ํ˜• ๋ชจ๋ธ'์„ ๋งŒ๋“œ๋Š” ๋ฐ๋งŒ ๋ชฐ์ž…ํ–ˆ์Šต๋‹ˆ๋‹ค. ๊ทธ๋ ‡๊ฒŒ v1์˜ ์ตœ์ข… ๋ฒ„์ „์ธ 541M ํฌ๊ธฐ์˜ ๋ชจ๋ธ๊นŒ์ง€ ๋ฐœ์ „์‹œ์ผฐ์Šต๋‹ˆ๋‹ค. ๋ช‡ ๊ฐ€์ง€ ๋ฒ„๊ทธ๊ฐ€ ๋ฐœ๊ฒฌ๋˜์—ˆ์ง€๋งŒ, ์šฐ์„  ๋ฒ„๊ทธ๋ฅผ ํ•ด๊ฒฐํ•œ ๋’ค ๊ณง๋ฐ”๋กœ ๋ชจ๋ธ์˜ ์ฒด๊ธ‰์„ ํ‚ค์šฐ๊ธฐ๋กœ ๊ฒฐ์‹ฌํ–ˆ์Šต๋‹ˆ๋‹ค. ์˜ค๋ฅ˜๋“ค์„ ์ˆ˜์ •ํ•˜๊ณ , ๋ชจ๋ธ ํฌ๊ธฐ๋ฅผ 2๋ฐฐ๊ฐ€๋Ÿ‰ ํ‚ค์›Œ ๋“œ๋””์–ด 1.09B(10์–ต 9์ฒœ๋งŒ ํŒŒ๋ผ๋ฏธํ„ฐ) ํฌ๊ธฐ์˜ ๋ชจ๋ธ์„ ๊ตฌ์ถ•ํ–ˆ์Šต๋‹ˆ๋‹ค. ๋ฐฉํ•™ ๊ธฐ๊ฐ„ ๋‚ด๋‚ด ์‹œ๊ฐ„์ด ๋‚  ๋•Œ๋งˆ๋‹ค ์ปดํ“จํ„ฐ๋ฅผ ์ผœ๊ณ  ํ•™์Šต์„ ๋Œ๋ ธ์Šต๋‹ˆ๋‹ค. ๊ทธ๋ ‡๊ฒŒ ๋์ด ๋ณด์ด์ง€ ์•Š๋˜ ํ•™์Šต์ด ์–ด๋А๋ง 44,000 ์Šคํ…์— ๋„๋‹ฌํ–ˆ์Šต๋‹ˆ๋‹ค. ์„ค๋ ˆ๋Š” ๋งˆ์Œ์œผ๋กœ ํ…Œ์ŠคํŠธ๋ฅผ ์œ„ํ•ด ์ฑ„ํŒ…์ฐฝ์— "์•ˆ๋…•?"์ด๋ผ๊ณ  ์ž…๋ ฅํ–ˆ์Šต๋‹ˆ๋‹ค.

"์•ˆ๋…•ํ•˜์„ธ์š”! ์˜ค๋Š˜์€ ๋ฌด์—‡์„ ๋„์™€๋“œ๋ฆด๊นŒ์š”?"

๋ชจ๋ธ์ด ์˜ฌ๋ฐ”๋ฅธ ๋‹ต๋ณ€์„ ํ™”๋ฉด์— ๋„์šด ๊ทธ ์ˆœ๊ฐ„, ๋ง๋กœ ํ‘œํ˜„ํ•  ์ˆ˜ ์—†์„ ๋งŒํผ ๊ธฐ๋ปค์Šต๋‹ˆ๋‹ค. ํ•˜์ง€๋งŒ ๊ธฐ์จ๋„ ์ž ์‹œ, ๋‹ค๋ฅธ ์งˆ๋ฌธ์„ ๋˜์ง€์ž ์ „ํ˜€ ์—‰๋šฑํ•œ ๋Œ€๋‹ต์„ ์Ÿ์•„๋‚ด๊ธฐ ์‹œ์ž‘ํ–ˆ์Šต๋‹ˆ๋‹ค. AI์™€ ํ•จ๊ป˜ ๋ฐค์ƒˆ ์ฝ”๋“œ๋ฅผ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, ๋ชจ๋ธ์ด ์‚ฌ์šฉ์ž์˜ ์ง€์‹œ์‚ฌํ•ญ์„ ๋ฌด์‹œํ•ด ๋ฒ„๋ฆฌ๋Š” ์น˜๋ช…์ ์ธ ๋ฒ„๊ทธ๊ฐ€ ๋ฐœ์ƒํ•œ ๊ฒƒ์ด์—ˆ์Šต๋‹ˆ๋‹ค. ๊ฐ€์Šด์ด ์•„ํŒ ์ง€๋งŒ, ๋” ์™„๋ฒฝํ•œ ๋ชจ๋ธ์„ ์œ„ํ•ด ์ง€๊ธˆ๊นŒ์ง€ ํ•™์Šตํ•œ ๊ฒฐ๊ณผ๋ฌผ์„ ๊ณผ๊ฐํžˆ ํ๊ธฐํ–ˆ์Šต๋‹ˆ๋‹ค. ๋‚™๋‹ดํ•˜์ง€ ์•Š๊ณ  v2์˜ ๋ฒ„๊ทธ๋ฅผ ์™„์ „ํžˆ ํ•ด๊ฒฐํ•œ ๋’ค, ๋‹ค์Œ ๋ฌธ์ œ์— ๋„์ „ํ–ˆ์Šต๋‹ˆ๋‹ค. 1B ์ฒด๊ธ‰์˜ ๋ชจ๋ธ์€ VRAM์„ ๋ฌด๋ ค 23GB๋‚˜ ์ฐจ์ง€ํ•˜์—ฌ ์ผ๋ฐ˜์ ์ธ ํ™˜๊ฒฝ์—์„œ ๋Œ๋ฆฌ๊ธฐ ๋„ˆ๋ฌด ๋ฌด๊ฑฐ์› ๊ธฐ ๋•Œ๋ฌธ์ž…๋‹ˆ๋‹ค. ์ด๋ฅผ 10GB ์ดํ•˜๋กœ ์ค„์—ฌ๋ณด๊ฒ ๋‹ค๋Š” ๋ชฉํ‘œ๋ฅผ ์„ธ์› ๊ณ , ๋งˆ์นจ๋‚ด v3์—์„œ ์–‘์žํ™”์— ์„ฑ๊ณตํ–ˆ์Šต๋‹ˆ๋‹ค. ๋ฐฉํ•™์ด ๋๋‚ฌ๋‹ค ๋ณด๋‹ˆ ์ œ๊ฐ€ ํ•™์Šตํ•˜๊ธฐ์—๋Š” ์–ด๋ ต์Šต๋‹ˆ๋‹ค. v2 ๋ฒ„๊ทธ ์ดํ›„ ํ•™์Šตํ•œ ๊ฑด ์•„์˜ˆ ์—†๊ณ  ์‹œ๊ฐ„์ด ๋‚ ๋•Œ ๋‹ค์‹œ ํ•™์Šตํ•ด๋ณด๊ฒ ์Šต๋‹ˆ๋‹ค. ์ด ํ”„๋กœ์ ํŠธ๋Š” ์˜ค์ง "๋‚ด ์†์œผ๋กœ ์ง์ ‘ LLM์„ ๋งŒ๋“ค๊ณ  ์‹ถ๋‹ค"๋Š” ๊ณ ์ง‘ ํ•˜๋‚˜๋กœ ์™„์„ฑํ•ด ๋‚ธ, ์ œ ์ธ์ƒ ์ตœ๊ณ ์˜ ์ž‘ํ’ˆ์ž…๋‹ˆ๋‹ค. ์ด ๋ชจ๋ธ์„ ์ž˜ ์‚ฌ์šฉํ•˜์‹œ๊ณ , ๋งˆ์Œ์— ๋“œ์…จ๋‹ค๋ฉด ์Šคํƒ€(โญ) ๋ฒ„ํŠผ ํ•œ ๋ฒˆ์”ฉ ๊ผญ ๋ˆŒ๋Ÿฌ์ฃผ์„ธ์š”! ๊ฐ์‚ฌํ•ฉ๋‹ˆ๋‹ค!

์ฒ˜์Œ๋ถ€ํ„ฐ ๋๊นŒ์ง€ ํ•œ๊ตญ์–ด๋กœ ํ•™์Šต๋œ 1.09B ํŒŒ๋ผ๋ฏธํ„ฐ LLM์œผ๋กœ, VRAM ์ตœ์ ํ™” ๊ธฐ๋ฒ•์„ ์ ๊ทน ํ™œ์šฉํ–ˆ์Šต๋‹ˆ๋‹ค.

๐Ÿ“‹ ์ฃผ์š” ํŠน์ง• โ€ข ๐Ÿš€ ๋น ๋ฅธ ์‹œ์ž‘ โ€ข ๐Ÿ’พ ๊ธฐ์ˆ  ์Šคํƒ โ€ข ๐Ÿ“Š ๋ฒ„์ „ ํžˆ์Šคํ† ๋ฆฌ

์‹ ๊ทœ ๋ฒ„์ „ V4 ์ถœ์‹œ


๐Ÿ“– ๊ฐœ์š”

Korean LLM Advanced v3๋Š” ํ•œ๊ตญ์–ด ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ์— ์ตœ์ ํ™”๋œ ๊ฒฝ๋Ÿ‰ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด๋ชจ๋ธ์ž…๋‹ˆ๋‹ค. ์ œํ•œ๋œ GPU ๋ฉ”๋ชจ๋ฆฌ ํ™˜๊ฒฝ์—์„œ๋„ ํšจ์œจ์ ์œผ๋กœ ํ•™์Šตํ•˜๊ณ  ์ถ”๋ก ํ•  ์ˆ˜ ์žˆ๋„๋ก ์„ค๊ณ„๋˜์—ˆ์Šต๋‹ˆ๋‹ค.

ํ•ต์‹ฌ ๋ชฉํ‘œ

  • โœ… ํ•œ๊ตญ์–ด ํ…์ŠคํŠธ ์ƒ์„ฑ ๋ฐ ์ดํ•ด ๋Šฅ๋ ฅ
  • โœ… VRAM ํšจ์œจ์„ฑ (9GB ๊ธฐ์ค€)
  • โœ… ๋น ๋ฅธ ํ•™์Šต ์†๋„
  • โœ… ์‰ฌ์šด ๋ฐฐํฌ ๋ฐ ํ™œ์šฉ

๐ŸŒŸ ์ฃผ์š” ํŠน์ง•

๐ŸŽฏ ๋ชจ๋ธ ๊ตฌ์กฐ

ํ•ญ๋ชฉ ์„ค๋ช…
๋ชจ๋ธ ํฌ๊ธฐ 1.09B ํŒŒ๋ผ๋ฏธํ„ฐ
์€๋‹‰์ธต ํฌ๊ธฐ 1,920์ฐจ์›
๋ ˆ์ด์–ด ์ˆ˜ 20๊ฐœ
์–ดํ…์…˜ ํ—ค๋“œ 10๊ฐœ
์ตœ๋Œ€ ์‹œํ€€์Šค ๊ธธ์ด 2,048 ํ† ํฐ
์–ดํœ˜์ง‘ ํฌ๊ธฐ ๋™์  (ํ† ํฌ๋‚˜์ด์ € ๊ธฐ์ค€)

๐Ÿ”ง ์ตœ์ ํ™” ๊ธฐ๋ฒ•

1๏ธโƒฃ BF16 ์ž๋™ ํ˜ผํ•ฉ ์ •๋ฐ€๋„

ํ‘œ์ค€ FP32์™€ ๋น„๊ตํ•ด ์•ฝ 50% VRAM ์ ˆ์•ฝ
- ๋ฉ”๋ชจ๋ฆฌ ํšจ์œจ: โฌ‡๏ธ 12GB โ†’ 6GB
- ์—ฐ์‚ฐ ์†๋„: โžก๏ธ ๋™๋“ฑ ๋˜๋Š” ํ–ฅ์ƒ

2๏ธโƒฃ 8๋น„ํŠธ AdamW ์˜ตํ‹ฐ๋งˆ์ด์ € (bitsandbytes)

์˜ตํ‹ฐ๋งˆ์ด์ € ์ƒํƒœ ๋ฉ”๋ชจ๋ฆฌ 75% ๊ฐ์†Œ
- ํ‘œ์ค€ AdamW: ~2.2GB (1B ๋ชจ๋ธ)
- 8-bit AdamW: ~0.55GB (1B ๋ชจ๋ธ)

3๏ธโƒฃ ์–‘์žํ™” (Quantization) โญ

๋ชจ๋ธ ๊ฐ€์ค‘์น˜ ๋™์  ์–‘์žํ™” ์ง€์›
- INT8 ์–‘์žํ™”: ํฌ๊ธฐ 4๋ฐฐ ๊ฐ์†Œ
- ์ถ”๋ก  ์†๋„: 1.5~2๋ฐฐ ํ–ฅ์ƒ

4๏ธโƒฃ ๊ทธ๋ž˜๋””์–ธํŠธ ๋ˆ„์  (Gradient Accumulation)

ํšจ๊ณผ์  ๋ฐฐ์น˜ ํฌ๊ธฐ ์ฆ๋Œ€
- ์„ค์ •: batch_size=2, accumulation_steps=8
- ํšจ๊ณผ: ๋ฐฐ์น˜ ํฌ๊ธฐ 16 ํšจ๊ณผ

5๏ธโƒฃ ๊ทธ๋ž˜๋””์–ธํŠธ ์ฒดํฌํฌ์ธํŒ…

ํ™œ์„ฑํ™”(Activation) ๋ฉ”๋ชจ๋ฆฌ ๊ฐ์†Œ
- ์žฌ๊ณ„์‚ฐ ๋น„์šฉ: ~30% ์†๋„ ์ €ํ•˜
- ๋ฉ”๋ชจ๋ฆฌ ์ ˆ์•ฝ: 30~40%

๐Ÿ’พ VRAM ์‚ฌ์šฉ๋Ÿ‰ ๋น„๊ต

๋ฒ„์ „ ํŒŒ๋ผ๋ฏธํ„ฐ VRAM ์‚ฌ์šฉ๋Ÿ‰ ์ตœ์ ํ™” ๊ธฐ๋ฒ•
v1 541M ~11GB ๊ธฐ๋ณธ FP32
v2 ~1.1B ~23GB BF16 + Gradient Checkpoint
v3 1.09B ~9GB โœจ BF16 + 8๋น„ํŠธ ์˜ตํ‹ฐ๋งˆ์ด์ € + ์–‘์žํ™”

v3์€ v2 ๋Œ€๋น„ VRAM 60% ๊ฐ์†Œ, v1๋ณด๋‹ค๋Š” ๋ชจ๋ธ ํฌ๊ธฐ 2๋ฐฐ ํ™•๋Œ€


๐Ÿš€ ๋น ๋ฅธ ์‹œ์ž‘

๐Ÿ“‹ ์‚ฌ์ „ ์š”๊ตฌ์‚ฌํ•ญ

Python 3.9 ์ด์ƒ
CUDA 11.8 ์ด์ƒ (GPU ํ•„์ˆ˜)
GPU ๋ฉ”๋ชจ๋ฆฌ: ์ตœ์†Œ 9GB ๊ถŒ์žฅ

1๏ธโƒฃ ์„ค์น˜

# ์ €์žฅ์†Œ ํด๋ก 
git clone https://github.com/seoan1024/korean-llm-v3.git
cd korean-llm-v3

# ํ•„์ˆ˜ ํŒจํ‚ค์ง€ ์„ค์น˜
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
pip install transformers datasets tqdm pandas matplotlib

# ์–‘์žํ™” ์ง€์› ๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ (์„ ํƒ)
pip install bitsandbytes

2๏ธโƒฃ ๋ฐ์ดํ„ฐ์…‹ ์ค€๋น„

์ฝ”๋“œ๊ฐ€ ์ž๋™์œผ๋กœ ๋‹ค์Œ ๋ฐ์ดํ„ฐ์…‹์„ ๋‹ค์šด๋กœ๋“œํ•ฉ๋‹ˆ๋‹ค:

  • ๐Ÿ”น nlpai-lab/kullm-v2 - ํ•œ๊ตญ์–ด ๋ช…๋ น์–ด ํŠœ๋‹ ๋ฐ์ดํ„ฐ
  • ๐Ÿ”น beomi/KoAlpaca-v1.1a - ํ•œ๊ตญ์‹ ์•ŒํŒŒ์นด ๋ฐ์ดํ„ฐ์…‹
# ๋ฐ์ดํ„ฐ์…‹์ด ์ž๋™ ๋‹ค์šด๋กœ๋“œ๋˜๋ฏ€๋กœ ๋ณ„๋„ ์ž‘์—… ๋ถˆํ•„์š”
# ์บ์‹œ ๋””๋ ‰ํ† ๋ฆฌ: ./datasets/cache/

3๏ธโƒฃ ํ•™์Šต ์‹คํ–‰

# ๊ธฐ๋ณธ ์„ค์ •์œผ๋กœ ํ•™์Šต ์‹œ์ž‘
python korean_llm_advanced_v3.py

# ๋˜๋Š” ์ปค์Šคํ…€ ์„ค์ •์œผ๋กœ ์‹คํ–‰
python korean_llm_advanced_v3.py \
    --batch-size 2 \
    --max-steps 50000 \
    --learning-rate 5e-5

4๏ธโƒฃ ๋ชจ๋‹ˆํ„ฐ๋ง

ํ•™์Šต ์ค‘ ์ž๋™์œผ๋กœ GUI ๋ชจ๋‹ˆํ„ฐ๋ง ์ฐฝ์ด ์—ด๋ฆฝ๋‹ˆ๋‹ค:

  • ๐Ÿ“Š ์‹ค์‹œ๊ฐ„ ์†์‹ค๊ฐ’(Loss) ๊ทธ๋ž˜ํ”„
  • ๐Ÿ’ฌ ์ธํ„ฐ๋ž™ํ‹ฐ๋ธŒ ์ฑ„ํŒ… (์ƒ์„ฑ ํ…Œ์ŠคํŠธ)
  • ๐Ÿ“ ๋กœ๊ทธ ๋ทฐ์–ด

๐Ÿ—๏ธ ํ”„๋กœ์ ํŠธ ๊ตฌ์กฐ

korean-llm-v3/
โ”œโ”€โ”€ korean_llm_advanced_v3.py    # ๋ฉ”์ธ ํ•™์Šต ์Šคํฌ๋ฆฝํŠธ
โ”œโ”€โ”€ README.md                      # ์ด ํŒŒ์ผ
โ”œโ”€โ”€ LICENSE                        # GPL-3.0 ๋ผ์ด์„ ์Šค
โ”‚
โ”œโ”€โ”€ checkpoints/                   # ์ €์žฅ๋œ ๋ชจ๋ธ ์ฒดํฌํฌ์ธํŠธ
โ”‚   โ””โ”€โ”€ korean_llm_*.pth
โ”‚
โ”œโ”€โ”€ datasets/                      # ๋ฐ์ดํ„ฐ์…‹ ์บ์‹œ
โ”‚   โ”œโ”€โ”€ cache/                     # ๋‹ค์šด๋กœ๋“œ๋œ ๋ฐ์ดํ„ฐ์…‹
โ”‚   โ””โ”€โ”€ datasets_manifest.json     # ๋ฉ”ํƒ€๋ฐ์ดํ„ฐ
โ”‚
โ””โ”€โ”€ logs/                          # ํ•™์Šต ๋กœ๊ทธ ๋ฐ ๊ทธ๋ž˜ํ”„
    โ”œโ”€โ”€ training.log               # ์ƒ์„ธ ๋กœ๊ทธ
    โ””โ”€โ”€ loss_history.json          # ์†์‹ค๊ฐ’ ๊ธฐ๋ก

๐Ÿ“Š ๋ฒ„์ „ ํžˆ์Šคํ† ๋ฆฌ

v1 (์ดˆ๊ธฐ ๋ฒ„์ „)

  • 541M ํŒŒ๋ผ๋ฏธํ„ฐ ๋ชจ๋ธ
  • VRAM ์‚ฌ์šฉ๋Ÿ‰: ~11GB
  • ๊ธฐ๋ณธ FP32 ํ•™์Šต

v2 (์ตœ์ ํ™” v1)

  • 1.1B ํŒŒ๋ผ๋ฏธํ„ฐ๋กœ ํ™•๋Œ€
  • VRAM ์‚ฌ์šฉ๋Ÿ‰: ~23GB (์ดˆ๊ธฐ 1.2๋ฐฐ ์ฆ๊ฐ€)
  • BF16 + Gradient Checkpoint ์ ์šฉ

v3 (ํ˜„์žฌ) โญ

  • 1.09B ํŒŒ๋ผ๋ฏธํ„ฐ (v2 ์ˆ˜์ค€)
  • VRAM ์‚ฌ์šฉ๋Ÿ‰: ~9GB (v2 ๋Œ€๋น„ 60% ๊ฐ์†Œ!)
  • ์ฃผ์š” ๊ฐœ์„ ์‚ฌํ•ญ:
    • 8๋น„ํŠธ AdamW ์˜ตํ‹ฐ๋งˆ์ด์ €
    • ๋™์  ์–‘์žํ™” ์ง€์›
    • ํ–ฅ์ƒ๋œ ๋ฉ”๋ชจ๋ฆฌ ๊ด€๋ฆฌ
    • ๋” ๋น ๋ฅธ ํ•™์Šต ์†๋„

๐Ÿ”ง ๊ธฐ์ˆ  ์Šคํƒ

ํ•ต์‹ฌ ๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ

๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ ๋ฒ„์ „ ์šฉ๋„
PyTorch 2.0+ ๋”ฅ๋Ÿฌ๋‹ ํ”„๋ ˆ์ž„์›Œํฌ
Transformers 4.30+ ํ† ํฌ๋‚˜์ด์ € ๋ฐ ์œ ํ‹ธ๋ฆฌํ‹ฐ
Datasets 2.10+ ํ•œ๊ตญ์–ด ๋ฐ์ดํ„ฐ์…‹ ๋กœ๋“œ
bitsandbytes 0.40+ 8๋น„ํŠธ ์–‘์žํ™” ์ตœ์ ํ™”
tqdm 4.60+ ์ง„ํ–‰๋ฅ  ํ‘œ์‹œ

์„ ํƒ ๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ

๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ ์šฉ๋„
matplotlib ์†์‹ค๊ฐ’ ๊ทธ๋ž˜ํ”„ ์‹œ๊ฐํ™”
tkinter GUI ๋ชจ๋‹ˆํ„ฐ๋ง (๋‚ด์žฅ)
pandas ๋ฐ์ดํ„ฐ ์ฒ˜๋ฆฌ

๐Ÿ’ก ์‚ฌ์šฉ ์˜ˆ์‹œ

๋ชจ๋ธ ๋กœ๋“œ ๋ฐ ํ…์ŠคํŠธ ์ƒ์„ฑ

import os, argparse
from pathlib import Path
from typing import Optional, Tuple, List
import torch, torch.nn as nn, torch.nn.functional as F
from transformers import AutoTokenizer

class RMSNorm(nn.Module):
    def __init__(self, dim, eps=1e-6):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))
    def forward(self, x):
        return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps) * self.weight

def precompute_freqs_cis(head_dim: int, end: int, theta: float = 10000.0) -> Tuple[torch.Tensor, torch.Tensor]:
    freqs = 1.0 / (theta ** (torch.arange(0, head_dim, 2)[:head_dim // 2].float() / head_dim))
    t = torch.arange(end, dtype=freqs.dtype)
    freqs = torch.outer(t, freqs)
    return torch.cos(freqs), torch.sin(freqs)

def apply_rotary_emb(x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
    head_dim_2 = cos.shape[-1]
    head_dim = head_dim_2 * 2
    x1 = x[..., :head_dim // 2]
    x2 = x[..., head_dim // 2:]
    cos = cos.unsqueeze(0).unsqueeze(0)
    sin = sin.unsqueeze(0).unsqueeze(0)
    return torch.cat([x1 * cos - x2 * sin, x1 * sin + x2 * cos], dim=-1)

class SwiGLU(nn.Module):
    def __init__(self, dim: int, hidden_dim: int):
        super().__init__()
        self.w1 = nn.Linear(dim, hidden_dim, bias=False)
        self.w2 = nn.Linear(hidden_dim, dim, bias=False)
        self.w3 = nn.Linear(dim, hidden_dim, bias=False)
    def forward(self, x):
        return self.w2(F.silu(self.w1(x)) * self.w3(x))

class Attention(nn.Module):
    def __init__(self, dim: int, n_heads: int):
        super().__init__()
        assert dim % n_heads == 0
        self.n_heads = n_heads
        self.head_dim = dim // n_heads
        self.wq = nn.Linear(dim, dim, bias=False)
        self.wk = nn.Linear(dim, dim, bias=False)
        self.wv = nn.Linear(dim, dim, bias=False)
        self.wo = nn.Linear(dim, dim, bias=False)
    def forward(self, x: torch.Tensor, f_cos: torch.Tensor, f_sin: torch.Tensor, kv_cache: Optional[Tuple[torch.Tensor, torch.Tensor]] = None):
        b, s, d = x.shape
        q = self.wq(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        k = self.wk(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        v = self.wv(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        q = apply_rotary_emb(q, f_cos, f_sin)
        k = apply_rotary_emb(k, f_cos, f_sin)
        if kv_cache is not None:
            pk, pv = kv_cache
            k = torch.cat([pk, k], dim=2)
            v = torch.cat([pv, v], dim=2)
        new_kv = (k.detach(), v.detach())
        out = F.scaled_dot_product_attention(q, k, v, attn_mask=None, is_causal=(s > 1))
        out = out.transpose(1, 2).contiguous().view(b, s, d)
        return self.wo(out), new_kv

class TransformerBlock(nn.Module):
    def __init__(self, dim: int, n_heads: int, hidden_dim: int):
        super().__init__()
        self.attention = Attention(dim, n_heads)
        self.feed_forward = SwiGLU(dim, hidden_dim)
        self.attention_norm = RMSNorm(dim)
        self.ffn_norm = RMSNorm(dim)
    def forward(self, x, f_cos, f_sin, kv_cache=None):
        normed_x = self.attention_norm(x)
        h, new_kv = self.attention(normed_x, f_cos, f_sin, kv_cache=kv_cache)
        x = x + h
        x = x + self.feed_forward(self.ffn_norm(x))
        return x, new_kv

class KoreanLLM(nn.Module):
    def __init__(self, vocab_size: int, pad_token_id: int, dim: int = 1920, n_layers: int = 20, n_heads: int = 10, max_seq_len: int = 512):
        super().__init__()
        self.vocab_size = vocab_size
        self.pad_token_id = pad_token_id
        self.dim = dim
        self.n_heads = n_heads
        self.head_dim = dim // n_heads
        self.max_seq_len = max_seq_len
        self.embed = nn.Embedding(vocab_size, dim)
        self.layers = nn.ModuleList([TransformerBlock(dim, n_heads, int(dim * 2.5)) for _ in range(n_layers)])
        self.norm = RMSNorm(dim)
        self.output = nn.Linear(dim, vocab_size, bias=False)
        self.output.weight = self.embed.weight
        f_cos, f_sin = precompute_freqs_cis(self.head_dim, max_seq_len * 2)
        self.register_buffer("f_cos", f_cos)
        self.register_buffer("f_sin", f_sin)
    def _get_freqs(self, f, start, length):
        end = start + length
        if end > f.shape[0]:
            raise ValueError(f"ํ˜„์žฌ ์ปจํ…์ŠคํŠธ๊ฐ€ ๋„ˆ๋ฌด ๊น๋‹ˆ๋‹ค: {end} > {f.shape[0]}")
        return f[start:end]
    @torch.no_grad()
    def forward(self, tokens: torch.Tensor, kv_caches=None):
        b, s = tokens.shape
        x = self.embed(tokens)
        start_pos = 0
        if kv_caches is not None and len(kv_caches) > 0 and kv_caches[0][0] is not None:
            start_pos = kv_caches[0][0].shape[2]
        f_cos = self._get_freqs(self.f_cos, start_pos, s)
        f_sin = self._get_freqs(self.f_sin, start_pos, s)
        new_kv_caches = []
        for i, layer in enumerate(self.layers):
            cache = kv_caches[i] if kv_caches is not None else None
            x, kv = layer(x, f_cos, f_sin, kv_cache=cache)
            new_kv_caches.append(kv)
        x = self.norm(x)
        logits = self.output(x)
        return logits, new_kv_caches

def find_latest_checkpoint(checkpoint_dir="checkpoints"):
    checkpoint_dir = Path(checkpoint_dir)
    if not checkpoint_dir.exists(): return None
    files = list(checkpoint_dir.glob("korean_llm_*.pth"))
    if not files: return None
    def step_number(path):
        try: return int(path.stem.split("_")[-1])
        except ValueError: return -1
    files.sort(key=step_number)
    return files[-1]

def load_checkpoint(model, checkpoint_path, device):
    print(f"๐Ÿ“ฆ ์ฒดํฌํฌ์ธํŠธ ๋กœ๋”ฉ:\n   {checkpoint_path}")
    checkpoint = torch.load(checkpoint_path, map_location=device)
    if "model_state_dict" in checkpoint:
        state_dict = checkpoint["model_state_dict"]
        step = checkpoint.get("step", "?")
    else:
        state_dict = checkpoint
        step = "?"
    model.load_state_dict(state_dict, strict=True)
    print(f"โœ… ๋ชจ๋ธ ๋กœ๋“œ ์™„๋ฃŒ\n   ํ•™์Šต step: {step}")
    return step

@torch.no_grad()
def generate(model, tokenizer, prompt, device, max_tokens=256, temperature=0.6, top_k=40, top_p=0.95, repetition_penalty=1.15, context_limit=512):
    model.eval()
    prompt_text = f"### ์งˆ๋ฌธ: {prompt}\n### ์‘๋‹ต:"
    tokens = tokenizer.encode(prompt_text, add_special_tokens=False, return_tensors="pt").to(device)
    if tokens.shape[1] >= context_limit:
        tokens = tokens[:, -context_limit + 1:]
    output_tokens = tokens
    kv_caches = None
    eos_id = tokenizer.eos_token_id
    for _ in range(max_tokens):
        input_tokens = output_tokens if kv_caches is None else output_tokens[:, -1:]
        logits, kv_caches = model(input_tokens, kv_caches=kv_caches)
        next_logits = logits[:, -1, :]
        temperature = max(float(temperature), 1e-5)
        next_logits = next_logits / temperature
        if repetition_penalty != 1.0:
            used_tokens = set(output_tokens[0].tolist())
            for token_id in used_tokens:
                if token_id < next_logits.shape[-1]:
                    if next_logits[0, token_id] < 0:
                        next_logits[0, token_id] *= repetition_penalty
                    else:
                        next_logits[0, token_id] /= repetition_penalty
        if top_k > 0:
            k = min(int(top_k), next_logits.shape[-1])
            threshold = torch.topk(next_logits, k).values[..., -1, None]
            next_logits = torch.where(next_logits < threshold, torch.full_like(next_logits, float("-inf")), next_logits)
        probs = F.softmax(next_logits, dim=-1)
        if 0 < top_p < 1.0:
            sorted_probs, sorted_indices = torch.sort(probs, descending=True, dim=-1)
            cumulative = torch.cumsum(sorted_probs, dim=-1)
            remove = cumulative > top_p
            remove[..., 0] = False
            indices_to_remove = torch.zeros_like(probs, dtype=torch.bool)
            indices_to_remove.scatter_(-1, sorted_indices, remove)
            probs = probs.masked_fill(indices_to_remove, 0.0)
            probs = probs / (probs.sum(dim=-1, keepdim=True) + 1e-10)
        if not torch.isfinite(probs).all():
            next_token = torch.argmax(next_logits, dim=-1, keepdim=True)
        else:
            next_token = torch.multinomial(probs, num_samples=1)
        output_tokens = torch.cat([output_tokens, next_token], dim=1)
        if eos_id is not None and next_token.item() == eos_id: break
        if output_tokens.shape[1] >= context_limit: break
    generated_text = tokenizer.decode(output_tokens[0], skip_special_tokens=True)
    if "### ์‘๋‹ต:" in generated_text:
        response = generated_text.split("### ์‘๋‹ต:", 1)[1]
    else:
        response = generated_text
    if "### ์งˆ๋ฌธ:" in response:
        response = response.split("### ์งˆ๋ฌธ:", 1)[0]
    return response.strip()

def main():
    parser = argparse.ArgumentParser(description="KoreanLLM ์ฒดํฌํฌ์ธํŠธ ์ฑ„ํŒ…")
    parser.add_argument("--checkpoint", type=str, default="latest", help="์ฒดํฌํฌ์ธํŠธ ๊ฒฝ๋กœ ๋˜๋Š” latest")
    parser.add_argument("--tokenizer", type=str, default="beomi/Llama-3-Open-Ko-8B")
    parser.add_argument("--max-tokens", type=int, default=256)
    parser.add_argument("--temperature", type=float, default=0.6)
    parser.add_argument("--top-k", type=int, default=40)
    parser.add_argument("--top-p", type=float, default=0.95)
    parser.add_argument("--repetition-penalty", type=float, default=1.15)
    parser.add_argument("--cpu", action="store_true", help="๊ฐ•์ œ๋กœ CPU ์‚ฌ์šฉ")
    args = parser.parse_args()
    if args.cpu:
        device = torch.device("cpu")
    else:
        device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    print("=" * 65)
    print("๐Ÿ‡ฐ๐Ÿ‡ท KoreanLLM v3 Checkpoint Chat")
    print("=" * 65)
    print(f"๐Ÿ–ฅ๏ธ Device: {device}")
    if device.type == "cuda":
        print(f"๐ŸŽฎ GPU: {torch.cuda.get_device_name(0)}")
    else:
        print("โš ๏ธ CPU ๋ชจ๋“œ์ž…๋‹ˆ๋‹ค. 1920-dim / 20-layer ๋ชจ๋ธ์ด๋ผ ์ƒ์„ฑ์ด ๋А๋ฆด ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.")
    print("\n๐Ÿ”ค ํ† ํฌ๋‚˜์ด์ € ๋กœ๋”ฉ...")
    tokenizer = AutoTokenizer.from_pretrained(args.tokenizer, clean_up_tokenization_spaces=False)
    if tokenizer.pad_token is None or tokenizer.pad_token_id == tokenizer.eos_token_id:
        tokenizer.add_special_tokens({"pad_token": "<|pad|>"})
    vocab_size = len(tokenizer)
    pad_token_id = tokenizer.pad_token_id
    print(f"   vocab_size = {vocab_size}")
    print(f"   eos_token_id = {tokenizer.eos_token_id}")
    print(f"   pad_token_id = {pad_token_id}")
    checkpoint_path = args.checkpoint
    if checkpoint_path.lower() == "latest":
        checkpoint_path = find_latest_checkpoint()
        if checkpoint_path is None:
            print("\nโŒ checkpoints ํด๋”์—์„œ ์ฒดํฌํฌ์ธํŠธ๋ฅผ ์ฐพ์ง€ ๋ชปํ–ˆ์Šต๋‹ˆ๋‹ค.")
            print("์˜ˆ: python chat_korean_llm.py --checkpoint checkpoints/korean_llm_50000.pth")
            return
    checkpoint_path = Path(checkpoint_path)
    if not checkpoint_path.exists():
        print(f"\nโŒ ์ฒดํฌํฌ์ธํŠธ๊ฐ€ ์—†์Šต๋‹ˆ๋‹ค:\n   {checkpoint_path}")
        return
    print("\n๐Ÿง  ๋ชจ๋ธ ์ƒ์„ฑ ์ค‘...")
    model_config = dict(vocab_size=vocab_size, pad_token_id=pad_token_id, dim=1920, n_layers=20, n_heads=10, max_seq_len=512)
    dtype = torch.bfloat16 if device.type == "cuda" else torch.float32
    model = KoreanLLM(**model_config).to(device)
    if device.type == "cuda":
        model = model.to(dtype=dtype)
    try:
        step = load_checkpoint(model, checkpoint_path, device)
    except RuntimeError as e:
        print("\nโŒ ์ฒดํฌํฌ์ธํŠธ์™€ ํ˜„์žฌ ๋ชจ๋ธ ๊ตฌ์กฐ๊ฐ€ ๋งž์ง€ ์•Š์Šต๋‹ˆ๋‹ค.")
        print("   ํŠนํžˆ tokenizer์˜ vocab_size / pad_token ์„ค์ •์„ ํ™•์ธํ•˜์„ธ์š”.")
        print(f"\n์ƒ์„ธ ์˜ค๋ฅ˜:\n{e}")
        return
    model.eval()
    params = sum(p.numel() for p in model.parameters())
    print(f"๐Ÿ“Š Parameters: {params / 1e6:.1f}M")
    print("\n" + "=" * 65)
    print("๐Ÿ’ฌ ์ฑ„ํŒ… ์‹œ์ž‘")
    print("   /exit  ์ข…๋ฃŒ")
    print("   /clear ๋Œ€ํ™” ์ž…๋ ฅ ๊ธฐ๋ก ์ดˆ๊ธฐํ™”")
    print("   /info  ๋ชจ๋ธ ์ •๋ณด")
    print("=" * 65)
    while True:
        try:
            prompt = input("\n๋‚˜ > ").strip()
        except (KeyboardInterrupt, EOFError):
            print("\n\n๐Ÿ‘‹ ์ข…๋ฃŒํ•ฉ๋‹ˆ๋‹ค.")
            break
        if not prompt: continue
        if prompt.lower() in {"/exit", "/quit", "exit", "quit"}:
            print("๐Ÿ‘‹ ์ข…๋ฃŒํ•ฉ๋‹ˆ๋‹ค.")
            break
        if prompt == "/clear":
            print("๐Ÿงน ์ž…๋ ฅ ์ƒํƒœ๋ฅผ ์ดˆ๊ธฐํ™”ํ–ˆ์Šต๋‹ˆ๋‹ค.")
            continue
        if prompt == "/info":
            print(f"\n์ฒดํฌํฌ์ธํŠธ : {checkpoint_path}\nํ•™์Šต step   : {step}\nDevice      : {device}\nParameters  : {params / 1e6:.1f}M\nTemperature : {args.temperature}\nTop-k       : {args.top_k}\nTop-p       : {args.top_p}")
            continue
        print("\n๋ชจ๋ธ > ", end="", flush=True)
        try:
            response = generate(model=model, tokenizer=tokenizer, prompt=prompt, device=device, max_tokens=args.max_tokens, temperature=args.temperature, top_k=args.top_k, top_p=args.top_p, repetition_penalty=args.repetition_penalty, context_limit=512)
            print(response)
        except torch.cuda.OutOfMemoryError:
            print("\nโŒ CUDA ๋ฉ”๋ชจ๋ฆฌ๊ฐ€ ๋ถ€์กฑํ•ฉ๋‹ˆ๋‹ค.")
            print("   --max-tokens ๊ฐ’์„ ๋‚ฎ์ถ”๊ฑฐ๋‚˜ ๋‹ค๋ฅธ GPU์—์„œ ์‹คํ–‰ํ•ด๋ณด์„ธ์š”.")
        except Exception as e:
            print(f"\nโŒ ์ƒ์„ฑ ์˜ค๋ฅ˜: {type(e).__name__}: {e}")

if __name__ == "__main__":
    main()

์ปค์Šคํ…€ ํ•™์Šต ์„ค์ •

from korean_llm_advanced_v3 import TrainingConfig, main

config = TrainingConfig(
    batch_size=4,                      # ๋ฐฐ์น˜ ํฌ๊ธฐ
    accumulation_steps=4,              # ๊ทธ๋ž˜๋””์–ธํŠธ ๋ˆ„์  ์Šคํ…
    max_steps=100000,                  # ์ตœ๋Œ€ ํ•™์Šต ์Šคํ…
    warmup_steps=1000,                 # ์›Œ๋ฐ์—… ์Šคํ…
    learning_rate=3e-5,                # ํ•™์Šต๋ฅ 
    eval_interval=5000,                # ํ‰๊ฐ€ ๊ฐ„๊ฒฉ
    use_bfloat16=True,                 # BF16 ์‚ฌ์šฉ ์—ฌ๋ถ€
    resume_from_checkpoint='latest'    # ์ตœ์‹  ์ฒดํฌํฌ์ธํŠธ์—์„œ ์žฌ๊ฐœ
)

main(config)

โ“ FAQ (์ž์ฃผ ๋ฌป๋Š” ์งˆ๋ฌธ)

Q1: ์ด ๋ชจ๋ธ์„ ์ถ”๋ก (inference)๋งŒ ํ•˜๋ ค๋ฉด?

A: ํ•™์Šต๋œ ์ฒดํฌํฌ์ธํŠธ๊ฐ€ ์žˆ๋‹ค๋ฉด ๋‹ค์Œ์ฒ˜๋Ÿผ ๊ฐ„๋‹จํžˆ:

import torch
from korean_llm_advanced_v3 import KoreanLLM, generate
from transformers import AutoTokenizer

model = KoreanLLM(...).to(device)
checkpoint = torch.load("checkpoints/korean_llm_50000.pth", map_location=device)
model.load_state_dict(checkpoint['model_state_dict'])
model.eval()

response = generate(model, tokenizer, prompt="์•ˆ๋…•?", max_tokens=50)

Q2: ๋‚ด GPU ๋ฉ”๋ชจ๋ฆฌ๊ฐ€ 9GB ๋ฏธ๋งŒ์ด๋ฉด?

A: ๋‹ค์Œ ๋ฐฉ๋ฒ•๋“ค์„ ์‹œ๋„ํ•ด๋ณด์„ธ์š”:

  • ๋ฐฐ์น˜ ํฌ๊ธฐ๋ฅผ 1๋กœ ๊ฐ์†Œ
  • ์‹œํ€€์Šค ๊ธธ์ด๋ฅผ 1024๋กœ ๋‹จ์ถ•
  • ๊ทธ๋ž˜๋””์–ธํŠธ ๋ˆ„์  ๋‹จ๊ณ„๋ฅผ 16์œผ๋กœ ์ฆ๊ฐ€
  • 8๋น„ํŠธ ์–‘์žํ™” ํ™œ์„ฑํ™”

Q3: ํ•™์Šต ์ค‘๋‹จ ํ›„ ์žฌ๊ฐœํ•˜๋ ค๋ฉด?

A: ์ž๋™์œผ๋กœ ์ตœ์‹  ์ฒดํฌํฌ์ธํŠธ๋ฅผ ๊ฐ์ง€ํ•ฉ๋‹ˆ๋‹ค:

config = TrainingConfig(
    resume_from_checkpoint='latest'  # ๋˜๋Š” ํŠน์ • ๊ฒฝ๋กœ
)
main(config)

Q4: ๋‹ค๋ฅธ ํ•œ๊ตญ์–ด ๋ฐ์ดํ„ฐ์…‹์„ ์‚ฌ์šฉํ•  ์ˆ˜ ์žˆ๋‚˜?

A: ๋„ค! DatasetManager ํด๋ž˜์Šค์˜ DATASETS_CONFIG๋ฅผ ์ˆ˜์ •ํ•˜๋ฉด ๋ฉ๋‹ˆ๋‹ค:

DATASETS_CONFIG = [
    {
        "name": "your-dataset/path",
        "split": "train",
        "text_keys": ["input", "output"]
    }
]

Q5: ์œˆ๋„์šฐ์—์„œ ์‹คํ–‰ํ•˜๋ฉด ์—๋Ÿฌ๊ฐ€ ๋‚˜์š”

A: num_workers ์„ค์ •์„ 0์œผ๋กœ ๋ณ€๊ฒฝํ•ด๋ณด์„ธ์š”:

loader = DataLoader(dataset, batch_size=2, num_workers=0)

Q6: VRAM ์‚ฌ์šฉ๋Ÿ‰์„ ๋” ์ค„์ผ ์ˆ˜ ์žˆ๋‚˜?

A: ๋‹ค์Œ ์˜ต์…˜์„ ์กฐํ•ฉํ•ด๋ณด์„ธ์š”:

  • ๋ฉ”๋ชจ๋ฆฌ ํšจ์œจ ๋ชจ๋“œ: use_bfloat16=True
  • ๋” ๊นŠ์€ ์–‘์žํ™”: INT4 (์ถ”๊ฐ€ ๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ ํ•„์š”)
  • LoRA ํŒŒ์ธํŠœ๋‹: ์„ ํƒ์  ๋ ˆ์ด์–ด๋งŒ ํ•™์Šต

Q7: ์ƒ์„ฑ๋œ ํ…์ŠคํŠธ ํ’ˆ์งˆ์ด ๋‚ฎ์œผ๋ฉด?

A: ๋‹ค์Œ์„ ํ™•์ธํ•˜์„ธ์š”:

  • ํ•™์Šต ์Šคํ…์ด ์ถฉ๋ถ„ํ•œ๊ฐ€? (์ตœ์†Œ 10,000 ์Šคํ… ๊ถŒ์žฅ)
  • Learning rate ์„ค์ •์ด ์ ์ ˆํ•œ๊ฐ€?
  • ๋ฐ์ดํ„ฐ์…‹ ํ’ˆ์งˆ์ด ์ข‹์€๊ฐ€?
  • temperature ํŒŒ๋ผ๋ฏธํ„ฐ ์กฐ์ • (0.5~1.0 ๊ถŒ์žฅ)

Q8: ๋ชจ๋ธ์„ ONNX๋‚˜ ๋‹ค๋ฅธ ํ˜•์‹์œผ๋กœ ๋ณ€ํ™˜ํ•˜๋ ค๋ฉด?

A: PyTorch์—์„œ ONNX๋กœ ๋ณ€ํ™˜ ๊ฐ€๋Šฅ:

import torch.onnx

dummy_input = torch.randint(0, 50000, (1, 2048)).to(device)
torch.onnx.export(
    model, dummy_input, "korean_llm.onnx",
    input_names=['input_ids'],
    output_names=['output']
)

Q9: ๊ฐœ๋ฐœ์ž๊ฐ€ ํ™œ๋ฐœํžˆ ์ง€์›ํ•˜๋‚˜?

A: ๋„ค! ์ด์Šˆ๋‚˜ ํ”ผ๋“œ๋ฐฑ์€ ์ด๋ฉ”์ผ(seoan102410@gmail.com)๋กœ ์—ฐ๋ฝ์ฃผ์„ธ์š”! ๐Ÿ’Œ

Q10: ์ƒ์šฉ ํ”„๋กœ์ ํŠธ์— ์‚ฌ์šฉ ๊ฐ€๋Šฅํ•œ๊ฐ€?

A: GPL-3.0 ๋ผ์ด์„ ์Šค์ด๋ฏ€๋กœ, ์ˆ˜์ • ์‚ฌํ•ญ์„ ๊ณต๊ฐœํ•ด์•ผ ํ•ฉ๋‹ˆ๋‹ค. ์ž์„ธํ•œ ๋‚ด์šฉ์€ LICENSE ํŒŒ์ผ์„ ํ™•์ธํ•˜์„ธ์š”.


๐Ÿ› ๏ธ ํŠธ๋Ÿฌ๋ธ”์ŠˆํŒ…

โŒ CUDA Out of Memory ์—๋Ÿฌ

์ฆ์ƒ: RuntimeError: CUDA out of memory

ํ•ด๊ฒฐ์ฑ…:

# ๋ฐฐ์น˜ ํฌ๊ธฐ ๊ฐ์†Œ
config.batch_size = 1

# ์ตœ๋Œ€ ์‹œํ€€์Šค ๊ธธ์ด ๊ฐ์†Œ
config.max_seq_len = 1024

# ๊ทธ๋ž˜๋””์–ธํŠธ ๋ˆ„์  ์ฆ๊ฐ€
config.accumulation_steps = 16

โŒ bitsandbytes ์„ค์น˜ ์‹คํŒจ

ํ•ด๊ฒฐ์ฑ…:

# CUDA Toolkit ๊ฒฝ๋กœ ๋ช…์‹œ
CUDA_HOME=/usr/local/cuda pip install bitsandbytes

โŒ ๋ฐ์ดํ„ฐ์…‹ ๋‹ค์šด๋กœ๋“œ ์‹คํŒจ

ํ•ด๊ฒฐ์ฑ…:

# ์บ์‹œ ์ดˆ๊ธฐํ™” ํ›„ ์žฌ์‹œ๋„
rm -rf datasets/cache/*
python korean_llm_advanced_v3.py

๐Ÿ“ˆ ์„ฑ๋Šฅ ์ตœ์ ํ™” ํŒ

  1. ๋ฐฐ์น˜ ํฌ๊ธฐ ์กฐ์ •: ๋„ˆ๋ฌด ์ž‘์œผ๋ฉด ํ•™์Šต์ด ๋А๋ฆฌ๊ณ , ๋„ˆ๋ฌด ํฌ๋ฉด VRAM ๋ถ€์กฑ
  2. ๊ทธ๋ž˜๋””์–ธํŠธ ๋ˆ„์ : ํšจ๊ณผ์ ์ธ ๋ฐฐ์น˜ ํฌ๊ธฐ ์ฆ๋Œ€์˜ ํ•ต์‹ฌ
  3. Learning Rate ์Šค์ผ€์ค„๋ง: Cosine Annealing์œผ๋กœ ์ˆ˜๋ ด ํ–ฅ์ƒ
  4. ํ˜ผํ•ฉ ์ •๋ฐ€๋„: BF16 ์‚ฌ์šฉ์œผ๋กœ ์†๋„์™€ ๋ฉ”๋ชจ๋ฆฌ ๋™์‹œ ๊ฐœ์„ 
  5. ์ฒดํฌํฌ์ธํŠธ: ์ •๊ธฐ์ ์œผ๋กœ ์ €์žฅํ•˜์—ฌ ํ•™์Šต ์žฌ๊ฐœ ๊ฐ€๋Šฅ

๐Ÿ“ž ์—ฐ๋ฝ์ฒ˜ ๋ฐ ์ •๋ณด


๐Ÿ“œ ๋ผ์ด์„ ์Šค

์ด ํ”„๋กœ์ ํŠธ๋Š” GPL-3.0 ๋ผ์ด์„ ์Šค ํ•˜์— ๋ฐฐํฌ๋ฉ๋‹ˆ๋‹ค.

GNU GENERAL PUBLIC LICENSE
Version 3, 29 June 2007

Copyright (C) 2024 seoan1024

This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.

๐Ÿ“– ์ „์ฒด ๋ผ์ด์„ ์Šค: LICENSE


๐Ÿค ๊ธฐ์—ฌํ•˜๊ธฐ

๋ฒ„๊ทธ ๋ฆฌํฌํŠธ, ๊ธฐ๋Šฅ ์ œ์•ˆ, ํ’€ ๋ฆฌํ€˜์ŠคํŠธ๋Š” ์–ธ์ œ๋‚˜ ํ™˜์˜ํ•ฉ๋‹ˆ๋‹ค!

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

๐Ÿ™ ๊ฐ์‚ฌ์˜ ๋ง

  • ๐ŸŽฏ ํ•œ๊ตญ์–ด LLM ์ปค๋ฎค๋‹ˆํ‹ฐ - ๊ท€์ค‘ํ•œ ํ”ผ๋“œ๋ฐฑ๊ณผ ๊ธฐ์—ฌ
  • ๐Ÿ“š Hugging Face - Transformers & Datasets ๋ผ์ด๋ธŒ๋Ÿฌ๋ฆฌ
  • ๐Ÿ”ง bitsandbytes - ์–‘์žํ™” ๋ฐ ์ตœ์ ํ™” ์†”๋ฃจ์…˜
  • ๐Ÿš€ PyTorch - ์˜คํ”ˆ์†Œ์Šค ๋”ฅ๋Ÿฌ๋‹ ํ”„๋ ˆ์ž„์›Œํฌ

๐Ÿ“š ์ฐธ๊ณ  ์ž๋ฃŒ

ํ•œ๊ตญ์–ด NLP

์ตœ์ ํ™” ๊ธฐ๋ฒ•

๋Œ€๊ทœ๋ชจ ์–ธ์–ด๋ชจ๋ธ


โญ ์ด ํ”„๋กœ์ ํŠธ๊ฐ€ ๋„์›€์ด ๋˜์—ˆ๋‹ค๋ฉด ๋ณ„โญ์„ ๋ˆŒ๋Ÿฌ์ฃผ์„ธ์š”!

โฌ† ์œ„๋กœ


๐Ÿ‡บ๐Ÿ‡ธ Korean LLM Advanced v3

License: GPL-3.0 Python 3.9+ PyTorch CUDA Model Size VRAM

Korean-Optimized Large Language Model - Scratch Implementation & Quantization Applied

๐Ÿš€ Project Development Journey

At first, free API quotas became tight, and there was a limit to "vibe coding." I tried running Ollama locally, but the models were too heavy for my computer to handle. Then suddenly, the thought struck me: "Why not just build it myself?"

However, I was only in 8th grade and knew almost nothing about AI models except the concept of 'B (Billion)' for model size. I couldn't code advanced concepts, only the basics.

So, as usual, I went into ChatGPT and boldly declared: "I want to create my own independent Korean LLM model!" and started this ambitious challenge.

While receiving help from GPT, I kept hitting limitations. Initially, I just collected code snippets from GPT and desperately hoped to avoid Dimension Errors. I spent entire days collecting and cleaning data, staring at my computer monitor.

The first version of my work (pre-v1) was an ultra-lightweight 50M (50 million parameter) model trained on Wikipedia dataโ€”it no longer exists. Although proper conversation was impossible, it did show signs of constructing grammatically correct sentences. I was so happy with that small success that I became obsessed with creating a "chatbot-type model." This led me to develop v1 to its final version: a 541M-sized model. Despite some bugs, I decided to fix them and immediately scale up the model.

After fixing errors and roughly doubling the model size, I finally built a 1.09B (1.09 billion parameter) model. Throughout the vacation, I kept my computer running for training whenever I had time.

Before I knew it, the seemingly endless training had reached 44,000 steps. With excitement, I typed "์•ˆ๋…•?" (Hello?) into the chat for testing.

"์•ˆ๋…•ํ•˜์„ธ์š”! ์˜ค๋Š˜์€ ๋ฌด์—‡์„ ๋„์™€๋“œ๋ฆด๊นŒ์š”?" (Hello! What can I help you with today?)

The moment the model displayed the correct response on screen, I felt indescribable joy.

But joy was short-lived. As I asked different questions, it started spitting out completely wrong answers. After analyzing code with AI all night, I discovered a critical bug: the model was ignoring user instructions. Heartbroken, I bravely discarded all training results for a more perfect model.

Without losing hope, I completely fixed v2's bugs and tackled the next challenge. A 1B-class model consumed a whopping 23GB of VRAM, making it too heavy to run in typical environments. I set a goal to reduce this to 10GB or less, and finally succeeded with quantization in v3.

As vacation ended, it became difficult for me to keep training. I haven't done any training since the v2 bug fix, and I'll resume when time allows. This project is my life's greatest work, completed solely through sheer stubbornness to build an LLM with my own hands.

Please use this model well, and if you like it, don't forget to click the star button (โญ)! Thank you!

A 1.09B parameter LLM trained entirely in Korean from scratch, making aggressive use of VRAM optimization techniques.

๐Ÿ“‹ Key Features โ€ข ๐Ÿš€ Quick Start โ€ข ๐Ÿ’พ Technology Stack โ€ข ๐Ÿ“Š Version History


๐Ÿ“– Overview

Korean LLM Advanced v3 is a lightweight large language model optimized for Korean natural language processing. It is designed to train and perform inference efficiently even in limited GPU memory environments.

Core Goals

  • โœ… Korean text generation and comprehension
  • โœ… VRAM efficiency (9GB baseline)
  • โœ… Fast training speed
  • โœ… Easy deployment and utilization

๐ŸŒŸ Key Features

๐ŸŽฏ Model Architecture

Item Description
Model Size 1.09B Parameters
Hidden Dimension 1,920
Number of Layers 20
Attention Heads 10
Max Sequence Length 2,048 Tokens
Vocabulary Size Dynamic (based on tokenizer)

๐Ÿ”ง Optimization Techniques

1๏ธโƒฃ BF16 Automatic Mixed Precision

~50% VRAM savings compared to standard FP32
- Memory efficiency: โฌ‡๏ธ 12GB โ†’ 6GB
- Computation speed: โžก๏ธ Equivalent or improved

2๏ธโƒฃ 8-bit AdamW Optimizer (bitsandbytes)

75% reduction in optimizer state memory
- Standard AdamW: ~2.2GB (1B model)
- 8-bit AdamW: ~0.55GB (1B model)

3๏ธโƒฃ Quantization โญ

Dynamic quantization of model weights
- INT8 Quantization: 4x size reduction
- Inference speed: 1.5~2x improvement

4๏ธโƒฃ Gradient Accumulation

Effective batch size increase
- Configuration: batch_size=2, accumulation_steps=8
- Effect: Equivalent to batch size 16

5๏ธโƒฃ Gradient Checkpointing

Activation memory reduction
- Recomputation cost: ~30% speed decrease
- Memory savings: 30~40%

๐Ÿ’พ VRAM Usage Comparison

Version Parameters VRAM Usage Optimization Techniques
v1 541M ~11GB Basic FP32
v2 ~1.1B ~23GB BF16 + Gradient Checkpoint
v3 1.09B ~9GB โœจ BF16 + 8-bit Optimizer + Quantization

v3 achieves 60% VRAM reduction compared to v2, with 2x model size expansion vs v1


๐Ÿš€ Quick Start

๐Ÿ“‹ Prerequisites

Python 3.9 or higher
CUDA 11.8 or higher (GPU required)
GPU Memory: Minimum 9GB recommended

1๏ธโƒฃ Installation

# Clone the repository
git clone https://github.com/seoan1024/korean-llm-v3.git
cd korean-llm-v3

# Install essential packages
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
pip install transformers datasets tqdm pandas matplotlib

# Quantization support library (optional)
pip install bitsandbytes

2๏ธโƒฃ Dataset Preparation

The code automatically downloads the following datasets:

  • ๐Ÿ”น nlpai-lab/kullm-v2 - Korean instruction-tuning data
  • ๐Ÿ”น beomi/KoAlpaca-v1.1a - Korean Alpaca dataset
# Datasets are automatically downloaded, no separate action needed
# Cache directory: ./datasets/cache/

3๏ธโƒฃ Start Training

# Start training with default settings
python korean_llm_advanced_v3.py

# Or run with custom configuration
python korean_llm_advanced_v3.py \
    --batch-size 2 \
    --max-steps 50000 \
    --learning-rate 5e-5

4๏ธโƒฃ Monitoring

A GUI monitoring window automatically opens during training:

  • ๐Ÿ“Š Real-time loss graph
  • ๐Ÿ’ฌ Interactive chat (generation testing)
  • ๐Ÿ“ Log viewer

๐Ÿ—๏ธ Project Structure

korean-llm-v3/
โ”œโ”€โ”€ korean_llm_advanced_v3.py    # Main training script
โ”œโ”€โ”€ README.md                      # This file
โ”œโ”€โ”€ LICENSE                        # GPL-3.0 License
โ”‚
โ”œโ”€โ”€ checkpoints/                   # Saved model checkpoints
โ”‚   โ””โ”€โ”€ korean_llm_*.pth
โ”‚
โ”œโ”€โ”€ datasets/                      # Dataset cache
โ”‚   โ”œโ”€โ”€ cache/                     # Downloaded datasets
โ”‚   โ””โ”€โ”€ datasets_manifest.json     # Metadata
โ”‚
โ””โ”€โ”€ logs/                          # Training logs and graphs
    โ”œโ”€โ”€ training.log               # Detailed log
    โ””โ”€โ”€ loss_history.json          # Loss history

๐Ÿ“Š Version History

v1 (Initial Version)

  • 541M parameter model
  • VRAM usage: ~11GB
  • Basic FP32 training

v2 (Optimization v1)

  • Expanded to 1.1B parameters
  • VRAM usage: ~23GB (1.2x initial increase)
  • BF16 + Gradient Checkpoint applied

v3 (Current) โญ

  • 1.09B parameters (v2 level)
  • VRAM usage: ~9GB (60% reduction from v2!)
  • Major Improvements:
    • 8-bit AdamW optimizer
    • Dynamic quantization support
    • Enhanced memory management
    • Faster training speed

๐Ÿ”ง Technology Stack

Core Libraries

Library Version Purpose
PyTorch 2.0+ Deep learning framework
Transformers 4.30+ Tokenizer and utilities
Datasets 2.10+ Korean dataset loading
bitsandbytes 0.40+ 8-bit quantization optimization
tqdm 4.60+ Progress display

Optional Libraries

Library Purpose
matplotlib Loss graph visualization
tkinter GUI monitoring (built-in)
pandas Data processing

๐Ÿ’ก Usage Examples

Model Loading and Text Generation

import os, argparse
from pathlib import Path
from typing import Optional, Tuple, List
import torch, torch.nn as nn, torch.nn.functional as F
from transformers import AutoTokenizer

class RMSNorm(nn.Module):
    def __init__(self, dim, eps=1e-6):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))
    def forward(self, x):
        return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps) * self.weight

def precompute_freqs_cis(head_dim: int, end: int, theta: float = 10000.0) -> Tuple[torch.Tensor, torch.Tensor]:
    freqs = 1.0 / (theta ** (torch.arange(0, head_dim, 2)[:head_dim // 2].float() / head_dim))
    t = torch.arange(end, dtype=freqs.dtype)
    freqs = torch.outer(t, freqs)
    return torch.cos(freqs), torch.sin(freqs)

def apply_rotary_emb(x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch.Tensor:
    head_dim_2 = cos.shape[-1]
    head_dim = head_dim_2 * 2
    x1 = x[..., :head_dim // 2]
    x2 = x[..., head_dim // 2:]
    cos = cos.unsqueeze(0).unsqueeze(0)
    sin = sin.unsqueeze(0).unsqueeze(0)
    return torch.cat([x1 * cos - x2 * sin, x1 * sin + x2 * cos], dim=-1)

class SwiGLU(nn.Module):
    def __init__(self, dim: int, hidden_dim: int):
        super().__init__()
        self.w1 = nn.Linear(dim, hidden_dim, bias=False)
        self.w2 = nn.Linear(hidden_dim, dim, bias=False)
        self.w3 = nn.Linear(dim, hidden_dim, bias=False)
    def forward(self, x):
        return self.w2(F.silu(self.w1(x)) * self.w3(x))

class Attention(nn.Module):
    def __init__(self, dim: int, n_heads: int):
        super().__init__()
        assert dim % n_heads == 0
        self.n_heads = n_heads
        self.head_dim = dim // n_heads
        self.wq = nn.Linear(dim, dim, bias=False)
        self.wk = nn.Linear(dim, dim, bias=False)
        self.wv = nn.Linear(dim, dim, bias=False)
        self.wo = nn.Linear(dim, dim, bias=False)
    def forward(self, x: torch.Tensor, f_cos: torch.Tensor, f_sin: torch.Tensor, kv_cache: Optional[Tuple[torch.Tensor, torch.Tensor]] = None):
        b, s, d = x.shape
        q = self.wq(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        k = self.wk(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        v = self.wv(x).view(b, s, self.n_heads, self.head_dim).transpose(1, 2)
        q = apply_rotary_emb(q, f_cos, f_sin)
        k = apply_rotary_emb(k, f_cos, f_sin)
        if kv_cache is not None:
            pk, pv = kv_cache
            k = torch.cat([pk, k], dim=2)
            v = torch.cat([pv, v], dim=2)
        new_kv = (k.detach(), v.detach())
        out = F.scaled_dot_product_attention(q, k, v, attn_mask=None, is_causal=(s > 1))
        out = out.transpose(1, 2).contiguous().view(b, s, d)
        return self.wo(out), new_kv

class TransformerBlock(nn.Module):
    def __init__(self, dim: int, n_heads: int, hidden_dim: int):
        super().__init__()
        self.attention = Attention(dim, n_heads)
        self.feed_forward = SwiGLU(dim, hidden_dim)
        self.attention_norm = RMSNorm(dim)
        self.ffn_norm = RMSNorm(dim)
    def forward(self, x, f_cos, f_sin, kv_cache=None):
        normed_x = self.attention_norm(x)
        h, new_kv = self.attention(normed_x, f_cos, f_sin, kv_cache=kv_cache)
        x = x + h
        x = x + self.feed_forward(self.ffn_norm(x))
        return x, new_kv

class KoreanLLM(nn.Module):
    def __init__(self, vocab_size: int, pad_token_id: int, dim: int = 1920, n_layers: int = 20, n_heads: int = 10, max_seq_len: int = 512):
        super().__init__()
        self.vocab_size = vocab_size
        self.pad_token_id = pad_token_id
        self.dim = dim
        self.n_heads = n_heads
        self.head_dim = dim // n_heads
        self.max_seq_len = max_seq_len
        self.embed = nn.Embedding(vocab_size, dim)
        self.layers = nn.ModuleList([TransformerBlock(dim, n_heads, int(dim * 2.5)) for _ in range(n_layers)])
        self.norm = RMSNorm(dim)
        self.output = nn.Linear(dim, vocab_size, bias=False)
        self.output.weight = self.embed.weight
        f_cos, f_sin = precompute_freqs_cis(self.head_dim, max_seq_len * 2)
        self.register_buffer("f_cos", f_cos)
        self.register_buffer("f_sin", f_sin)
    def _get_freqs(self, f, start, length):
        end = start + length
        if end > f.shape[0]:
            raise ValueError(f"ํ˜„์žฌ ์ปจํ…์ŠคํŠธ๊ฐ€ ๋„ˆ๋ฌด ๊น๋‹ˆ๋‹ค: {end} > {f.shape[0]}")
        return f[start:end]
    @torch.no_grad()
    def forward(self, tokens: torch.Tensor, kv_caches=None):
        b, s = tokens.shape
        x = self.embed(tokens)
        start_pos = 0
        if kv_caches is not None and len(kv_caches) > 0 and kv_caches[0][0] is not None:
            start_pos = kv_caches[0][0].shape[2]
        f_cos = self._get_freqs(self.f_cos, start_pos, s)
        f_sin = self._get_freqs(self.f_sin, start_pos, s)
        new_kv_caches = []
        for i, layer in enumerate(self.layers):
            cache = kv_caches[i] if kv_caches is not None else None
            x, kv = layer(x, f_cos, f_sin, kv_cache=cache)
            new_kv_caches.append(kv)
        x = self.norm(x)
        logits = self.output(x)
        return logits, new_kv_caches

def find_latest_checkpoint(checkpoint_dir="checkpoints"):
    checkpoint_dir = Path(checkpoint_dir)
    if not checkpoint_dir.exists(): return None
    files = list(checkpoint_dir.glob("korean_llm_*.pth"))
    if not files: return None
    def step_number(path):
        try: return int(path.stem.split("_")[-1])
        except ValueError: return -1
    files.sort(key=step_number)
    return files[-1]

def load_checkpoint(model, checkpoint_path, device):
    print(f"๐Ÿ“ฆ ์ฒดํฌํฌ์ธํŠธ ๋กœ๋”ฉ:\n   {checkpoint_path}")
    checkpoint = torch.load(checkpoint_path, map_location=device)
    if "model_state_dict" in checkpoint:
        state_dict = checkpoint["model_state_dict"]
        step = checkpoint.get("step", "?")
    else:
        state_dict = checkpoint
        step = "?"
    model.load_state_dict(state_dict, strict=True)
    print(f"โœ… ๋ชจ๋ธ ๋กœ๋“œ ์™„๋ฃŒ\n   ํ•™์Šต step: {step}")
    return step

@torch.no_grad()
def generate(model, tokenizer, prompt, device, max_tokens=256, temperature=0.6, top_k=40, top_p=0.95, repetition_penalty=1.15, context_limit=512):
    model.eval()
    prompt_text = f"### ์งˆ๋ฌธ: {prompt}\n### ์‘๋‹ต:"
    tokens = tokenizer.encode(prompt_text, add_special_tokens=False, return_tensors="pt").to(device)
    if tokens.shape[1] >= context_limit:
        tokens = tokens[:, -context_limit + 1:]
    output_tokens = tokens
    kv_caches = None
    eos_id = tokenizer.eos_token_id
    for _ in range(max_tokens):
        input_tokens = output_tokens if kv_caches is None else output_tokens[:, -1:]
        logits, kv_caches = model(input_tokens, kv_caches=kv_caches)
        next_logits = logits[:, -1, :]
        temperature = max(float(temperature), 1e-5)
        next_logits = next_logits / temperature
        if repetition_penalty != 1.0:
            used_tokens = set(output_tokens[0].tolist())
            for token_id in used_tokens:
                if token_id < next_logits.shape[-1]:
                    if next_logits[0, token_id] < 0:
                        next_logits[0, token_id] *= repetition_penalty
                    else:
                        next_logits[0, token_id] /= repetition_penalty
        if top_k > 0:
            k = min(int(top_k), next_logits.shape[-1])
            threshold = torch.topk(next_logits, k).values[..., -1, None]
            next_logits = torch.where(next_logits < threshold, torch.full_like(next_logits, float("-inf")), next_logits)
        probs = F.softmax(next_logits, dim=-1)
        if 0 < top_p < 1.0:
            sorted_probs, sorted_indices = torch.sort(probs, descending=True, dim=-1)
            cumulative = torch.cumsum(sorted_probs, dim=-1)
            remove = cumulative > top_p
            remove[..., 0] = False
            indices_to_remove = torch.zeros_like(probs, dtype=torch.bool)
            indices_to_remove.scatter_(-1, sorted_indices, remove)
            probs = probs.masked_fill(indices_to_remove, 0.0)
            probs = probs / (probs.sum(dim=-1, keepdim=True) + 1e-10)
        if not torch.isfinite(probs).all():
            next_token = torch.argmax(next_logits, dim=-1, keepdim=True)
        else:
            next_token = torch.multinomial(probs, num_samples=1)
        output_tokens = torch.cat([output_tokens, next_token], dim=1)
        if eos_id is not None and next_token.item() == eos_id: break
        if output_tokens.shape[1] >= context_limit: break
    generated_text = tokenizer.decode(output_tokens[0], skip_special_tokens=True)
    if "### ์‘๋‹ต:" in generated_text:
        response = generated_text.split("### ์‘๋‹ต:", 1)[1]
    else:
        response = generated_text
    if "### ์งˆ๋ฌธ:" in response:
        response = response.split("### ์งˆ๋ฌธ:", 1)[0]
    return response.strip()

def main():
    parser = argparse.ArgumentParser(description="KoreanLLM ์ฒดํฌํฌ์ธํŠธ ์ฑ„ํŒ…")
    parser.add_argument("--checkpoint", type=str, default="latest", help="์ฒดํฌํฌ์ธํŠธ ๊ฒฝ๋กœ ๋˜๋Š” latest")
    parser.add_argument("--tokenizer", type=str, default="beomi/Llama-3-Open-Ko-8B")
    parser.add_argument("--max-tokens", type=int, default=256)
    parser.add_argument("--temperature", type=float, default=0.6)
    parser.add_argument("--top-k", type=int, default=40)
    parser.add_argument("--top-p", type=float, default=0.95)
    parser.add_argument("--repetition-penalty", type=float, default=1.15)
    parser.add_argument("--cpu", action="store_true", help="๊ฐ•์ œ๋กœ CPU ์‚ฌ์šฉ")
    args = parser.parse_args()
    if args.cpu:
        device = torch.device("cpu")
    else:
        device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    print("=" * 65)
    print("๐Ÿ‡ฐ๐Ÿ‡ท KoreanLLM v3 Checkpoint Chat")
    print("=" * 65)
    print(f"๐Ÿ–ฅ๏ธ Device: {device}")
    if device.type == "cuda":
        print(f"๐ŸŽฎ GPU: {torch.cuda.get_device_name(0)}")
    else:
        print("โš ๏ธ CPU ๋ชจ๋“œ์ž…๋‹ˆ๋‹ค. 1920-dim / 20-layer ๋ชจ๋ธ์ด๋ผ ์ƒ์„ฑ์ด ๋А๋ฆด ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.")
    print("\n๐Ÿ”ค ํ† ํฌ๋‚˜์ด์ € ๋กœ๋”ฉ...")
    tokenizer = AutoTokenizer.from_pretrained(args.tokenizer, clean_up_tokenization_spaces=False)
    if tokenizer.pad_token is None or tokenizer.pad_token_id == tokenizer.eos_token_id:
        tokenizer.add_special_tokens({"pad_token": "<|pad|>"})
    vocab_size = len(tokenizer)
    pad_token_id = tokenizer.pad_token_id
    print(f"   vocab_size = {vocab_size}")
    print(f"   eos_token_id = {tokenizer.eos_token_id}")
    print(f"   pad_token_id = {pad_token_id}")
    checkpoint_path = args.checkpoint
    if checkpoint_path.lower() == "latest":
        checkpoint_path = find_latest_checkpoint()
        if checkpoint_path is None:
            print("\nโŒ checkpoints ํด๋”์—์„œ ์ฒดํฌํฌ์ธํŠธ๋ฅผ ์ฐพ์ง€ ๋ชปํ–ˆ์Šต๋‹ˆ๋‹ค.")
            print("์˜ˆ: python chat_korean_llm.py --checkpoint checkpoints/korean_llm_50000.pth")
            return
    checkpoint_path = Path(checkpoint_path)
    if not checkpoint_path.exists():
        print(f"\nโŒ ์ฒดํฌํฌ์ธํŠธ๊ฐ€ ์—†์Šต๋‹ˆ๋‹ค:\n   {checkpoint_path}")
        return
    print("\n๐Ÿง  ๋ชจ๋ธ ์ƒ์„ฑ ์ค‘...")
    model_config = dict(vocab_size=vocab_size, pad_token_id=pad_token_id, dim=1920, n_layers=20, n_heads=10, max_seq_len=512)
    dtype = torch.bfloat16 if device.type == "cuda" else torch.float32
    model = KoreanLLM(**model_config).to(device)
    if device.type == "cuda":
        model = model.to(dtype=dtype)
    try:
        step = load_checkpoint(model, checkpoint_path, device)
    except RuntimeError as e:
        print("\nโŒ ์ฒดํฌํฌ์ธํŠธ์™€ ํ˜„์žฌ ๋ชจ๋ธ ๊ตฌ์กฐ๊ฐ€ ๋งž์ง€ ์•Š์Šต๋‹ˆ๋‹ค.")
        print("   ํŠนํžˆ tokenizer์˜ vocab_size / pad_token ์„ค์ •์„ ํ™•์ธํ•˜์„ธ์š”.")
        print(f"\n์ƒ์„ธ ์˜ค๋ฅ˜:\n{e}")
        return
    model.eval()
    params = sum(p.numel() for p in model.parameters())
    print(f"๐Ÿ“Š Parameters: {params / 1e6:.1f}M")
    print("\n" + "=" * 65)
    print("๐Ÿ’ฌ ์ฑ„ํŒ… ์‹œ์ž‘")
    print("   /exit  ์ข…๋ฃŒ")
    print("   /clear ๋Œ€ํ™” ์ž…๋ ฅ ๊ธฐ๋ก ์ดˆ๊ธฐํ™”")
    print("   /info  ๋ชจ๋ธ ์ •๋ณด")
    print("=" * 65)
    while True:
        try:
            prompt = input("\n๋‚˜ > ").strip()
        except (KeyboardInterrupt, EOFError):
            print("\n\n๐Ÿ‘‹ ์ข…๋ฃŒํ•ฉ๋‹ˆ๋‹ค.")
            break
        if not prompt: continue
        if prompt.lower() in {"/exit", "/quit", "exit", "quit"}:
            print("๐Ÿ‘‹ ์ข…๋ฃŒํ•ฉ๋‹ˆ๋‹ค.")
            break
        if prompt == "/clear":
            print("๐Ÿงน ์ž…๋ ฅ ์ƒํƒœ๋ฅผ ์ดˆ๊ธฐํ™”ํ–ˆ์Šต๋‹ˆ๋‹ค.")
            continue
        if prompt == "/info":
            print(f"\n์ฒดํฌํฌ์ธํŠธ : {checkpoint_path}\nํ•™์Šต step   : {step}\nDevice      : {device}\nParameters  : {params / 1e6:.1f}M\nTemperature : {args.temperature}\nTop-k       : {args.top_k}\nTop-p       : {args.top_p}")
            continue
        print("\n๋ชจ๋ธ > ", end="", flush=True)
        try:
            response = generate(model=model, tokenizer=tokenizer, prompt=prompt, device=device, max_tokens=args.max_tokens, temperature=args.temperature, top_k=args.top_k, top_p=args.top_p, repetition_penalty=args.repetition_penalty, context_limit=512)
            print(response)
        except torch.cuda.OutOfMemoryError:
            print("\nโŒ CUDA ๋ฉ”๋ชจ๋ฆฌ๊ฐ€ ๋ถ€์กฑํ•ฉ๋‹ˆ๋‹ค.")
            print("   --max-tokens ๊ฐ’์„ ๋‚ฎ์ถ”๊ฑฐ๋‚˜ ๋‹ค๋ฅธ GPU์—์„œ ์‹คํ–‰ํ•ด๋ณด์„ธ์š”.")
        except Exception as e:
            print(f"\nโŒ ์ƒ์„ฑ ์˜ค๋ฅ˜: {type(e).__name__}: {e}")

if __name__ == "__main__":
    main()

Custom Training Configuration

from korean_llm_advanced_v3 import TrainingConfig, main

config = TrainingConfig(
    batch_size=4,                      # Batch size
    accumulation_steps=4,              # Gradient accumulation steps
    max_steps=100000,                  # Maximum training steps
    warmup_steps=1000,                 # Warmup steps
    learning_rate=3e-5,                # Learning rate
    eval_interval=5000,                # Evaluation interval
    use_bfloat16=True,                 # Whether to use BF16
    resume_from_checkpoint='latest'    # Resume from latest checkpoint
)

main(config)

โ“ FAQ (Frequently Asked Questions)

Q1: What if I only want to run inference?

A: If you have a trained checkpoint, it's simple:

import torch
from korean_llm_advanced_v3 import KoreanLLM, generate
from transformers import AutoTokenizer

model = KoreanLLM(...).to(device)
checkpoint = torch.load("checkpoints/korean_llm_50000.pth", map_location=device)
model.load_state_dict(checkpoint['model_state_dict'])
model.eval()

response = generate(model, tokenizer, prompt="์•ˆ๋…•?", max_tokens=50)

Q2: What if my GPU memory is less than 9GB?

A: Try these approaches:

  • Reduce batch size to 1
  • Shorten sequence length to 1024
  • Increase gradient accumulation steps to 16
  • Enable 8-bit quantization

Q3: How do I resume training after interruption?

A: It automatically detects the latest checkpoint:

config = TrainingConfig(
    resume_from_checkpoint='latest'  # Or specify a specific path
)
main(config)

Q4: Can I use a different Korean dataset?

A: Yes! Modify DATASETS_CONFIG in the DatasetManager class:

DATASETS_CONFIG = [
    {
        "name": "your-dataset/path",
        "split": "train",
        "text_keys": ["input", "output"]
    }
]

Q5: I get errors when running on Windows

A: Try changing num_workers setting to 0:

loader = DataLoader(dataset, batch_size=2, num_workers=0)

Q6: Can I reduce VRAM usage even more?

A: Try combining these options:

  • Memory Efficient Mode: use_bfloat16=True
  • Deeper Quantization: INT4 (additional library required)
  • LoRA Fine-tuning: Train only selective layers

Q7: Generated text quality is low

A: Check these:

  • Is the training step count sufficient? (Minimum 10,000 steps recommended)
  • Is the learning rate setting appropriate?
  • Is dataset quality good?
  • Adjust temperature parameter (0.5~1.0 recommended)

Q8: How do I convert the model to ONNX or other formats?

A: Convert from PyTorch to ONNX:

import torch.onnx

dummy_input = torch.randint(0, 50000, (1, 2048)).to(device)
torch.onnx.export(
    model, dummy_input, "korean_llm.onnx",
    input_names=['input_ids'],
    output_names=['output']
)

Q9: Is the developer actively supporting this?

A: Yes! Contact via email (seoan102410@gmail.com) for issues or feedback! ๐Ÿ’Œ

Q10: Can I use this in commercial projects?

A: It's under GPL-3.0 license, so you must disclose modifications. See the LICENSE file for details.


๐Ÿ› ๏ธ Troubleshooting

โŒ CUDA Out of Memory Error

Symptom: RuntimeError: CUDA out of memory

Solution:

# Reduce batch size
config.batch_size = 1

# Reduce maximum sequence length
config.max_seq_len = 1024

# Increase gradient accumulation
config.accumulation_steps = 16

โŒ bitsandbytes Installation Failure

Solution:

# Specify CUDA Toolkit path
CUDA_HOME=/usr/local/cuda pip install bitsandbytes

โŒ Dataset Download Failure

Solution:

# Clear cache and retry
rm -rf datasets/cache/*
python korean_llm_advanced_v3.py

๐Ÿ“ˆ Performance Optimization Tips

  1. Batch Size Adjustment: Too small slows training, too large causes VRAM shortage
  2. Gradient Accumulation: Key to increasing effective batch size
  3. Learning Rate Scheduling: Cosine Annealing improves convergence
  4. Mixed Precision: BF16 improves both speed and memory
  5. Checkpointing: Regular saving enables training resumption

๐Ÿ“ž Contact & Information


๐Ÿ“œ License

This project is distributed under GPL-3.0 License.

GNU GENERAL PUBLIC LICENSE
Version 3, 29 June 2007

Copyright (C) 2024 seoan1024

This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.

๐Ÿ“– Full License: LICENSE


๐Ÿค Contributing

Bug reports, feature suggestions, and pull requests are always welcome!

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

๐Ÿ™ Acknowledgments

  • ๐ŸŽฏ Korean LLM Community - Valuable feedback and contributions
  • ๐Ÿ“š Hugging Face - Transformers & Datasets libraries
  • ๐Ÿ”ง bitsandbytes - Quantization and optimization solutions
  • ๐Ÿš€ PyTorch - Open-source deep learning framework

๐Ÿ“š References

Korean NLP

Optimization Techniques

Large Language Models


โญ If this project was helpful, please click the star button!

โฌ† Back to Top

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Papers for seoan1024/korean-llm-v3