ASTERIZER · LUNA Model Family

LUNA-300M Program

A ~303M-parameter, English-first causal language model — pretrained from scratch on the 4.5B-token LUNA_PreTrain corpus. The scaled-up sibling of LUNA-100M.

≈303M
Parameters
31,052
Pretrain steps
4.5B
Tokens trained
1,024
Context window

Status

✓ TRAINED (pretraining complete) — final fp32 weights available at pretrained/final/lit_model.pth. Instruction tuning (RAG + MCP SFT, mirroring LUNA-100M) is planned on top of this checkpoint.

Architecture

PropertyValue
Layers20
Hidden size1024
Heads16
Context1,024 tokens
Vocab50,304 (Pythia-160m tokenizer)
Precisionbf16 / fp16

Training

PropertyValue
CorpusLUNA_PreTrain (4,515,286,950 tokens)
OptimizerAdamW (wd 0.1, clip 1.0)
LR3e-4 → 3e-5 cosine
Global batch120
Steps31,052

Checkpoints

FileDescription
pretrained/final/lit_model.pthFinal fp32 weights (1.2 GB) — recommended
pretrained/step-00031052/lit_model.pthFinal logged step
pretrained/step-00031000/lit_model.pthMilestone checkpoint
pretrained/latest.ptWeights + optimizer state (3.6 GB)