Add a 4/16/4 recurrent 1B architecture and 100-step trial results

#1

This adds a recurrent Nano-1B candidate with a 4/16/4 prelude, shared core, and coda.

The model stores 24 independent layers as a 4-layer prelude, 16-layer shared core, and 4-layer coda. The default K=2 executes 40 layers while retaining 1,076,700,672 parameters. K=1 uses the original sequential schedule. Both core passes participate in backpropagation; cached decoding uses separate slots for each execution and rejects depth changes during continuation. The tokenizer, global GQA, DeepSeek V4 components, and weight shapes are retained.

100-step trial

Both runs used the same initialization seed, token order, 6,553,600 training tokens, 2048-token sequences, optimizer settings, fixed evaluation samples, and CoreWeave 4×GB200 node. The recurrent run trained at fixed K=2 and evaluated both depths every 20 steps.

Metric at step 100 Base 1B (Vanilla) Recurrent 1B (K=2)
Prelude / core / coda — 4 / 16 / 4
Physical / effective layers 24 / 24 24 / 40
Training CE 6.9774 7.1601
Overall validation CE 7.2836 7.4369
Overall validation BPB 2.2104 2.2568
Mean seconds/step, steps 11–100 0.741 0.920
Peak allocated GiB/GPU 29.97 33.13

On the recurrent checkpoint, K=1 evaluation gives CE 7.5839 / BPB 2.3012. The second pass helps that checkpoint, but this short trial does not beat the separately trained vanilla model. This is a matched-token comparison, not a matched-compute experiment.

Dense baseline W&B · Recurrent W&B.

Validation

  • 11 local behavioral tests passed against the exact PR payload: K=1 equivalence with the original dense initialization; K=2 versus independent unrolling and summed shared gradients; checkpointing; packed/padded attention; cached generation; save/reload and optimizer continuation; invalid settings; FlashAttention varlen dispatch with external kernels stubbed.
  • Full-size meta count passed at both K=1 and K=2. The count script reports 3 GiB and 5 GiB of BF16 KV storage at 128K, respectively, without duplicating parameter counts.
  • Full-size 100-step job 101752 completed with finite losses and nonzero gradient/update probes in all three sections. Independent audit and W&B server readback passed: 100 training points and 72 evaluation records.
  • All files are architecture/configuration/documentation assets. No tests, training data, credentials, or model/optimizer weights are uploaded to this model repository.

Actual FlashAttention kernels, 128K training, and long-run convergence are outside this trial's validation.

The independent padding/EOS fix is tracked in PR #2. This PR keeps the existing padding configuration.

bowang0911 changed pull request title from Add a 4/16/4 recurrent 1B candidate and preserve EOS supervision with dedicated padding to Add a 4/16/4 recurrent 1B architecture and 100-step trial results
Cannot merge
This branch has merge conflicts in the following files:
  • README.md

Sign up or log in to comment