Instructions to use bowang0911/Nano-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bowang0911/Nano-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bowang0911/Nano-1B", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("bowang0911/Nano-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bowang0911/Nano-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bowang0911/Nano-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bowang0911/Nano-1B
- SGLang
How to use bowang0911/Nano-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bowang0911/Nano-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bowang0911/Nano-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bowang0911/Nano-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bowang0911/Nano-1B with Docker Model Runner:
docker model run hf.co/bowang0911/Nano-1B
Add a 4/16/4 recurrent 1B architecture and 100-step trial results
This adds a recurrent Nano-1B candidate with a 4/16/4 prelude, shared core, and coda.
The model stores 24 independent layers as a 4-layer prelude, 16-layer shared core, and 4-layer coda. The default K=2 executes 40 layers while retaining 1,076,700,672 parameters. K=1 uses the original sequential schedule. Both core passes participate in backpropagation; cached decoding uses separate slots for each execution and rejects depth changes during continuation. The tokenizer, global GQA, DeepSeek V4 components, and weight shapes are retained.
100-step trial
Both runs used the same initialization seed, token order, 6,553,600 training tokens, 2048-token sequences, optimizer settings, fixed evaluation samples, and CoreWeave 4×GB200 node. The recurrent run trained at fixed K=2 and evaluated both depths every 20 steps.
| Metric at step 100 | Base 1B (Vanilla) | Recurrent 1B (K=2) |
|---|---|---|
| Prelude / core / coda | — | 4 / 16 / 4 |
| Physical / effective layers | 24 / 24 | 24 / 40 |
| Training CE | 6.9774 | 7.1601 |
| Overall validation CE | 7.2836 | 7.4369 |
| Overall validation BPB | 2.2104 | 2.2568 |
| Mean seconds/step, steps 11–100 | 0.741 | 0.920 |
| Peak allocated GiB/GPU | 29.97 | 33.13 |
On the recurrent checkpoint, K=1 evaluation gives CE 7.5839 / BPB 2.3012. The second pass helps that checkpoint, but this short trial does not beat the separately trained vanilla model. This is a matched-token comparison, not a matched-compute experiment.
Dense baseline W&B · Recurrent W&B.
Validation
- 11 local behavioral tests passed against the exact PR payload: K=1 equivalence with the original dense initialization; K=2 versus independent unrolling and summed shared gradients; checkpointing; packed/padded attention; cached generation; save/reload and optimizer continuation; invalid settings; FlashAttention varlen dispatch with external kernels stubbed.
- Full-size meta count passed at both K=1 and K=2. The count script reports 3 GiB and 5 GiB of BF16 KV storage at 128K, respectively, without duplicating parameter counts.
- Full-size 100-step job 101752 completed with finite losses and nonzero gradient/update probes in all three sections. Independent audit and W&B server readback passed: 100 training points and 72 evaluation records.
- All files are architecture/configuration/documentation assets. No tests, training data, credentials, or model/optimizer weights are uploaded to this model repository.
Actual FlashAttention kernels, 128K training, and long-run convergence are outside this trial's validation.
The independent padding/EOS fix is tracked in PR #2. This PR keeps the existing padding configuration.