MoWorld: A Flash World Model
Team Moxin
Camera-controlled world generation Β· MoWorld Base checkpoint
We present MoWorld, a cost-efficient, high-performance Flash World Model for real-time interactive video generation on Neural Processing Units (NPUs). MoWorld jointly optimizes data, algorithms, systems, and hardware across an end-to-end pipeline spanning data generation and pretraining.
3D-Native Data Engine: A scalable data pipeline constructs geometrically consistent training samples from real-world and synthetic environments through geometry completion, quality control, trajectory verification, and vision-language annotation.
Curriculum Cross-Frame Pretraining: Training progressively extends from short clips to sequences of up to 2,000 frames, improving camera control, temporal consistency, and long-horizon spatial memory.
Few-Step Autoregressive Generation: Autoregressive flow-matching pretraining and Self-Forcing distillation reduce the standard 50-step sampling process to 4-step generation.
Real-Time NPU Inference: Pipeline-, parallelism-, and kernel-level optimizations enable up to 50 FPS on NPUs in the configuration reported in the paper, with an average inference cost reported at 30%β50% of existing world-model solutions.
Given a first frame, a text prompt, and camera controls, it generates long-horizon video while preserving visual quality, spatial consistency, and control responsiveness. Visit the project page to explore more world-generation, reconstruction, and downstream application results.
π¦ Model Release
| Model | Component | Format | Weight size | Download |
|---|---|---|---|---|
| MoWorld Base | High-noise expert | Safetensors | 59.17 GB / 55.11 GiB | Hugging Face |
Release scope: This release provides the high-noise expert checkpoint for the training workflow below. The capabilities described above refer to the full MoWorld system. It does not include a standalone inference pipeline, the low-noise expert, a VAE, a text encoder, or tokenizer assets. The cover illustrates MoWorld scenes; it is not a benchmark of this individual checkpoint.
βοΈ Quick Start
Prepare a Linux environment with Python 3.10+, PyTorch 2.7.1, and compatible CANN and torch_npu packages.
Installation
Clone the repository:
git clone --recursive https://github.com/Moxin-Tech/moworld1.0.git
cd moworld1.0
export MOWORLD_ROOT="$PWD"
Install dependencies in your Ascend environment:
python -m pip install -e ./MindSpeed
git clone --branch core_v0.12.1 https://github.com/NVIDIA/Megatron-LM.git ../Megatron-LM
export PYTHONPATH="$(cd ../Megatron-LM && pwd):${PYTHONPATH:-}"
python -m pip install -e ./moworld
Download the MoWorld Checkpoint:
python -m pip install -U huggingface_hub
export MOWORLD_WEIGHTS="$MOWORLD_ROOT/models/moworld"
hf download "moworld1/moworld_base" --local-dir "$MOWORLD_WEIGHTS"
Prepare your training data and any required tokenizer/encoder assets in the layout required by the MoWorld source code, then run the commands below from moworld/:
export MODEL_PATH="$MOWORLD_ROOT/models/base"
export MINDSPEED_PATH="$MOWORLD_ROOT/MindSpeed"
cd "$MOWORLD_ROOT/moworld"
Training
Prepare your feature data and examples/moworld/local_data/train_data.json following the data preparation instructions in the MoWorld source repository. The published checkpoint contains the high-noise expert only. Link it into the converter's expected layout, convert it to DCP, then train the high-noise expert:
mkdir -p "$MODEL_PATH/high_noise_model"
ln -s "$MOWORLD_WEIGHTS/high_noise_model.safetensors" \
"$MODEL_PATH/high_noise_model/diffusion_pytorch_model.safetensors"
EXPERTS=high_noise_model SOURCE_ROOT="$MODEL_PATH" \
bash examples/moworld/convert_weights.sh
export MM_DATA="$PWD/examples/moworld/local_data/train_data.json"
export TRAIN_ITERS=10 SAVE_INTERVAL=10
bash examples/moworld/pretrain_high.sh
This is a short pretraining run on 16 NPUs. Checkpoints are saved under checkpoints/train/high_noise_model/.
The root config.json in this Hugging Face repository is a checkpoint manifest containing the weight filename, size, and SHA256. It is not the expert architecture configuration; prepare any additional configurations required by the converter and training code separately.
π License
The original MoWorld content in this repository is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0).
You may use, share, and adapt the licensed material for non-commercial purposes with appropriate attribution. Commercial use requires separate permission from Team Moxin. Third-party components remain subject to their original licenses and usage terms.
π Citation
If you find this work useful for your research, please cite our paper:
@article{moworld2026,
title = {MoWorld: A Flash World Model},
author = {{Team Moxin} and Deyi Ji and Tianrun Chen and Xin Zhang and Jiale Yang and Qi Zhu and An Zhao and Zihao Xie and Han Wang and Xuanyi Liu and Yixiang Zhou and Pei Liu and Yi Tan and Cheng Chen and Dayi Zhu and Mingyu Wei and Hanjie Xu and Jun Liao and Siqi Li and Lingyu Lu and Hongye Fang and Hongming Tan and Youjiang Zhu and Taiyu Zhang and Zejian Li and Chaotao Ding and Zhipeng Liang and Wenxuan Song and Yi Li and Baochuan Yang and Xin Jiang and Ben Feng and Jingyuan Zou and Yanlin Liu and Rong Shi and Lingfeng Li and Liyi Yao and Lanyun Zhu and Yunhe Pan and Lingyun Sun},
journal = {arXiv preprint arXiv:2607.06216},
year = {2026}
}
- Downloads last month
- 10,281