ASIL Qwen3.5-9B RL

This model is part of the ASIL paper release (Findings of EMNLP 2026).
Project page: https://sharryxr.github.io/ASIL/
Code: https://github.com/sharryXR/ASIL

This repository contains the ASIL v0.1.0 paper release checkpoint for ASIL Qwen3.5-9B RL.

  • Release: v0.1.0
  • Selected checkpoint: global_step_140_actor_hf
  • Source path: /public/LLM_model_dataset/rl_temp/asil_operational_benchmark_a100_20260514_144617/results/rl/qwen35_9b_agentic_a800/qwen35_9b_rl_from_sft543_resume80_p100_307542_20260523_224030/checkpoints/global_step_140_actor_hf
  • Base/init checkpoint: /public/LLM_model_dataset/rl_temp/asil_sft_rl_a100_20260513_173133/results/sft_train/qwen35_9b_base_sft_v0_v1_6ep_20260520_000159/checkpoints/global_step_543; resumed from Verl global_step_80 in the first-stage RL run
  • Training data: rl_learnable_v4_320_80; 320 train / 80 valid task prompts
  • Prepared at: 2026-07-30T18:02:32+08:00

The repo root contains the HF-loadable checkpoint files (config.json, tokenizer files, generation_config.json, and *.safetensors). Training-only artifacts such as optimizer state, scheduler state, trainer state, logs, caches, wandb output, and credentials are intentionally excluded.

See checkpoint_metadata.json and SHA256SUMS for provenance and file checksums.

Downloads last month
211
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sharryXR/asil-qwen35-9b-rl

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(858)
this model

Collection including sharryXR/asil-qwen35-9b-rl

Paper for sharryXR/asil-qwen35-9b-rl