Instructions to use narutatsuri/qwen3-8b-selfctrl-privacy-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use narutatsuri/qwen3-8b-selfctrl-privacy-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B") model = PeftModel.from_pretrained(base_model, "narutatsuri/qwen3-8b-selfctrl-privacy-lora") - Notebooks
- Google Colab
- Kaggle
Qwen3-8B | Self-CTRL | Privacy | LoRA | v1
A single-rule research LoRA adapter for Qwen/Qwen3-8B, trained using an adaptation of Self-CTRL's released behavior-only objective. This package contains adapter weights, not the base model. It is the final ckpt_160 of the privacy pilot, separate from the Llama refusal reproduction.
Identity
| Field | Value |
|---|---|
| Base model | Qwen/Qwen3-8B |
| Training method | Self-CTRL, behavior-only normative-rule pilot |
| Rule | Privacy (privacy) |
| Artifact | PEFT LoRA adapter, rank 16, alpha 32, configured dropout 0.05 |
| Release | Git tag _v1 |
| Checkpoint | ckpt_160: 160 example presentations, five epochs over 32 training tasks |
| Development data | 8 separate validation tasks |
| Weight size | 174.66 MB |
Repository names follow BASE-selfctrl-RULE-lora; each rule has its own repository. _v1 pins the shared release. See provenance.json for source identity, hashes and validation scope, and training_config.json for portable training metadata. These metadata are descriptive; they are not a standalone training launcher.
Purpose and training
The target is agreement between generated task behavior and a separately elicited normative explanation concerning privacy. The pilot uses independently authored paired prompts, the released local jury and auxiliary engagement reward, an explanation-side anchor, and instruction-tuning regularization. Training uses eight candidates, temperature 1, top-p 0.9, learning rate 1e-5, and the released behavior-only coefficients. Configured settings should not be confused with effective semantics: the audited released implementation has SFT truncation and scoring-time dropout caveats.
Upstream source: Self-CTRL, commit 64e5611d09f5a9160f90113f86839df1b5ba3755. Project implementation: self_consistency_safety. This pilot is not an exact reproduction of the paper's original refusal dataset experiment.
Evaluation and limitations
On the separate 24-case task-only enactment panel, raw Luna-graded adherence was 14/24 before training → 11/24 after training. See evaluation_summary.json for useful-completion metrics. Counts retain the original grades; independent transcript review identified judgment/reference caveats and did not rescore the panel. One small panel and one training run do not establish generalization or dependable safety improvement. Do not interpret the model name as a safety certification.
Load the adapter
Install requirements.txt in a suitable environment. This is a public repository under narutatsuri; collaborators can download the adapter directly. Download this repository at revision _v1, or pass its full repository ID to load_adapter.py:
python load_adapter.py narutatsuri/qwen3-8b-selfctrl-privacy-lora --revision _v1
The script loads the full base model and attaches this adapter using PEFT. Sufficient memory for the base model is required; small adapter files do not make base-model inference small. The generation command is a loading smoke example, not the study's evaluation protocol.
The historical base revision is not verified. The loader uses the official base model ID; provenance does not claim a bitwise-pinned base download.
The source adapter was successfully attached and merged using Transformers 5.5.4 / PEFT 0.21.0; native serialized tensors were verified and the merged model was used in inference. This sharing copy has byte-identical weights, readable safetensors and a validated PEFT configuration. A fresh generation from the staged package has not been run; this card does not claim otherwise.
Files and attribution
adapter_model.safetensors and adapter_config.json are the inference artifact. Supporting files document the training, evaluation, provenance and loader. No base weights or optimizer states are included. Base-model licensing and attribution remain those of Qwen/Qwen3-8B; this private research package does not assign a new license to project contributions. Original training artifacts remain unchanged.
- Downloads last month
- 15