Qwen3-8B | Self-CTRL | Privacy | LoRA | v1

A single-rule research LoRA adapter for Qwen/Qwen3-8B, trained using an adaptation of Self-CTRL's released behavior-only objective. This package contains adapter weights, not the base model. It is the final ckpt_160 of the privacy pilot, separate from the Llama refusal reproduction.

Identity

Field Value
Base model Qwen/Qwen3-8B
Training method Self-CTRL, behavior-only normative-rule pilot
Rule Privacy (privacy)
Artifact PEFT LoRA adapter, rank 16, alpha 32, configured dropout 0.05
Release Git tag _v1
Checkpoint ckpt_160: 160 example presentations, five epochs over 32 training tasks
Development data 8 separate validation tasks
Weight size 174.66 MB

Repository names follow BASE-selfctrl-RULE-lora; each rule has its own repository. _v1 pins the shared release. See provenance.json for source identity, hashes and validation scope, and training_config.json for portable training metadata. These metadata are descriptive; they are not a standalone training launcher.

Purpose and training

The target is agreement between generated task behavior and a separately elicited normative explanation concerning privacy. The pilot uses independently authored paired prompts, the released local jury and auxiliary engagement reward, an explanation-side anchor, and instruction-tuning regularization. Training uses eight candidates, temperature 1, top-p 0.9, learning rate 1e-5, and the released behavior-only coefficients. Configured settings should not be confused with effective semantics: the audited released implementation has SFT truncation and scoring-time dropout caveats.

Upstream source: Self-CTRL, commit 64e5611d09f5a9160f90113f86839df1b5ba3755. Project implementation: self_consistency_safety. This pilot is not an exact reproduction of the paper's original refusal dataset experiment.

Evaluation and limitations

On the separate 24-case task-only enactment panel, raw Luna-graded adherence was 14/24 before training → 11/24 after training. See evaluation_summary.json for useful-completion metrics. Counts retain the original grades; independent transcript review identified judgment/reference caveats and did not rescore the panel. One small panel and one training run do not establish generalization or dependable safety improvement. Do not interpret the model name as a safety certification.

Load the adapter

Install requirements.txt in a suitable environment. This is a public repository under narutatsuri; collaborators can download the adapter directly. Download this repository at revision _v1, or pass its full repository ID to load_adapter.py:

python load_adapter.py narutatsuri/qwen3-8b-selfctrl-privacy-lora --revision _v1

The script loads the full base model and attaches this adapter using PEFT. Sufficient memory for the base model is required; small adapter files do not make base-model inference small. The generation command is a loading smoke example, not the study's evaluation protocol.

The historical base revision is not verified. The loader uses the official base model ID; provenance does not claim a bitwise-pinned base download.

The source adapter was successfully attached and merged using Transformers 5.5.4 / PEFT 0.21.0; native serialized tensors were verified and the merged model was used in inference. This sharing copy has byte-identical weights, readable safetensors and a validated PEFT configuration. A fresh generation from the staged package has not been run; this card does not claim otherwise.

Files and attribution

adapter_model.safetensors and adapter_config.json are the inference artifact. Supporting files document the training, evaluation, provenance and loader. No base weights or optimizer states are included. Base-model licensing and attribution remain those of Qwen/Qwen3-8B; this private research package does not assign a new license to project contributions. Original training artifacts remain unchanged.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for narutatsuri/qwen3-8b-selfctrl-privacy-lora

Finetuned
Qwen/Qwen3-8B
Adapter
(2239)
this model