Instructions to use whisperle/harnessGRPO-checkpoints with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use whisperle/harnessGRPO-checkpoints with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("whisperle/harnessGRPO-checkpoints", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Multi-Harness RL checkpoints
Every trained checkpoint behind What Does Multi-Harness RL Learn? Credit Assignment and
Portability in Coding Agents. All are full-parameter fine-tunes of Qwen/Qwen3-8B in bf16,
saved as sharded safetensors with their tokenizer β load any of them with subfolder=:
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "whisperle/harnessGRPO-checkpoints"
SUB = "crossed/R0-sft-cross-s17"
model = AutoModelForCausalLM.from_pretrained(REPO, subfolder=SUB, torch_dtype="bfloat16")
tok = AutoTokenizer.from_pretrained(REPO, subfolder=SUB)
Traces, verdicts and frozen analysis outputs live in the companion dataset
whisperle/harnessGRPO-traces.
What the arms mean
The study separates exposure (which harnesses' trajectories the policy trains on) from credit (how outcomes from different harnesses compete inside the GRPO advantage group). Records, tokens and rewards are held fixed across arms so the credit rule is the only difference.
Each arm is identified by the manifest stream it consumed, not by a training flag: the
advantage is precomputed into the stream, so crossdown and residmatched both run with
grouping=cross and differ from PlainCross only in the advantage column. Check
run_config.json -> samples to identify any folder unambiguously.
| folder | paper name | stream | grouping |
|---|---|---|---|
*within*, s2within_* |
Within | samples_within.jsonl |
within |
*cross*, s4cross_* |
Cross / PlainCross | samples_cross.jsonl |
cross |
*crossdown*, s4down_* |
coverage-matched | samples_crossdown.jsonl |
cross |
*residmatched*, s5resid_* |
residualized | samples_resid.jsonl |
cross |
t14_placebo_seed0 |
placebo control | samples_resid_placebo.jsonl |
cross |
ON1, ON2 |
online round 1 / 2 | samples_cross.jsonl (re-collected) |
cross |
Every arm above trains at lr=3e-7, objective=grpo, one epoch over its frozen manifest.
crossed/ β the primary study (Verified-500 roster, one RL seed per arm)
| folder | what it is |
|---|---|
S1-final-s17 |
the SFT start every RL arm branches from (LR 1e-6) |
S1-pilot-lr3e-6-s17, S1-pilot-lr1e-5-s17 |
the two LR pilots not selected |
R0-sft-{within,cross,crossdown,residmatched}-s17 |
the four credit rules, seed 17 |
R0-sft-{within,cross}-s{29,41} |
seed-variance arms |
ON1-cross-s17, ON2-cross-s17 |
the online control, rounds 1 and 2 |
three-seed/ β the secondary study (audit-200 roster, three RL seeds per arm)
sft is the start; s2within_seed{0,1,2}, s4cross_seed{0,1,2}, s4down_seed{0,1,2} and
s5resid_seed{0,1,2} are the four credit rules at three seeds. single_{aider,openhands, qwen_code,swe_agent} and weakonly* are single-harness exposure ablations; t14_placebo_seed0
is the placebo control.
Provenance
Each folder carries run_config.json, run_metadata.json, diagnostics.jsonl and
heartbeat.log from its training run: objective, grouping rule, learning rate, KL coefficient,
seed, step count, the git commit, and the SHA-256 of both the frozen sample manifest
(input_hash / manifest_hash) and the task roster. Those hashes tie a checkpoint to the exact
manifest in the dataset repo. Absolute cluster paths in the two JSON files were replaced with
logical <release>/β¦ and <checkpoint>/β¦ names; nothing else was altered.
Caveats
- Read the paper before reading the deltas. Rollout generation is not bit-reproducible (vLLM continuous batching flips ~6% of per-task outcomes between identical reruns), so no single-run difference under Β±6pp between two of these checkpoints is evidence of anything.
- These are research artifacts from a study whose headline is a bounded null on held-out-harness
portability. They are not intended or tuned as general-purpose coding models, and none of them
should be expected to outperform the
Qwen3-8Bbase outside this study's setup.