Multi-Harness RL checkpoints

Every trained checkpoint behind What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents. All are full-parameter fine-tunes of Qwen/Qwen3-8B in bf16, saved as sharded safetensors with their tokenizer β€” load any of them with subfolder=:

from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "whisperle/harnessGRPO-checkpoints"
SUB  = "crossed/R0-sft-cross-s17"
model = AutoModelForCausalLM.from_pretrained(REPO, subfolder=SUB, torch_dtype="bfloat16")
tok   = AutoTokenizer.from_pretrained(REPO, subfolder=SUB)

Traces, verdicts and frozen analysis outputs live in the companion dataset whisperle/harnessGRPO-traces.

What the arms mean

The study separates exposure (which harnesses' trajectories the policy trains on) from credit (how outcomes from different harnesses compete inside the GRPO advantage group). Records, tokens and rewards are held fixed across arms so the credit rule is the only difference.

Each arm is identified by the manifest stream it consumed, not by a training flag: the advantage is precomputed into the stream, so crossdown and residmatched both run with grouping=cross and differ from PlainCross only in the advantage column. Check run_config.json -> samples to identify any folder unambiguously.

folder paper name stream grouping
*within*, s2within_* Within samples_within.jsonl within
*cross*, s4cross_* Cross / PlainCross samples_cross.jsonl cross
*crossdown*, s4down_* coverage-matched samples_crossdown.jsonl cross
*residmatched*, s5resid_* residualized samples_resid.jsonl cross
t14_placebo_seed0 placebo control samples_resid_placebo.jsonl cross
ON1, ON2 online round 1 / 2 samples_cross.jsonl (re-collected) cross

Every arm above trains at lr=3e-7, objective=grpo, one epoch over its frozen manifest.

crossed/ β€” the primary study (Verified-500 roster, one RL seed per arm)

folder what it is
S1-final-s17 the SFT start every RL arm branches from (LR 1e-6)
S1-pilot-lr3e-6-s17, S1-pilot-lr1e-5-s17 the two LR pilots not selected
R0-sft-{within,cross,crossdown,residmatched}-s17 the four credit rules, seed 17
R0-sft-{within,cross}-s{29,41} seed-variance arms
ON1-cross-s17, ON2-cross-s17 the online control, rounds 1 and 2

three-seed/ β€” the secondary study (audit-200 roster, three RL seeds per arm)

sft is the start; s2within_seed{0,1,2}, s4cross_seed{0,1,2}, s4down_seed{0,1,2} and s5resid_seed{0,1,2} are the four credit rules at three seeds. single_{aider,openhands, qwen_code,swe_agent} and weakonly* are single-harness exposure ablations; t14_placebo_seed0 is the placebo control.

Provenance

Each folder carries run_config.json, run_metadata.json, diagnostics.jsonl and heartbeat.log from its training run: objective, grouping rule, learning rate, KL coefficient, seed, step count, the git commit, and the SHA-256 of both the frozen sample manifest (input_hash / manifest_hash) and the task roster. Those hashes tie a checkpoint to the exact manifest in the dataset repo. Absolute cluster paths in the two JSON files were replaced with logical <release>/… and <checkpoint>/… names; nothing else was altered.

Caveats

  • Read the paper before reading the deltas. Rollout generation is not bit-reproducible (vLLM continuous batching flips ~6% of per-task outcomes between identical reruns), so no single-run difference under Β±6pp between two of these checkpoints is evidence of anything.
  • These are research artifacts from a study whose headline is a bounded null on held-out-harness portability. They are not intended or tuned as general-purpose coding models, and none of them should be expected to outperform the Qwen3-8B base outside this study's setup.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for whisperle/harnessGRPO-checkpoints

Finetuned
Qwen/Qwen3-8B
Finetuned
(2130)
this model