hindsight-v4

A small, calibrated, non-generative verifier for coding-agent sessions. Given the canonical event trace of a session (user turns, agent messages, tool calls and results, diffs) it predicts, at every user-turn boundary, whether the user's next turn will be a correction, and of which kind (not-done, unrequested, rule-violation, blind-action, drift, communication, damage, quality, ungrounded). Trained on hindsight labels (what the user actually did next), never on the agent's own claims. Encoder: UniXCoder-base fine-tuned end to end; level 2: a GRU over the last 24 event vectors with an attention readout conditioned on the latest user request; heads Platt-calibrated on a held-out split. Inputs are the event texts and a structural feature vector only: no labelers, no source id (the unknown-source slot), so the model runs on any harness's trace.

Held-out AUROC (Platt-scaled, test split by repository and user)

head all sources Claude Code local Codex local SWE-chat
correction 0.732 0.769 0.758 0.722
not_done 0.798 0.817 0.667 0.793
unrequested 0.794 0.885 0.653 0.811
rule_violation 0.735 0.884 0.806 0.727
blind_action 0.675 0.297 0.842 0.662
communication 0.630 0.679 0.852 0.597
quality 0.699 0.786 0.754 0.662
ungrounded 0.652 0.566 0.773 0.657

Coverage at 80% precision on the any-correction head is 0.04; use it as a rank or reject signal. The other head is not usable.

Files

encoder/ (transformers, safetensors), model.pt (level-2 core and heads, keys prefixed ver.), config.json (dims, window, source list), results.json (metrics, Platt parameters), quant_bench.json (fp32 vs int8 vs ONNX on CPU: keep fp32).

onnx/ is the deployment export (gpu/export_onnx.py): encoder.onnx (token ids and mask to a normalized event vector), level2.onnx (a 24-event window to the any-correction logit and the nine class logits), tokenizer.json, hindsight.json (feature layout, window, Platt parameters). It is what octolib::hindsight (Rust, ONNX Runtime, CPU) loads; it runs without the tier-1 labelers (those feature slots are zero) and feeds the unknown-source slot.

Downloads last month
27
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for muvon/hindsight

Quantized
(6)
this model