hindsight-v4
A small, calibrated, non-generative verifier for coding-agent sessions. Given the canonical event trace of a session (user turns, agent messages, tool calls and results, diffs) it predicts, at every user-turn boundary, whether the user's next turn will be a correction, and of which kind (not-done, unrequested, rule-violation, blind-action, drift, communication, damage, quality, ungrounded). Trained on hindsight labels (what the user actually did next), never on the agent's own claims. Encoder: UniXCoder-base fine-tuned end to end; level 2: a GRU over the last 24 event vectors with an attention readout conditioned on the latest user request; heads Platt-calibrated on a held-out split. Inputs are the event texts and a structural feature vector only: no labelers, no source id (the unknown-source slot), so the model runs on any harness's trace.
Held-out AUROC (Platt-scaled, test split by repository and user)
| head | all sources | Claude Code local | Codex local | SWE-chat |
|---|---|---|---|---|
| correction | 0.732 | 0.769 | 0.758 | 0.722 |
| not_done | 0.798 | 0.817 | 0.667 | 0.793 |
| unrequested | 0.794 | 0.885 | 0.653 | 0.811 |
| rule_violation | 0.735 | 0.884 | 0.806 | 0.727 |
| blind_action | 0.675 | 0.297 | 0.842 | 0.662 |
| communication | 0.630 | 0.679 | 0.852 | 0.597 |
| quality | 0.699 | 0.786 | 0.754 | 0.662 |
| ungrounded | 0.652 | 0.566 | 0.773 | 0.657 |
Coverage at 80% precision on the any-correction head is 0.04; use it as a rank or reject signal.
The other head is not usable.
Files
encoder/ (transformers, safetensors), model.pt (level-2 core and heads, keys prefixed ver.),
config.json (dims, window, source list), results.json (metrics, Platt parameters),
quant_bench.json (fp32 vs int8 vs ONNX on CPU: keep fp32).
onnx/ is the deployment export (gpu/export_onnx.py): encoder.onnx (token ids and mask to a
normalized event vector), level2.onnx (a 24-event window to the any-correction logit and the nine
class logits), tokenizer.json, hindsight.json (feature layout, window, Platt parameters). It is
what octolib::hindsight (Rust, ONNX Runtime, CPU) loads; it runs without the tier-1 labelers
(those feature slots are zero) and feeds the unknown-source slot.
- Downloads last month
- 27
Model tree for muvon/hindsight
Base model
microsoft/unixcoder-base