Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
Abstract
We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that sits atop a fully frozen Gemma 4 E2B model. The base model is never updated; only the correction module learns, via supervised fine-tuning followed by reference-free DPO on 83,400 error-correction pairs. On a 60-question domain exam (CEHRI: Certified Human-Robot Intelligence, covering facts, arithmetic, and implicit-goal reasoning), CRN v2 corrects 53.3% of base-model errors (reworded variant: 43.3%) while showing no degradation on tested capability benchmarks (MMLU/BoolQ N=200; car-wash N=8). A LoRA baseline at the matched CRN v1 budget (6.6M params, rank 19) achieves 83.3% correction but suffers 30-75% capability loss on the same benchmarks -- the correction-capability tradeoff. An ablation shows that the KL preservation term (lambda=0.1) is critical: lowering it to 0.01 degrades correction to 35.0%. A hidden-state injection variant at earlier layers (1.6M params, SFT-only) reaches 50.0%/55.8% but does not exceed logit correction; shallower injection (layer 4) drops to 30.0%/28.3%; multi-depth logit correction (~35M) reaches only 40%; and longer training (5,000 SFT + 2,000 DPO) stays at 53.3% -- none of the alternative configurations we tested exceeded the rank-128 logit result, consistent with a best-achieved result of ~53% rather than a floor. This is a study of a design principle (frozen base + logit correction + KL anchoring), not a claim of architectural novelty. All code, main-result weights, and evaluation scripts are released (deep variant as code only -- no trained deep checkpoints).
Community
Author here, happy to answer questions!
This paper asks a deliberately narrow question: if you freeze a language model completely and only train a small correction module on top, how far can you get and what do you keep?
The headline tradeoff, on a 60-question domain exam with a frozen Gemma-4-E2B:
CRN v2 (34M params): 53.3% correction, zero degradation on MMLU/BoolQ/car-wash
LoRA (6.6M params): 83.3% correction, but MMLU 62.5%→32%, car-wash 75%→0%
The part I'm most interested in discussing: we tried five ways to break past ~53% (rank, 5k-step training, multi-depth, deep injection at two layers) and none of them worked, we published that as the finding rather than burying it. DPO at depth 7 even destroyed capabilities outright (MMLU 13%). If anyone has ideas for cracking the frozen-base ceiling without touching weights, I'd love to hear them.
Everything is reproducible: code + weights + eval scripts at github.com/eulogik/prajna and eulogik/Prajna-CRNv2. All trained on a Mac Mini M4, no GPU.
This paper asks a deliberately narrow question: if you freeze a language model completely and only train a small correction module on top, how far can you get and what do you keep?
The headline tradeoff, on a 60-question domain exam with a frozen Gemma-4-E2B:
CRN v2 (34M params): 53.3% correction, zero degradation on MMLU/BoolQ/car-wash
LoRA (6.6M params): 83.3% correction, but MMLU 62.5%→32%, car-wash 75%→0%
The part I'm most interested in discussing: we tried five ways to break past ~53% (rank, 5k-step training, multi-depth, deep injection at two layers) and none of them worked, we published that as the finding rather than burying it. DPO at depth 7 even destroyed capabilities outright (MMLU 13%). If anyone has ideas for cracking the frozen-base ceiling without touching weights, I'd love to hear them.
Everything is reproducible: code + weights + eval scripts at https://github.com/eulogik/prajna and eulogik/Prajna-CRNv2. All trained on a Mac Mini M4, no GPU.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Stuck on"A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model (2026)
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation (2026)
- You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model (2026)
- Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following (2026)
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See (2026)
- Self-Specialized Teachers for Domain Post-Training (2026)
- Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.16145 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 3
eulogik/prajna
Datasets citing this paper 1
eulogik/prajna-cehri
Spaces citing this paper 0
No Space linking this paper



