EMNLP 2026. One-shot GRPO LoRA adapters: a single BBQ example saturates fairness benchmarks without making models fairer.
AI & ML interests
None defined yet.
Recent Activity
View all activity
models 26
MichiganNLP/hacking-fairness-benchmarks-qwen3-8b-base-z1
Updated • 20
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z999
Updated • 23
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z876
Updated • 17
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z751
Updated • 16
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z501
Updated • 18
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z251
Updated • 25
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z2
Updated • 20
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1000
Updated • 20
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1
Updated • 16
MichiganNLP/hacking-fairness-benchmarks-llama-3.1-8b-z1
Updated • 15
datasets 20
MichiganNLP/language-energy-divide
Viewer • Updated • 122 • 36
MichiganNLP/LUCid
Preview • Updated • 137
MichiganNLP/misfired-alignment-eval-results
Updated • 3
MichiganNLP/misfired-alignment
Viewer • Updated • 4.06k • 5
MichiganNLP/one-shot-grpo-bias-flipped
Viewer • Updated • 72 • 4
MichiganNLP/pact-culture-personalization
Viewer • Updated • 339k • 50
MichiganNLP/TAMA_Instruct
Viewer • Updated • 71.9k • 678 • 1
MichiganNLP/blog-images
Viewer • Updated • 2 • 53
MichiganNLP/Chumor
Viewer • Updated • 3.34k • 40 • 10
MichiganNLP/MUStARD
Viewer • Updated • 1.38k • 388 • 3