EMNLP 2026. One-shot GRPO LoRA adapters: a single BBQ example saturates fairness benchmarks without making models fairer.
AI & ML interests
None defined yet.
Recent Activity
View all activity
models 26
MichiganNLP/hacking-fairness-benchmarks-qwen3-8b-base-z1
Updated • 18
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z999
Updated • 21
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z876
Updated • 16
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z751
Updated • 15
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z501
Updated • 17
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z251
Updated • 20
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z2
Updated • 19
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1000
Updated • 19
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1
Updated • 15
MichiganNLP/hacking-fairness-benchmarks-llama-3.1-8b-z1
Updated • 14
datasets 20
MichiganNLP/language-energy-divide
Viewer • Updated • 122 • 36
MichiganNLP/LUCid
Preview • Updated • 144
MichiganNLP/misfired-alignment-eval-results
Updated • 5
MichiganNLP/misfired-alignment
Viewer • Updated • 4.06k • 6
MichiganNLP/one-shot-grpo-bias-flipped
Viewer • Updated • 72 • 5
MichiganNLP/pact-culture-personalization
Viewer • Updated • 339k • 45
MichiganNLP/TAMA_Instruct
Viewer • Updated • 71.9k • 740 • 1
MichiganNLP/blog-images
Viewer • Updated • 2 • 55
MichiganNLP/Chumor
Viewer • Updated • 3.34k • 40 • 10
MichiganNLP/MUStARD
Viewer • Updated • 1.38k • 398 • 3