An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and withou
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “An Empirical Study of Reward Specification and Benchmark Rel” ≈ “jeinlee1991/chinese-llm-benchmark””
- LinkedLinked via arxiv author · 85%Rubén Balbastre →
“An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning”
- LinkedLinked via arxiv author · 85%Juan Manuel Orduña →
“An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning”
- LinkedLinked via arxiv author · 85%Mariano Pérez →
“An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning”
- PossiblePossibly related (embedding) · 49%Same GRPO recipe on three from-scratch LLMs (353M/316M/672M) gave three different outcomes, with no clean relationship to scale [P] →
