Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
When a hint turns a failing generated program into a passing one, does it provide missing information or merely steer the model toward a solution it could already produce? We test these hypotheses on HumanEval+ and MBPP+ using executable evaluation. For Qwen2.5-3B-Instruct, adaptive relevant hints rescue 36 of 79 selected failures; an unrelated hint rescues 19, while eight unhinted samples solve 46 and recover 31 of the 36 relevant-hint rescues. Phi-3.5-mini shows the same pattern: relevant hints rescue 42 of 101 failures, an unrelated hint rescues 17, and unhinted sampling solves 57, includin
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Liquid AI - Antidoom (the doom loop remover) →
- PossiblePossibly related (embedding) · 46%Competence Gate: gating tool-use on a small model's internal confidence signal instead of its verbalised one — Qwen3.5-4B, open weights [P] →
- FuzzySimilar title/name (fuzzy) · 59%KKKKhazix/khazix-skills →
“Fuzzy title match (0.73): “Hints Help But Do They Teach? Evaluating Skills Transfer in ” ≈ “KKKKhazix/khazix-skills””
- LinkedLinked via arxiv author · 85%Will Badr →
“Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation”
