Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry before compliance degrades, and how much context a model can hold before recall and honesty degrade. We report two controlled experiments crossing all three factors on one held, contamination-free synthetic corpus (the "Book of Veyra," 8,780 uniquely-named entities, deterministically regenerable from a fixed seed), evaluated across five models. Experiment
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Clinical drug report generation using multi-phase prompt large language models - Nature →
- PossiblePossibly related (embedding) · 49%A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P] →
- FuzzySimilar title/name (fuzzy) · 59%NirDiamant/Prompt_Engineering →
“Fuzzy title match (0.73): “Prompt Design at Scale: How Format, Instruction Count, and C” ≈ “NirDiamant/Prompt_Engineering””
- FuzzySimilar title/name (fuzzy) · 59%linshenkx/prompt-optimizer →
“Fuzzy title match (0.73): “Prompt Design at Scale: How Format, Instruction Count, and C” ≈ “linshenkx/prompt-optimizer””
- LinkedLinked via arxiv author · 85%Netanel Eliav →
“Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination ”
