Learning a Size-Weight Frontier for Synthetic-Augmented Inference
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%From virtual experiments to biomedical insight with synthetic data →
- PossiblePossibly related (embedding) · 49%Inference →
- FuzzySimilar title/name (fuzzy) · 84%amitness/learning →
“Fuzzy title match (0.92): “Learning a Size-Weight Frontier for Synthetic-Augmented Infe” ≈ “amitness/learning””
- FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference →
“Fuzzy title match (0.92): “Learning a Size-Weight Frontier for Synthetic-Augmented Infe” ≈ “xorbitsai/inference””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Chengpiao Huang →
“Learning a Size-Weight Frontier for Synthetic-Augmented Inference”
- LinkedLinked via arxiv author · 85%Kaizheng Wang →
“Learning a Size-Weight Frontier for Synthetic-Augmented Inference”
