newsReddit r/MachineLearningTrust 52 · CommunityPublished 4d agoLive · 2d ago
A dataset with 52 Text to image model evaluation [P]
I created a simple text to image benchmark. I curated 192 prompts that are difficult for T2I models in various ways: text rendering, spatial reasoning, human realism, negations, etc... I then asked a VLM to judge every output against a pre-specified binary question with the ground truth baked in. I'm publishing all the results including the images. (Most public T2I leaderboards don't publish the actual images and that's a sh
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures →
- PossiblePossibly related (embedding) · 51%GenRouter: Unified Workflow Routing for Agentic Image Generation →
- PossiblePossibly related (embedding) · 50%Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation →
- PossiblePossibly related (embedding) · 49%Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA →
- PossiblePossibly related (embedding) · 48%Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation →
- PossiblePossibly related (embedding) · 49%When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images →
Covers
paperHow Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific FigurespaperGenRouter: Unified Workflow Routing for Agentic Image GenerationpaperMulti-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image GenerationpaperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperRead It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image GenerationpaperWhen Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
Related across the graph
paperJust how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQApaperMulti-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image GenerationpaperWhen Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated ImagespaperGenRouter: Unified Workflow Routing for Agentic Image GenerationpaperRead It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image GenerationpaperHow Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
