New Reasoning Image Generation Benchmarks End The Beauty Contest
R2I-Bench fires 3,068 reasoning-heavy prompts at text-to-image models, and pretty can't save you anymore. It's a reasoning image generation benchmark built to score what demo galleries never show: whether the model actually thought before it drew.
Here's the direct answer to what'