

- Published on 9 Feb 2026
- Last updated on 9 Feb 2026
- Reading Time: 6 minutes
Image Arena Improvements: New Categories & Quality Filtering
After analyzing over 4 million user prompts, it is clear that a single global leaderboard no longer captures the full picture. Today we introduce new categories and quality filtering.
Text-to-image models have advanced rapidly, leading to a surge in the diversity of their applications. After analyzing over 4 million user prompts—ranging from fantasy art to practical uses like logo and poster design—it is clear that a single global leaderboard no longer captures the full picture. Instead, category-specific rankings provide a clearer understanding of how different models perform across specific domains.
Furthermore, not all prompts are created equal when it comes to evaluation. Underspecified or malformed prompts introduce noise that can distort model rankings. To address these challenges, we are introducing two complementary updates to the Text-to-Image Arena:
Furthermore, not all prompts are created equal when it comes to evaluation. Underspecified or malformed prompts introduce noise that can distort model rankings. To address these challenges, we are introducing two complementary updates to the Text-to-Image Arena:
- Prompt Categories: Enabling category-specific leaderboards for more granular performance tracking.
- Quality Filtering: Reducing noise to ensure higher ranking quality and statistical reliability.
Together, these enhancements provide a more interpretable and dependable framework for evaluating the state of the art in text-to-image generation.
Categories in Text-to-Image Arena
Categories are the primary way we organize prompts on Arena. Each text-to-image prompt can be tagged with one or more categories, though tagging remains optional. This flexibility ensures that we capture the richness of real-world use cases without forcing every conversation into a rigid box.
When you view a category leaderboard, you are seeing the same evaluation methodology as the main Text-to-Image Arena leaderboard, just filtered for that specific prompt domain. This makes categories a powerful tool for comparing model performance on specialized tasks, such as 3D generation, portraits, or commercial logo designs.
After conducting an extensive clustering analysis on a massive collection of image generation prompts, we are introducing seven distinct categories to the Text-to-Image Arena. While these categories capture a wide range of intent, their frequencies vary substantially: Photorealistic & Cinematic Imagery is the most common at roughly 43%, while specialized tasks like 3D Imaging & Modeling account for about 10%. Because these categories are not mutually exclusive, prompts can often overlap—for instance, a Product, Branding & Commercial Design prompt may frequently require high-quality Text Rendering. The figure below illustrates the distribution of these domains.











