paper-with-me

Image Generation 벤치마크

Image Generation on WISE

15개 결과 · ⬇ CSV · JSON

Overall

0.23 0.35 0.47 0.59 0.71 2023-07 2026-09 stable-diffusion-xl-base-0.9 — 0.43 (2023-07-04) PixArt-XL-2-1024-MS — 0.47 (2023-09-30) Playground-v2.5-1024px-aesthetic — 0.49 (2024-02-27) stable-diffusion-3.5-large — 0.46 (2024-03-05) Show-o — 0.35 (2024-08-22) Emu3-gen — 0.39 (2024-09-27) Janus-pro — 0.35 (2025-01-29) Janus — 0.23 (2025-01-29) MindOmni (w/ cot) — 0.71 (2025-05-19) MindOmni (w/o cot) — 0.43 (2025-05-19) Bagel (w/ cot) — 0.7 (2025-05-20) Bagel — 0.52 (2025-05-20) UniWorld-V1 — 0.55 (2025-06-03) stable-diffusion-xl-base-0.9 — 0.43 (2023-07-04) PixArt-XL-2-1024-MS — 0.47 (2023-09-30) Playground-v2.5-1024px-aesthetic — 0.49 (2024-02-27) MindOmni (w/ cot) — 0.71 (2025-05-19)
RankModel OverallCulturalTimeSpaceBiologyPhysicsChemistry PaperCodeYear
1 MindOmni (w/ cot) 0.710.750.700.760.760.720.52 MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO easonxiao-888/mindomni 2025
2 Bagel (w/ cot) 0.700.760.690.750.650.750.58 Emerging Properties in Unified Multimodal Pretraining ByteDance-Seed/Bagel · neverbiasu/ComfyUI-BAGEL 2025
3 MetaQuery-XL 0.550.560.550.620.490.630.41
3 UniWorld-V1 0.550.530.550.730.450.590.41 UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation PKU-YuanGroup/UniWorld-V1 · pku-yuangroup/imgedit 2025
5 Bagel 0.520.440.550.680.440.600.39 Emerging Properties in Unified Multimodal Pretraining ByteDance-Seed/Bagel · neverbiasu/ComfyUI-BAGEL 2025
6 Playground-v2.5-1024px-aesthetic 0.490.490.580.550.430.480.33 Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation 2024
7 PixArt-XL-2-1024-MS 0.470.450.500.480.490.560.34 PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis PixArt-alpha/PixArt-alpha · Karine-Huang/T2I-CompBench · swookey-thinky/image_diffusion 2023
8 stable-diffusion-3.5-large 0.460.440.500.580.440.520.31 Scaling Rectified Flow Transformers for High-Resolution Image Synthesis Karine-Huang/T2I-CompBench · hxixixh/adaflow 2024
9 stable-diffusion-xl-base-0.9 0.430.430.480.470.440.450.27 SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis stability-ai/generative-models · compvis/fm-boosting · yuchen413/text2image_safety · +6 2023
9 MindOmni (w/o cot) 0.430.400.380.620.360.520.32 MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO easonxiao-888/mindomni 2025
11 Emu3-gen 0.390.340.450.480.410.450.27 Emu3: Next-Token Prediction is All You Need baaivision/emu3 · flagopen/flagscale 2024
12 Janus-pro 0.350.300.370.490.360.420.26 Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling deepseek-ai/janus 2025
12 Show-o 0.350.280.400.480.300.460.30 Show-o: One Single Transformer to Unify Multimodal Understanding and Generation showlab/show-o 2024
14 Janus 0.230.160.260.350.280.300.14 Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling deepseek-ai/janus 2025
15 Mural 자동 추출 0.66 Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers 2026
1–15 / 15 페이지당 10 20 50 100