paper-with-me

홈 › Papers

SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model

2026-02-21 · Luca Cazzaniga arxiv

This paper presents SCHEMA (Structured Components for Harmonized Engineered Modular Architecture), a structured prompt engineering methodology specifically developed for Google Gemini 3 Pro Image. Unlike generic prompt guidelines or model-agnostic tips, SCHEMA is an engineered framework built on systematic professional practice encompassing 850 verified API predictions within an estimated corpus of approximately 4,800 generated images, spanning six professional domains: real estate photography, commercial product photography, editorial content, storyboards, commercial campaigns, and information design. The methodology introduces a three-tier progressive system (BASE, MEDIO, AVANZATO) that scales practitioner control from exploratory (approximately 5%) to directive (approximately 95%), a modular label architecture with 7 core and 5 optional structured components, a decision tree with explicit routing rules to alternative tools, and systematically documented model limitations with corresponding workarounds. Key findings include an observed 91% Mandatory compliance rate and 94% Prohibitions compliance rate across 621 structured prompts, a comparative batch consistency test demonstrating substantially higher inter-generation coherence for structured prompts, independent practitioner validation (n=40), and a dedicated Information Design validation demonstrating >95% first-generation compliance for spatial and typographical control across approximately 300 publicly verifiable infographics. Previously published on Zenodo (doi:10.5281/zenodo.18721380).

📄 PDF Abstract BibTeX arXiv:2602.18903

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt EngineeringImage Generation

Similar Papers 제목 키워드 기반

ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction

2026-02-12 · Nick Ferguson, Josh Pennington, Narek Beghian, Aravind Mohan 외 arxiv

Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed to automate this extraction, making acc…

Evidence-Guided Schema Normalization for Temporal Tabular Reasoning

2025-11-29 · Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta, Vivek Gupta arxiv

Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose a SQL-based approach that involves (1) generating a 3NF schema from Wikipedia infoboxes, (2) generating SQL quer…

Runtime Burden Allocation for Structured LLM Routing in Agentic Expert Systems: A Full-Factorial Cross-Backend Methodology

2026-03-26 · Zhou Hanlin, Chan Huah Yong arxiv

Structured LLM routing is often treated as a prompt-engineering problem. We argue that it is, more fundamentally, a systems-level burden-allocation problem. As large language models (LLMs) become core control components …

Generating Structured Outputs from Language Models: Benchmark and Studies

2025-01-18 · Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins 외

Reliably generating structured outputs has become a critical capability for modern language model (LM) applications. Constrained decoding has emerged as the dominant technology across sectors for enforcing structured out…

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

2026-04-25 · Duaa Alim, Mogtaba Alim, Liam Chalcroft arxiv

Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times and consult the literature before writing a report. We introduce GAZE (…

Edge Detection