paper-with-me

Papers

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

2026-08-20 · Josef Chen, Erim Hayretci arxiv

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.

📄 PDF Abstract BibTeX arXiv:2608.20574

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task

2024-11-04 · Hoonick Lee, Mogan Gim, Donghyeon Park, Donghee Choi 외

The advent of Large Language Models (LLMs) have shown promise in various creative domains, including culinary arts. However, many LLMs still struggle to deliver the desired level of culinary creativity, especially when t…

Recipe Generation

Control Under Compression: Reliability Frontiers for Tool-Using Agents

2026-08-02 · Yinghan Hou, Zongyou Yang arxiv

Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent c…

A generative grammar of cooking

2022-10-12 · Ganesh Bagler

Cooking is a uniquely human endeavor for transforming raw ingredients into delicious dishes. Over centuries, cultures worldwide have evolved diverse cooking practices ingrained in their culinary traditions. Recipes, thus…

NutritionRecipe Generation

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

2026-06-24 · Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios 외 arxiv

Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. …

CookingSense: A Culinary Knowledgebase with Multidisciplinary Assertions

2024-05-01 · Donghee Choi, Mogan Gim, Donghyeon Park, Mujeen Sung 외

This paper introduces CookingSense, a descriptive collection of knowledge assertions in the culinary domain extracted from various sources, including web data, scientific papers, and recipes, from which knowledge coverin…

DescriptiveLanguage ModelingLanguage ModellingRetrieval