paper-with-me

홈 › Papers

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

2026-08-17 · Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen, Yang Zhang arxiv

Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically optimized pages visibility disproportionate to their authority or relevance and even make weak or false information appear well supported. Unlike conventional search, generative search synthesizes information into direct answers rather than presenting competing sources, which can further amplify these risks, as assessing source provenance and authority requires additional user interaction. Despite these concerns, systematic methods for detecting GEO-optimized webpages remain underexplored. We introduce \texttt{GEOFlagBench}, a benchmark of 3,200 webpages spanning 400 queries, four domains, and eight GEO optimizer families, and use it to systematically evaluate existing GEO detection methods. Although the strongest baseline achieves an aggregate F1 of 0.880, method-level and authorship-conditioned evaluations reveal substantial weaknesses and potential reliance on authorship-related shortcuts. We therefore propose \emph{Intervention-Paired Training} (IPT), which supervises detector responses to GEO interventions and non-GEO AI polishing; on ModernBERT, IPT improves F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883. We develop a GEO-gated Agent system for auditing the Source Tier and verifiability of Citation URLs in detected GEO pages. Finally, we deploy the complete pipeline on released Google Search and Gemini-grounded retrieval results for 1,000 real-user queries. Across 10,095 available pages, we estimate an overall GEO prevalence of 8.90\%, reaching 16.36\% among pages modified in 2026. Our results establish a foundation for systematically detecting, auditing, and measuring GEO in real-world search ecosystems.

📄 PDF Abstract BibTeX arXiv:2608.16824

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI

2025-11-29 · Yanhui Li, Qi Zhou, Zhihong Xu, Huizhong Guo 외 arxiv

Large vision-language models (LVLMs) are increasingly used for tasks where detecting multimodal harmful content is crucial, such as online content moderation. However, real-world harmful content is often camouflaged, rel…

Scene UnderstandingVisual Reasoning

Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification

2023-02-17 · Luke Bates, Iryna Gurevych

Few-shot text classification systems have impressive capabilities but are infeasible to deploy and use reliably due to their dependence on prompting and billion-parameter language models. SetFit (Tunstall et al., 2022) i…

ClassificationContrastive LearningFew-Shot Text ClassificationSentence+2

Countering Malicious Content Moderation Evasion in Online Social Networks: Simulation and Detection of Word Camouflage

2022-12-27 · Álvaro Huertas-García, Alejandro Martín, Javier Huertas Tato, David Camacho

Content moderation is the process of screening and monitoring user-generated content online. It plays a crucial role in stopping content resulting from unacceptable behaviors such as hate speech, harassment, violence aga…

Multilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok

2025-07-30 · Minh Duc Chu, Kshitij Pawar, Zihao He, Roxanna Sharifi 외 arxiv

Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness…

Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms

2022-11-11 · Vashist Avadhanula, Omar Abdul Baki, Hamsa Bastani, Osbert Bastani 외

We describe the current content moderation strategy employed by Meta to remove policy-violating content from its platforms. Meta relies on both handcrafted and learned risk models to flag potentially violating content fo…