paper-with-me

Papers

WebPII: Benchmarking Visual PII Detection for Computer-Use Agents

2026-03-18 · Nathan Zhao arxiv

Computer use agents create new privacy risks: training data collected from real websites inevitably contains sensitive information, and cloud-hosted inference exposes user screenshots. Detecting personally identifiable information in web screenshots is critical for privacy-preserving deployment, but no public benchmark exists for this task. We introduce WebPII, a fine-grained synthetic benchmark of 44,865 annotated e-commerce UI images designed with three key properties: extended PII taxonomy including transaction-level identifiers that enable reidentification, anticipatory detection for partially-filled forms where users are actively entering data, and scalable generation through VLM-based UI reproduction. Experiments validate that these design choices improve layout-invariant detection across diverse interfaces and generalization to held-out page types. We train WebRedact to demonstrate practical utility, more than doubling text-extraction baseline accuracy (0.753 vs 0.357 mAP@50) at real-time CPU latency (20ms). We release the dataset and model to support privacy-preserving computer use research.

📄 PDF Abstract BibTeX arXiv:2603.17357

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents

2024-08-06 · Qiang Sun, Yuanyi Luo, Sirui Li, Wenxiao Zhang 외

Multimodal conversational agents are highly desirable because they offer natural and human-like interaction. However, there is a lack of comprehensive end-to-end solutions to support collaborative development and benchma…

BenchmarkingRetrieval-augmented GenerationSpeech-to-Text

Benchmarking Suite for Synthetic Aperture Radar Imagery Anomaly Detection (SARIAD) Algorithms

2025-04-10 · Lucian Chauvina, Somil Guptac, Angelina Ibarrac, Joshua Peeples

Anomaly detection is a key research challenge in computer vision and machine learning with applications in many fields from quality control to radar imaging. In radar imaging, specifically synthetic aperture radar (SAR),…

Anomaly DetectionBenchmarking

macOSWorld: A Multilingual Interactive Benchmark for GUI Agents

2025-06-04 · Pei Yang, Hai Ci, Mike Zheng Shou

Graphical User Interface (GUI) agents show promising capabilities for automating computer-use tasks and facilitating accessibility, but existing interactive benchmarks are mostly English-only, covering web-use or Windows…

BenchmarkingDomain Adaptation

OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents

2025-06-19 · Reyna Abhyankar, Qi Qi, Yiying Zhang

Generative AI is being leveraged to solve a variety of computer-use tasks involving desktop applications. State-of-the-art systems have focused solely on improving accuracy on leading benchmarks. However, these systems a…

Benchmarking

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

2026-06-15 · Anqi Zou, Han Deng, Chengyu Zhang, Junquan Hu 외 arxiv

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven pa…