paper-with-me

Papers

Structural Certification for Reliable Physical Design with Language Models

2026-06-29 · Nakul Vyas, Iliya D. Stoev arxiv

An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the model proposes, and a deterministic engine alone certifies, returning certified, impossible, or unknown. We introduce Physics-Anchored Certification (PHACT), a propose-certify loop spanning five scientific domains, and identify what makes such a certificate trustworthy. A checker that accepts a model-supplied value can be forged; deriving the certified quantity from fixed inputs instead makes forgery impossible by construction. Across eighty adversarial trials spanning two models, two decoding temperatures, and a deliberately faulted engine, this contract produced zero false certifications.

📄 PDF Abstract BibTeX arXiv:2606.30107

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GeoCert: Certified Geometric AI for Reliable Forecasting

2026-04-25 · Regina Zhang, Zongru Li, Honggang Wen, Xiaofeng Liu 외 arxiv

Forecasting systems in science must be accurate, physically consistent, and certifiably reliable. Most existing models address prediction, constraint enforcement, and verification separately, limiting scalability and int…

World Models in Pieces: Structural Certification for General Agents

2026-06-23 · Yikai Lu, Yifei Wu, Xinyu Lu, Tongxin Li arxiv

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish between the understa…

Fast Certification of Vision-Language Models Using Incremental Randomized Smoothing

2023-11-15 · A K Nirala, A Joshi, C Hegde, S Sarkar

A key benefit of deep vision-language models such as CLIP is that they enable zero-shot open vocabulary classification; the user has the ability to define novel class labels via natural language prompts at inference time…

Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design

2026-05-19 · Elias Berger, Muhammad Usama, Jan Mehlstäubl, Bernhard Saske 외 arxiv

Large Language Models (LLMs) can generate Computer-Aided Design (CAD), yet lack physical comprehension required for reliable engineering design. Instead of attempting to implicitly learn physical laws from data, we propo…

Decision Making

Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design

2025-09-16 · Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma arxiv

Large language models often produce confident but incorrect outputs, creating a critical need for reliable uncertainty quantification with formal abstention guarantees. We introduce information-lift certificates that com…