paper-with-me

Papers

Aletheia tackles FirstProof autonomously

2026-02-24 · Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong arxiv

We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed timeframe of the challenge, Aletheia autonomously solved 6 problems (2, 5, 7, 8, 9, 10) out of 10 according to majority expert assessments; we note that experts were not unanimous on Problem 8 (only). For full transparency, we explain our interpretation of FirstProof and disclose details about our experiments as well as our evaluation. Raw prompts and outputs are available at https://github.com/google-deepmind/superhuman/tree/main/aletheia.

📄 PDF Abstract BibTeX arXiv:2602.21201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

2026-07-10 · Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely Bérczi 외 arxiv

Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practic…

Verify as You Go: An LLM-Powered Browser Extension for Fake News Detection

2026-02-03 · Dorsaf Sallami, Esma Aïmeur arxiv

The rampant spread of fake news in the digital age poses serious risks to public trust and democratic institutions, underscoring the need for effective, transparent, and user-centered detection tools. Existing browser ex…

Fake News Detection

Towards Autonomous Mathematics Research

2026-02-10 · Tony Feng, Trieu H. Trinh, Garrett Bingham, Dawsen Hwang 외 arxiv

Recent advances in foundational models have yielded reasoning systems capable of achieving a gold-medal standard at the International Mathematical Olympiad. The transition from competition-level problem-solving to profes…

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures

2026-04-04 · Abdulmalek Saket arxiv

Low-Rank Adaptation (LoRA) has become the dominant parameter-efficient fine-tuning method for large language models, yet standard practice applies LoRA adapters uniformly to all transformer layers regardless of their rel…

parameter-efficient fine-tuning

Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings

2026-07-14 · Joseph Walusimbi, Ann Move Oguti, Abubakhari Sserwadda, Precious Boss Kasasira 외 arxiv

Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where physician-to-patient ratios can fall below 1:25,000 in rural settings. Existing AI-assisted diagnostic tools predominantly…