paper-with-me

Weekly Digest 2026년 38주차

2026-09-14 ~ 2026-09-20 신규 논문 5편을 분야별로 묶었습니다. 정렬: 인기 신호(HF 업보트·GitHub 스타·구현 수).

← 이전 주 2026-W38

General 3편

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

2026-09-14 · Ling Yang, Zhenfei Yin, Yingcheng Wu · ▲ 12 hf

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: fro…

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

2026-09-14 · Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He 외 · ▲ 11 hf

Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, h…

Omni-Streaming Thinking

2026-09-14 · Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li 외 · ▲ 8 hf

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If…

Natural Language Processing 1편

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

2026-09-14 · Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi 외 · 구현 1개 · ▲ 3 · ★ 5 hf

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an age…

Question Answering

Methodology 1편

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

2026-09-14 · Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang 외 · ▲ 3 hf

Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals…

Reinforcement Learning