paper-with-me

머신러닝 최신 연구, 코드와 함께.

718,557논문
302,733코드 링크
2,088벤치마크 task
15,008데이터셋
8,725방법론

Trending 지금 뜨는 연구

HF 업보트 · GitHub 스타 · 공개 구현 수 기준

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

2026-09-02 · Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu 외 · 구현 3개 · ▲ 534 · ★ 119 hf

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still …

StudentSim: Training LLM-based Student Simulators

2026-09-01 · Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh 외 · 구현 3개 · ▲ 482 · ★ 114 hf

AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learne…

Reinforcement Learning

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

2026-09-08 · NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo 외 · ▲ 387 hf

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-nat…

Instruction Following

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

2026-09-03 · Yuntian Deng, Pengyu Nie, Stuart Shieber · 구현 3개 · ▲ 378 · ★ 114 hf

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present com…

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

2026-09-09 · NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng 외 · 구현 8개 · ▲ 304 · ★ 115 hf

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to pre…

Domain Adaptation

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

2026-09-03 · Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang 외 · ▲ 289 hf

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: …

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

2026-09-01 · Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei 외 · 구현 1개 · ▲ 260 hf

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model we…

SenseNova-U1.5: Towards Native Unified Visual Intelligence

2026-09-10 · Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng 외 · 구현 3개 · ▲ 250 · ★ 115 hf

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface throu…

Reinforcement LearningInstruction FollowingImage Editing

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

2026-09-03 · Chuyan Chen, Haoxing Chen, Kun Chen, Zhenglin Cheng 외 · 구현 5개 · ▲ 232 · ★ 8 hf

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbo…

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

2026-09-08 · Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang 외 · ▲ 218 hf

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we const…

Reinforcement Learning

Latest 최신 논문

Omni-Streaming Thinking

2026-09-14 · Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li 외 hf

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If…

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

2026-09-14 · Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He 외 hf

Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, h…

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

2026-09-14 · Yolo Y. Tang, Daiki Shimada, Jiayue Meng, Jing Bi 외 hf

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an age…

Question Answering

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

2026-09-14 · Ling Yang, Zhenfei Yin, Yingcheng Wu hf

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: fro…

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

2026-09-14 · Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang 외 hf

Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals…

Reinforcement Learning

StepAudio 3 Gen Technical Report

2026-09-11 · Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang 외 hf

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types…

Audio Generation

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

2026-09-11 · Zhiwei Li, Lei Zhu, Hao Gu, Xiang Hu 외 hf

Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods us…

Long-Context Understanding

SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image

2026-09-11 · Yu-Rou Tuan, Hao-Tang Tsui, Nicolas Ugrinovic, Kris Kitani 외 hf

Part-aware 3D asset generation enables applications such as editing, articulation, simulation, and fabrication, yet existing methods can generate visually complete individual parts without ensuring that they form a valid…

3D Generation

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

2026-09-11 · Jianman Lin, Shailesh Shailesh, Zhongyi Luo, Jiafei Duan hf

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irr…

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

2026-09-10 · Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang 외 hf

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. W…

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

2026-09-10 · Logesh Kumar Umapathi hf

We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the <=2B parameter division with 0.8279 on the held-out test set. Our system is a single 2B vision-language mo…

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

2026-09-10 · Logesh Kumar Umapathi hf

We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide…

Visual Grounding

Feature Recovery for Object Understanding After Irreversible Fire Damage

2026-09-10 · Aditi Tiwari, Sofia Stoica, Savya Khosla, David Forsyth 외 hf

Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, material state, and visual appearance. Detecting and identifying these remnants is critical for locating h…

Object Detection

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

2026-09-10 · Aditi Tiwari, Aashrith Bandaru, Heng Ji hf

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be…

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

2026-09-10 · Shilong Zou, Shilin Zhang, Yingji Zhang, Yuhang Huang 외 hf

In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and d…

Decision Making

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

2026-09-10 · Changbo Yan, Zhongbo Zhang, Zaibin Zhang, Yifan Wang 외 hf

3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, stand…

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

2026-09-10 · Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang 외 hf

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video g…

Instruction FollowingVideo Generation

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

2026-09-10 · Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao 외 hf

Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. …

World in World: Explore the World with World Models

2026-09-10 · Chenxi Song, Yanming Yang, Chi Zhang hf

Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised…