Papers Form
“Form” 태그가 달린 논문 1,618편 · 필터 해제
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-aligned audio-text pairs, existing T2A metho…
Audio GenerationData AugmentationFormControlled Retrieval-augmented Context Evaluation for Long-form RAG
Retrieval-augmented generation (RAG) enhances large language models by incorporating context retrieved from external knowledge sources. While the effectiveness of the retrieval module is typically evaluated with relevanc…
DiagnosticFormRAGRetrieval+1FormGym: Doing Paperwork with Agents
Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without access to OCR, typeset PDF text, or a DOM. For computer agents, it requires multip…
FormInformation RetrievalOptical Character Recognition (OCR)FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
Semantic querying in complex 3D scenes through free-form language presents a significant challenge. Existing 3D scene understanding methods use large-scale training data and CLIP to align text queries with 3D semantic fe…
FormGraph GenerationRelational ReasoningScene UnderstandingDirect Reasoning Optimization: LLMs Can Reward And Refine Their Own Reasoning for Open-Ended Tasks
Recent advances in Large Language Models (LLMs) have showcased impressive reasoning abilities in structured tasks like mathematics and programming, largely driven by Reinforcement Learning with Verifiable Rewards (RLVR),…
FormMathARGUS: Hallucination and Omission Evaluation in Video-LLMs
Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinat…
DescriptiveFormHallucinationMultiple-choice+2LLM Unlearning Should Be Form-Independent
Large Language Model (LLM) unlearning aims to erase or suppress undesirable knowledge within the model, offering promise for controlling harmful or private information to prevent misuse. However, recent studies highlight…
FormLarge Language ModelWriting-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Recent advances in Large Language Models (LLMs) have enabled strong performance in long-form writing, yet existing supervised fine-tuning (SFT) approaches suffer from limitations such as data saturation and restricted le…
FormSchedulingToward Better SSIM Loss for Unsupervised Monocular Depth Estimation
Unsupervised monocular depth learning generally relies on the photometric relation among temporally adjacent images. Most of previous works use both mean absolute error (MAE) and structure similarity index measure (SSIM)…
Depth EstimationFormMonocular Depth EstimationSSIM+1Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
Humans have a remarkable ability to acquire and understand grammatical phenomena that are seen rarely, if ever, during childhood. Recent evidence suggests that language models with human-scale pretraining data may posses…
FormSuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
Long-form text generation remains a significant challenge for large language models (LLMs), particularly in maintaining coherence, ensuring logical consistency, and preserving text quality as sequence length increases. T…
FormText GenerationFifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observ…
FormAutomated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
Web applications are critical to modern software ecosystems, yet ensuring their reliability remains challenging due to the complexity and dynamic nature of web interfaces. Recent advances in large language models (LLMs) …
FormScript GenerationBrain-Like Processing Pathways Form in Models With Heterogeneous Experts
The brain is made up of a vast set of heterogeneous regions that dynamically organize into pathways as a function of task demands. Examples of such pathways can be seen in the interactions between cortical and subcortica…
FormMixture-of-ExpertsDo Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
Evaluations of LLMs' ethical risks and value inclinations often rely on short-form surveys and psychometric tests, yet real-world use involves long-form, open-ended responses -- leaving value-related risks and preference…
FormSpecificitySelf-supervised Latent Space Optimization with Nebula Variational Coding
Deep learning approaches process data in a layer-by-layer way with intermediate (or latent) features. We aim at designing a general solution to optimize the latent manifolds to improve the performance on classification, …
FormMetric LearningVariational InferenceExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
This paper introduces ExpertLongBench, an expert-level benchmark containing 11 tasks from 9 domains that reflect realistic expert workflows and applications. Beyond question answering, the application-driven tasks in Exp…
BenchmarkingFormQuestion AnsweringFormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one click", existing tools remain largely ru…
BenchmarkingFormNexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
Summarizing long-form narratives--such as books, movies, and TV scripts--requires capturing intricate plotlines, character interactions, and thematic coherence, a task that remains challenging for existing LLMs. We intro…
DescriptiveFormLong-Form Narrative SummarizationBeyond Multiple Choice: Evaluating Steering Vectors for Adaptive Free-Form Summarization
Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference time. So far, steering vectors have predominantly been evaluated in multiple-c…
FormLanguage ModelingLanguage ModellingMultiple-choice