ONOTE: Hypergraph-Grounded Omnimodal Reasoning for Computational Music Science
Omnimodal notation processing, centered on sheet music, is a controlled scientific setting in which auditory, visual, symbolic, and physical representations must encode the same musical events. Yet existing work remains fragmented across recognition and transcription, rarely testing structural consistency across notation systems. Western-staff bias and underspecified model judges further conceal errors in pitch, timing, ordering, and instrument-specific constraints. We introduce ONOTE, a unified framework that treats music as a scientifically structured domain of measurable cross-representation correspondences. Its test-only benchmark draws on a diverse collection of musical sources covering staff, Jianpu, and tablature across varied genres, instruments, and structural conditions, with aligned multimodal derivatives. Four complementary tasks cover score understanding, notation conversion, audio transcription, and symbolic generation, testing pitch and duration ordering, output syntax, and disclosed instrument-specific constraints. ONOTE also constructs a provenance-bearing proposition hypergraph from external music-theory materials for entity- and hyperedge-based evidence retrieval. Deterministic validity checks, disclosed structural-compliance SMG scoring, and controlled RAG comparisons reveal gaps between visual recognition and structure-preserving outputs. Results separate perception from music-theory application and structural or physical constraint satisfaction. ONOTE provides an auditable framework for studying representation invariance and knowledge-grounded intervention in computational music science.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Hypergraph Transformer: Weakly-supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering
Knowledge-based visual question answering (QA) aims to answer a question which requires visually-grounded external knowledge beyond image content itself. Answering complex questions that require multi-hop reasoning under…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
Unified multimodal models (UMMs) have emerged as a powerful paradigm for seamlessly unifying text and image understanding and generation. However, prevailing evaluations treat these abilities in isolation, such that task…
Question AnsweringUni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating…
Computational EfficiencyImage GenerationTowards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To exte…
Multimodal ReasoningBetter with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social platforms. However, they still reset at every post, leaving useful co…