paper-with-me

Papers

Rethinking Implicit Neural Representations for Vision Learners

2022-11-22 · Yiran Song, Qianyu Zhou, Lizhuang Ma

Implicit Neural Representations (INRs) are powerful to parameterize continuous signals in computer vision. However, almost all INRs methods are limited to low-level tasks, e.g., image/video compression, super-resolution, and image generation. The questions on how to explore INRs to high-level tasks and deep networks are still under-explored. Existing INRs methods suffer from two problems: 1) narrow theoretical definitions of INRs are inapplicable to high-level tasks; 2) lack of representation capabilities to deep networks. Motivated by the above facts, we reformulate the definitions of INRs from a novel perspective and propose an innovative Implicit Neural Representation Network (INRN), which is the first study of INRs to tackle both low-level and high-level tasks. Specifically, we present three key designs for basic blocks in INRN along with two different stacking ways and corresponding loss functions. Extensive experiments with analysis on both low-level tasks (image fitting) and high-level vision tasks (image classification, object detection, instance segmentation) demonstrate the effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:2211.12040

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage ClassificationImage GenerationInstance Segmentationobject-detectionObject DetectionSemantic SegmentationSuper-ResolutionVideo Compression

Similar Papers 제목 키워드 기반

Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

2024-03-03 · Kun-Yu Lin, Henghui Ding, Jiaming Zhou, Yu-Ming Tang 외

Building upon the impressive success of CLIP (Contrastive Language-Image Pretraining), recent pioneer works have proposed to adapt the powerful CLIP to video data, leading to efficient and effective video learners for op…

Action RecognitionOpen Vocabulary Action Recognition

Interactive Disentanglement: Learning Concepts by Interacting with their Prototype Representations

2021-12-04 · CVPR 2022 1 · Wolfgang Stammer, Marius Memmel, Patrick Schramowski, Kristian Kersting

Learning visual concepts from raw images without strong supervision is a challenging task. In this work, we show the advantages of prototype representations for understanding and revising the latent space of neural conce…

Disentanglement

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

2025-10-15 · Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang 외 arxiv

Multimodal semantic cues, such as textual descriptions, have shown strong potential in enhancing target perception for tracking. However, existing methods rely on static textual descriptions from large language models, w…

Multi-Object Tracking

Rethinking Annotation: Can Language Learners Contribute?

2022-10-13 · Haneul Yoo, Rifki Afina Putri, Changyoon Lee, Youngin Lee 외

Researchers have traditionally recruited native speakers to provide annotations for widely used benchmark datasets. However, there are languages for which recruiting native speakers can be difficult, and it would help to…

Machine Reading Comprehensionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3

Rethinking Generative Human Video Coding with Implicit Motion Transformation

2025-06-12 · Bolin Chen, Ru-Ling Liao, Jie Chen, Yan Ye

Beyond traditional hybrid-based video codec, generative video codec could achieve promising compression performance by evolving high-dimensional signals into compact feature representations for bitstream compactness at t…

DecoderVideo Compression