paper-with-me

홈 › Papers

Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression

2026-03-08 · Mridankan Mandal arxiv

Accurate estimation of pasture biomass from agricultural imagery is critical for sustainable livestock management, yet existing methods are limited by the small, imbalanced, and sparsely annotated datasets typical of real world monitoring. In this study, adaptation of vision foundation models to agricultural regression is systematically evaluated on the CSIRO Pasture Biomass benchmark, a 357 image dual view dataset with laboratory validated, component wise ground truth for five biomass targets, through 17 configurations spanning four backbones (EfficientNet-B3 to DINOv3-ViT-L), five cross view fusion mechanisms, and a 4x2 metadata factorial. A counterintuitive principle, termed "fusion complexity inversion", is uncovered: on scarce agricultural data, a two layer gated depthwise convolution (R^2 = 0.903) outperforms cross view attention transformers (0.833), bidirectional SSMs (0.819), and full Mamba (0.793, below the no fusion baseline). Backbone pretraining scale is found to monotonically dominate all architectural choices, with the DINOv2 -> DINOv3 upgrade alone yielding +5.0 R^2 points. Training only metadata (species, state, and NDVI) is shown to create a universal ceiling at R^2 ~ 0.829, collapsing an 8.4 point fusion spread to 0.1 points. Actionable guidelines for sparse agricultural benchmarks are established: backbone quality should be prioritized over fusion complexity, local modules preferred over global alternatives, and features unavailable at inference excluded.

📄 PDF Abstract BibTeX arXiv:2603.07819

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Effective Real Image Editing with Accelerated Iterative Diffusion Inversion

2023-09-10 · ICCV 2023 1 · Zhihong Pan, Riccardo Gherardi, Xiufeng Xie, Stephen Huang

Despite all recent progress, it is still challenging to edit and manipulate natural images with modern generative models. When using Generative Adversarial Network (GAN), one major hurdle is in the inversion process mapp…

Computational EfficiencyDenoisingGenerative Adversarial Network

An Iteration-Free Fixed-Point Estimator for Diffusion Inversion

2025-12-09 · Yifei Chen, Kaiyu Song, Yan Pan, Jianxing Yu 외 arxiv

Diffusion inversion aims to recover the initial noise corresponding to a given image such that this noise can reconstruct the original image through the denoising diffusion process. The key component of diffusion inversi…

Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models

2023-09-14 · James Burgess, Kuan-Chieh Wang, Serena Yeung-Levy

Text-to-image diffusion models generate impressive and realistic images, but do they learn to represent the 3D world from only 2D supervision? We demonstrate that yes, certain 3D scene representations are encoded in the …

Image GenerationNovel View SynthesisText to Image GenerationText-to-Image Generation

Novel View Synthesis using DDIM Inversion

2025-08-14 · Sehajdeep Singh, A V Subramanyam, Aditya Gupta, Sahil Gupta arxiv

Synthesizing novel views from a single input image is a challenging task. It requires extrapolating the 3D structure of a scene while inferring details in occluded regions, and maintaining geometric consistency across vi…

Novel View Synthesis

A Low-Complexity Plug-and-Play Deep Learning Model for Massive MIMO Precoding Across Sites

2025-02-12 · Ali Hasanzadeh Karkan, Ahmed Ibrahim, Jean-François Frigon, François Leduc-Primeau

Massive multiple-input multiple-output (mMIMO) technology has transformed wireless communication by enhancing spectral efficiency and network capacity. This paper proposes a novel deep learning-based mMIMO precoder to ta…

Computational EfficiencyDomain GeneralizationMeta-Learning