paper-with-me

홈 › Papers

TCFormer: Visual Recognition via Token Clustering Transformer

2024-07-16 · Wang Zeng, Sheng Jin, Lumin Xu, Wentao Liu, Chen Qian, Wanli Ouyang, Ping Luo, Xiaogang Wang

Transformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid region with a vision token. However, fixed token distribution disregards the semantic meaning of different image regions, resulting in sub-optimal performance. To address this issue, we propose the Token Clustering Transformer (TCFormer), which generates dynamic vision tokens based on semantic meaning. Our dynamic tokens possess two crucial characteristics: (1) Representing image regions with similar semantic meanings using the same vision token, even if those regions are not adjacent, and (2) concentrating on regions with valuable details and represent them using fine tokens. Through extensive experimentation across various applications, including image classification, human pose estimation, semantic segmentation, and object detection, we demonstrate the effectiveness of our TCFormer. The code and models for this work are available at https://github.com/zengwang430521/TCFormer.

📄 PDF Abstract BibTeX arXiv:2407.11321

Code (1)

zengwang430521/tcformer 공식 구현 pytorch

Tasks

Clusteringimage-classificationImage Classificationobject-detectionObject DetectionPose EstimationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Residual Connection 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer

2022-04-19 · CVPR 2022 1 · Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian 외

Vision transformers have achieved great successes in many computer vision tasks. Most methods generate vision tokens by splitting an image into a regular and fixed grid and treating each cell as a token. However, not all…

2D Human Pose Estimation3D Human Pose EstimationAllClustering+1

FTCFormer: Fuzzy Token Clustering Transformer for Image Classification

2025-07-14 · Muyi Bao, Changyu Zeng, Yifan Wang, Zhengni Yang 외 arxiv

Transformer-based deep neural networks have achieved remarkable success across various computer vision tasks, largely attributed to their long-range self-attention mechanism and scalability. However, most transformer arc…

Image Classification

TCFormer: A 5M-Parameter Transformer with Density-Guided Aggregation for Weakly-Supervised Crowd Counting

2025-12-21 · Qiang Guo, Rubo Zhang, Bingbing Zhang, Junjie Liu 외 arxiv

Crowd counting typically relies on labor-intensive point-level annotations and computationally intensive backbones, restricting its scalability and deployment in resource-constrained environments. To address these challe…

Crowd Counting

PointCFormer: a Relation-based Progressive Feature Extraction Network for Point Cloud Completion

2024-12-11 · Yi Zhong, Weize Quan, Dong-Ming Yan, Jie Jiang 외

Point cloud completion aims to reconstruct the complete 3D shape from incomplete point clouds, and it is crucial for tasks such as 3D object detection and segmentation. Despite the continuous advances in point cloud anal…

3D Object DetectionPoint Cloud CompletionRelation

3D Human Pose Estimation With Spatio-Temporal Criss-Cross Attention

2023-01-01 · CVPR 2023 1 · Zhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong 외

Recent transformer-based solutions have shown great success in 3D human pose estimation. Nevertheless, to calculate the joint-to-joint affinity matrix, the computational cost has a quadratic growth with the increasin…

3D Human Pose EstimationPose Estimation