CALiT · Cross Aggregation and Local Mixing Transformer

Vision Transformers Need Cross Aggregation for Patch Semantics

Ajad Chhatkuli1Pramish Paudel1Deheng Zhang1Kanzhi Wu2Luc Van Gool1Danda Pani Paudel1

1INSAIT, Sofia University "St. Kliment Ohridski"2vivo Mobile Communication Co., Ltd.

INSAIT

Paper (soon)arXiv (soon)Code (soon)BibTeX
Figure 1 of the paper: (a) Cityscapes linear-probe gains of CALiT over ViT for five pre-training setups at S and B scale; (b) gradient-agreement maps on two ImageNet images for ViT-S and CALiT-S; (c) mean gradient agreement inside and outside the object box for ViT-S, CAT-S and CALiT-S
(a) Frozen-backbone linear probe on Cityscapes (mIoU) at S and B scale: circles are ViTs, squares CALiT, one colour per pre-training. (b) Cosine between the mean gradient inside the object box and each patch's gradient at block L−1, for supervised ViT-S and CALiT-S after 200 epochs (green: aligned, indigo: opposing). ViT's gradient agrees with the object nearly everywhere, background included. (c) The same agreement inside and outside the box, mean and standard deviation over all 38.7k ImageNet-val images, for ViT-S, CAT-S (cross aggregation only) and CALiT-S.

Abstract

A Vision Transformer (ViT) is tasked to learn both a useful global representation through the CLS token and distinct local semantics through patch tokens. Yet a loss on the CLS token does not always induce learning of discriminative patch tokens. We theoretically observe that a CLS-based loss induces low rank patch gradients and in practice, most patches receive similar gradients. Adding patch-based losses may not compensate for the strong unidirectional CLS induced gradients. Consequently, even under global-local objectives, patch updates can remain insufficiently differentiated and spatially correlated. We therefore introduce cross-aggregation, which redefines the CLS token by aggregating patch-level concept heads in the expanded FFN space of every block. This parameter-free operation adds only 0.6% FLOPs to ViT-S and more directly couples global semantics to patch tokens. Combined with local depth-wise convolutions, it further improves the patch semantics, in a variety of pre-training methods. Together, these components form the Cross Aggregation and Local Mixing Transformer (CALiT). CALiT consistently improves semantic probes under both global-only and global-local pre-training objectives and vision-language finetuning. In particular, CALiT improves thin class semantics mean IoU by double digits in many pre-training/linear probe pairs and provides moderate gains in all other local semantics.

Qualitative results

Two Cityscapes val images with a zoomed region: patch-feature PCA and segmentation for DINO ViT-S/16 against CALiT-S (DINO), and iBOT ViT-S/16 against CALiT-S (iBOT)
Thin structures, DINO and iBOT at S/16. Each row shows a Cityscapes val image, a zoomed region (yellow box), the region's patch-feature PCA (top three components of the frozen last-layer features, as RGB) and its segmentation (ground truth and each model's linear-probe prediction). The baselines miss the bicycle and the person; CALiT recovers them.

Patch features on random Cityscapes frames

Six val frames drawn at random with a fixed seed, not selected. Pick a pre-training and a frame.

Cityscapes val frame munster_000106_000019, 1024 by 1024 crop
input
Patch-feature PCA of the baseline ViT
DeiT-S/16
Patch-feature PCA of CALiT
CALiT-S

Top three principal components of the frozen last-layer patch features, as RGB, on a 1024×1024 crop (64×64 patches). PCA is fitted per model and per image, so colours carry no meaning across panels: compare the structure, such as poles, pedestrians and lane markings.

Method

CALiT changes only the FFN. After fc1 and GELU, each token's 4D-wide hidden is read as S slots of concepts, and two light modules act on it in parallel before fc2.

C′ = MHAH(C, P, P), heada = softmax  (CaP⊤a√dh) Pa
C ∈ ℝS×4D/S: the S slots of the CLS hidden. P ∈ ℝNS×4D/S: the S slots of each of the N patch hiddens. dh = 4D/(SH).
CALiT block, pseudocode
# x: (1+N, D) tokens, CLS first, patches on a √N×√N grid.  S slots, H heads (as in SA).
y       = x + SA(LN(x))
h       = GELU(W1(LN(y)))              # (1+N, 4D)
h_c     = h[:1];  h_p = h[1:]          # split the CLS row from the N patch rows

# cross aggregation: parameter-free, MHA without projection weights
C       = h_c.reshape(S, 4D/S)         # CLS slots: queries
P       = h_p.reshape(N*S, 4D/S)       # patch slots: keys and values (un-mixed h_p)
h_c_new = MHA_H(C, P, P).reshape(1, 4D)

# local mixing: depthwise 3×3, Dirac-initialised, patches only
G       = h_p.T.reshape(4D, √N, √N)
h_p_new = DWConv3x3(G).reshape(4D, N).T

z       = y + W2(concat[h_c_new, h_p_new])
return z

Patch gradients, measured

We take the gradient of each model's own CLS loss with respect to the 196 patch tokens entering the last block, on all 38,651 ImageNet-val images whose object box leaves at least one patch outside. Two questions: how many directions does that gradient span, and does it single out the object?

pre-trainingCLS loss2nheffective rank ↑contrast ↑CALiT > ViT
ViTCALiTViTCALiT
supervised IN-1k, ViT-S/16cross-entropy1212.137.00.220.5691 %
iBOT IN-1k, ViT-S/16CLS self-distillation1213.044.80.130.2783 %
iBOT IN-1k, ViT-B/16CLS self-distillation2420.353.10.120.2582 %

Structure of the patch gradient at the input of the last block. Effective rank: entropy-effective rank of the 196×D gradient matrix; a plain ViT is bounded by 2nh (two routes per head). Contrast: agreement with the mean in-box gradient direction, inside minus outside the box, each patch weighted by its gradient norm. Last column: share of images on which CALiT's contrast exceeds ViT's; its effective rank exceeds ViT's on ≥ 99.9 % of images.

Gradient agreement, per model

Gradient agreement maps and inside/outside bar chart for supervised ViT-S, CAT-S and CALiT-S at epoch 100
Supervised ViT-S/16 family, epoch 100
Gradient agreement maps and inside/outside bar chart for supervised ViT-S, CAT-S and CALiT-S at epoch 300
Supervised ViT-S/16 family, epoch 300
Gradient agreement maps and bar chart for iBOT ViT-S and CALiT-S under iBOT's CLS self-distillation loss
iBOT, ViT-S/16, CLS self-distillation loss
Gradient agreement maps and bar chart for iBOT ViT-B and CALiT-B under iBOT's CLS self-distillation loss
iBOT, ViT-B/16, CLS self-distillation loss

Results

26/26

dense-probe comparisons won by CALiT over its ViT: 7 pre-training setups on 4 datasets, wherever the baseline is reported.

+24.0

mIoU on Cityscapes for supervised ViT-S, the largest single gain.

300

pre-training epochs for CALiT-S with DINO and iBOT, against 800 for the public ViT-S checkpoints it outperforms.

modelpre-training (ep.)IN-1k top-1ADE20K ↑VOC ↑Cityscapes ↑COCO-Stuff ↑
ViT-S/16 · supervised IN-1k
DeiT-S/16sup. 30079.8524.4860.5127.5329.22
LocalViT-Ssup. 30080.8320.2349.8540.0122.41
CAT-S (CAg only)sup. 30079.9629.6267.6846.1034.79
CALiT-S (ours)sup. 30080.5731.6571.5051.5636.68
Δ CALiT − ViT+0.72+7.17+10.99+24.03+7.46
ViT-S/16 · self-supervised IN-1k, iBOT
iBOT ViT-S/16iBOT 80077.9 lp29.7866.6749.0934.71
CALiT-S (ours)iBOT 30076.3 lp32.8173.7855.9637.97
Δ−1.6+3.03+7.11+6.87+3.26
ViT-S/16 · self-supervised IN-1k, DINO
DINO ViT-S/16DINO 80077.0 lp23.7950.4145.8926.08
CALiT-S (ours)DINO 30075.9 lp27.4561.4749.3631.48
Δ−1.1+3.66+11.06+3.47+5.40
ViT-B/16 · supervised IN-1k
DeiT-III-B*sup. 40081.825.461.748.9n/a
CALiT-B (ours)sup. 30082.2032.2371.8853.0937.31
Δ+0.40+6.83+10.18+4.19n/a
ViT-B/16 · self-supervised IN-1k, iBOT
iBOT ViT-B/16iBOT 40079.5 lp36.2972.9954.8938.90
CALiT-B (ours)iBOT 30079.3 lp37.2575.8259.6641.08
Δ−0.2+0.96+2.83+4.77+2.18
ViT-B/16 · self-supervised IN-22K, DINOv2
DINOv2 ViT-B/16*DINOv2 600k80.4 lp38.376.658.4n/a
CALiT-B (ours)DINOv2 600k77.6 lp38.7377.3764.4441.51
Δ−2.8+0.43+0.77+6.04n/a
ViT-B/16 · iBOT IN-1k, then vision-language fine-tuning (LeVLJEPA, DataComp-12M)
iBOT ViT-B/16iBOT 400 + VL 5k44.78 zs36.9375.5254.5439.63
CALiT-B (ours)iBOT 300 + VL 5k44.65 zs38.5777.9360.1441.01
Δ−0.13+1.64+2.42+5.60+1.38

Semantic segmentation with a frozen backbone (Table 1 of the paper). DINOv2-style linear probe, a BatchNorm and a 1×1 convolution on the last-layer patch tokens, mIoU ↑. IN-1k top-1 is measured as each baseline does: supervised accuracy, a linear probe (lp) for self-supervised models, zero-shot (zs) after vision-language fine-tuning. Δ: CALiT minus the ViT of the same block. * mIoU as reported by Marouani et al. with the same probe. Bold: best per block. CALiT gives up a little ImageNet top-1 in 5 of the 7 settings, by 2.8 points at most (DINOv2-B).

Thin structures

The largest gains are on thin classes: fence, pole, traffic light, traffic sign, rider, motorcycle and bicycle. We also probe ACDC, urban driving under fog, rain and snow, with the Cityscapes head (zero-shot) and with a head trained on ACDC.

modelpre-training (ep.)Cityscapes valACDC zero-shotACDC trained
allthinallthinallthin
ViT-S/16 · supervised IN-1k
DeiT-S/16sup. 30027.4112.2115.416.4821.6210.33
CAT-S (CAg only)sup. 30046.1029.9025.4412.2237.4220.01
CALiT-S (ours)sup. 30051.0540.0628.3418.7740.7326.42
Δ CALiT − ViT+23.64+27.85+12.93+12.29+19.10+16.09
ViT-S/16 · self-supervised IN-1k, iBOT
iBOT ViT-S/16iBOT 80048.7331.0132.8417.5440.2521.41
CALiT-S (ours)iBOT 30055.8043.8137.9725.2548.3431.24
Δ+7.07+12.80+5.13+7.71+8.09+9.83
ViT-B/16 · self-supervised IN-1k, iBOT
iBOT ViT-B/16iBOT 40054.7537.1536.2822.3047.0026.77
CALiT-B (ours)iBOT 30058.4647.6638.4526.1553.0337.91
Δ+3.71+10.51+2.17+3.85+6.04+11.13

Thin-structure segmentation and robustness (mIoU ↑, linear probe as above). thin: the 7 thin classes. ACDC zero-shot: the Cityscapes-trained head applied to ACDC val. ACDC trained: a head trained from scratch on the 1600 ACDC training images for 4k iterations, evaluated on the 400 val images.

Ablations

modelIN-1k top-1ADE20K ↑VOC ↑Cityscapes ↑
CALiT-S (reference)80.5731.6571.5051.56
one change to CALiT-S, same recipe
CAg in the last block only79.5425.9960.7843.86
half-width CAg80.6429.8367.2950.76
half-width CAg, no LMix80.1629.1767.2447.72
mean-pool CAg (no attention)80.7329.5466.4748.19
DeiT-S/16, for reference79.8524.4860.5127.53

Single-variable ablations of CALiT-S (supervised IN-1k, linear probes as above; bold best, italic second). Placed only in the last block, CAg loses most of the dense gain: the earlier blocks still get the narrow gradient, so CAg belongs in every block. Applying CAg to half of the FFN width recovers most of the gain. A uniform mean in place of attention loses patch quality, more than half-width CAg does.

BibTeX

@article{chhatkuli2026calit,
  title   = {Vision Transformers Need Cross Aggregation for Patch Semantics},
  author  = {Chhatkuli, Ajad and Paudel, Pramish and Zhang, Deheng and
             Wu, Kanzhi and Van Gool, Luc and Paudel, Danda Pani},
  journal = {arXiv preprint},
  year    = {2026}
}