Varifocal Loss: Learning IoU-Aware Classification Scores for Detection
· 2 min read · 382 words
Varifocal Loss (VFL) trains a dense object detector to predict an IoU-aware classification score: a score that represents both the presence of an object and the localization quality of its predicted box.
For the target class, the training target is the IoU between the predicted box and its matched ground-truth box. For background and non-target classes, . If is the predicted class probability, the paper defines
Why it differs from focal loss
Focal loss was designed to reduce the influence of numerous easy negatives. Varifocal Loss keeps that behavior for negatives through , but treats positives asymmetrically:
- Positive targets are continuous IoU values rather than one-hot labels.
- High-quality positives receive more weight through the outer factor .
- Positive examples are not down-weighted with the focal factor used for negatives.
This trains the classification score itself to rank well-localized boxes above poorly localized ones. That alignment matters because detectors select top candidates and apply non-maximum suppression using their scores.
Numerical and training details
The paper uses and in its main experiments. Those values belong to its detector and training recipe rather than to the mathematical definition, so a different detector or assignment strategy may need different settings.
Scope
Varifocal Loss was proposed as one part of VarifocalNet, alongside a star-shaped feature representation and box-refinement branch. Do not attribute the detector's complete reported improvement to the loss alone: the paper's ablation separates these components.
Primary source
- Haoyang Zhang et al., VarifocalNet: An IoU-Aware Dense Object Detector, CVPR 2021.
Keep reading
- Squared ReLU: The Tiny Activation Change Used by Primer
Squared ReLU computes the square of a rectified activation. This note explains why Primer used it, how it changes gradients, and when not to adopt it blindly.
- Mix-FFN in SegFormer: Adding Local Context to Transformer MLPs
SegFormer Mix-FFN inserts a depthwise 3×3 convolution between two feed-forward projections, giving image tokens local spatial context without explicit positional embeddings.
- LayerScale: A Small Change That Stabilizes Deep Vision Transformers
LayerScale multiplies each Transformer residual branch by a learned per-channel gain initialized near zero, helping very deep image Transformers begin close to an identity mapping.
