Small object detection.

computer-vision
object-detection
Author

Mirza Tupkušić

Published

August 8, 2026

1 Introduction

Two weeks ago, I began a post on object detection. I never finished it completely, but here I am starting another one anyway. The plan at initiative was to focus on the particular difficulties of detecting small objects, but I got completely sidetracked. Why focus on the little ones? I recently built a detector for small objects, and it failed miserably. I took many off-the-shelf models, fine-tuned them, and ended up disappointed in them again and again. Let’s see why that has happened.

In an image, object sizes theoretically range from occupying every pixel down to occupying a single one. Intuitively, larger objects should be easier to detect there’s simply more to look at. And if that’s true, then where’s the line where things start to fall apart? Answering that can be as easy as citing the COCO dataset (Lin et al. 2014) or as hard as writing a chapter of a book. The size categories and thresholds COCO uses:

  • Small Objects: Area < 32 \times 32 pixels (or < 1,024 pixels total area)
  • Medium Objects: 32 \times 32 ≤ Area ≤ 96 \times 96 pixels (between 1,024 and 9,216 pixels area)
  • Large Objects: Area > 96 \times 96 pixels (or > 9,216 pixels total area)

Dataset Distribution of COCO dataset:

  • Small: Roughly 41% of all object instances
  • Medium: Roughly 34% of all object instances
  • Large: Roughly 24% of all object instances

What do these numbers actually mean for us? Let’s see how RT-DETR (Zhao et al. 2024), a perfectly respectable detector, actually holds up across these categories and whether “there’s more data for big objects” survives contact with reality.

Model Scale AP AP_S AP_M AP_L
RT-DETR-R18 S 49.2 33.2 52.3 64.8
RT-DETR-R50 M 55.3 37.9 59.9 71.8
RT-DETR-R101 L 56.2 38.3 60.5 73.5

Looks like as the object gets smaller the average precission worsens. Though, size of the model improves the score with some visible saturation. This now explains why I have been failing to detect small objects. Well, the truth is that I had a quite interesting problem. According to how COCO sees small objects, my objects weren’t small, but tiny, with areas between 5 \times 5 and 10 \times 10 pixels. So what is happening? I have mentioned in my earlier object detection post that todays object detection models follow a architectral patter whith 3 main parts: the backbone, the neck and the detection head. The backbone computes image features, the neck should acomplish any feature interaction, information sharing between feature scales and the detection head with it should detect objects. As it sounds great, there are couple of challenges introduced down the line. So, what can we do to tackle these challenges? Well, in one blog post we certainly won’t reinvent the wheel when it comes to modern object detection standards. That job needs years of experience, research time, founding, etc. Instead, we can look at our existing options and tweak them to push the model’s accuracy as far as possible.

Figure 1: AI Generated – Object detection standard architecture.

2 Standard object detection pipeline

2.1 Backbone changes

Object detector must extract features from an input image so that the neck and detection heads have meaningful data to work with. While some modern backbones rely on attention mechanisms (like Swin Transformers (Liu et al. 2021)), most today are still based on convolutions, such as ResNet (He et al. 2016) or ConvNeXt (Liu et al. 2022).

Let’s focus on convolutions. A typical ResNet architecture is divided into four main stages, producing four distinct feature maps at cumulative strides of 4, 8, 16, and 32. The term stride refers to how many pixels a convolutional filter shifts as it moves across the input image. If the stride is 1, the filter moves by 1 pixel at a time. If the stride is 2, the filter skips by 2 pixels, which halves the spatial dimensions (both width and height) of the output feature map. As the image passes through ResNet’s four stages, this downsampling accumulates. For a standard 640 \times 640 input image, the spatial resolutions of the four consecutive feature maps shrink to:

  • Stage 1 (Stride 4): 160 \times 160
  • Stage 2 (Stride 8): 80 \times 80
  • Stage 3 (Stride 16): 40 \times 40
  • Stage 4 (Stride 32): 20 \times 20

Now, think back to the COCO dataset definition of a small object: anything under 32 \times 32 pixels. By the time the image data reaches that final 20 \times 20 feature map, those small objects have been completely annihilated. We are here talking about an 32 times downscale! Note, that convloutions benefit from receptive field, so the information about these objects is ussually saved in other feature maps cells. Because of this, passing only the final feature map to the detection head is a terrible idea for small objects. Instead, using the earlier, higher resolution feature maps is important.

Shallow feature maps preserve precise spatial information like sharp edges and fine corners. Deeper feature maps capture abstract, contextual information. To detect objects across a range of sizes, it helps to combine feature maps from different stages: the shallow, high resolution maps retain the detail small objects need, while the deep maps carry the semantic strength that helps with larger ones. This is exactly what a neck in the architecture is doing. There are different ones, FPN-style (Lin et al. 2017) and its variations are quite popular today.

So what can we do about small object detection here? We already said we’re not going to play researcher here, and with that in mind, we don’t have that many options on the backbone side. A few of them:

  • Increasing the input image size, which yields richer, higher-resolution feature maps. Sadly, this raises the network’s complexity and slows down both training and inference quite a bit.
  • Scaling up the backbone to a larger model. Same trade off as before a bigger model means more computation.
  • Trying different backbones and tuning their parameters.

2.2 Neck

The next element of the network is a interesting one. As mentioned earlier, the core philosophy of a modern neck is to provide a way to merge multi-scale feature maps so the model can detect objects of varying sizes. Ideally, the neck must seamlessly blend deep contextual features with high-resolution spatial details into a unified representation for the detection head.

This feature fusion can be achieved in several ways. Some architectures rely entirely on CNN-based fusion paths such as FPN (Lin et al. 2017), PANet (Liu et al. 2018), or BiFPN (Tan et al. 2020) while others use attention-based transformers, and some combine both approaches.

Which one is the best? There is no single right answer, as they all perform reasonably well depending on the use case. CNN-based networks use elegant top-down and bottom-up pathways to propagate semantic information down to spatial layers, and spatial details back up to semantic ones. On the other hand, transformer encoders use self-attention across every input token to capture rich, global context.

So, what can we actually do to improve small object detection? A great starting point is to experiment with different neck architectures, tweak their structures, introduce a mechanisms to force the network to focus on smaller objects, etc. While there are countless strategies to try, the real challenge lies in finding out which one will actually deliver results. We slowly but surely find ourself entering the territory of extremely demanding research and development. It quickly becomes an endless loop of brainstorming, implementing, and testing until something finally works. In this domain, the approach of throwing everything at the wall to see what sticks works wonders.

For many engineers in the industry and myself in the past this is exactly where things drop like a hot potato. Whatever I tried, I got average results. I didn’t got anything that stood out compared to off-the-shelf models. Months and months of experimenting, only to come to the conclusion to just use YOLO and chill. This was the moment I understood the current state of object detection and accepted that beating a SOTA model by a noticeable margin requires something truly special in terms of architectural change.

With that in mind I didn’t lose the hope, but my perspective on the problem shifted.

2.3 Detection head

The detection head is the last element in the standard pipeline, and here the story is much the same as before: research, brainstorm, implement, and test, over and over. That said, when the goal is small-object performance specifically, most of the use lives in the earlier stages, the backbone and neck rather than the head. The head still matters, but it’s rarely the bottleneck for tiny objects.

3 Conclusion

This post could go further with topics about loss functions, processing ideas, special network elements, etc. I will end it here with the standard pipeline so this doesn’t go out of hand like before. And I don’t want to torture myself further. Starting a blog post is kinda easy at the beginning, but finding the time, right words to express something and ending it drains me.

What can we do about small-object detection? From my experience, these are the things that produced the biggest gains in past projects, and I hope they make a difference for others too:

  • Increase the number of small-object samples. This is the most straightforward one. Real examples, hard use cases in the training and validation sets combined with the right architecture will push the network to learn features that actually fire on small objects.

  • Increase the input image size. The problem, as the image passes through the backbone, is that it gets progressively downscaled, which is what makes small objects fade in the deeper feature maps. To counter this, we can upsample the input first. This is probably the next most valuable fix for small-object detection. It demands more compute, but more pixels per image push detection accuracy higher. The catch with upsampling is that it is an artificial increase in detail, so it can hallucinate pixels and lead to odd artefacts. Modern methods like super resolution do a fairly good job. So if the added compute is not a problem it’s worth a shot. However, I would first advise using a higher resolution camera or getting a high quality optical lense with a focal length that satisfies the problem in terms of needed field of view.

  • Run inference on tiled windows. Also known as SAHI (Slicing Aided Hyper Inference) (Akyon et al. 2022): the input image is cut into multiple windows the size of the network’s input, and inference is run on each. This helps when the input image is larger than the network input. Resizing the whole image down would throw away a lot of information and hurt performance. Slicing into crops preserves the original resolution, so the model gets to work at full quality.

  • Use higher resolution feature maps. As explained earlier, these retain the fine spatial information that makes a big difference for small objects.

  • Apply spatial or local attention modules to prioritize fine grained, small-object features. If we’re building a small-object detector, we want the network to put more focus on features coming from the small ones. I hope to write something in future about this.

  • Use a loss that penalizes errors on small objects more heavily. This pushes training to emphasize small-object features. I also plan this for a future post.

  • Tune the anchors for small objects. If the network is anchor-based (older YOLO versions (Redmon et al. 2016), Faster R-CNN (Ren et al. 2015), etc.) it’s important to refine the anchor scales and aspect ratios so the priors actually cover small objects. This is one of the biggest fixes for such models, because it addresses the root cause directly: an anchor is only assigned as a positive sample when it overlaps a ground truth box above an IoU threshold, and if every anchor is far larger than a tiny object, almost none of them clear that threshold, so the small object contributes little or no training signal. Right sizing the anchors puts positives back on the board.

That’s pretty much it for now. I hope this information helped you in some way. See ya!

References

Akyon, Fatih Cagatay, Sinan Onur Altinuc, and Alptekin Temizel. 2022. “Slicing Aided Hyper Inference and Fine-Tuning for Small Object Detection.” Proceedings of the IEEE International Conference on Image Processing (ICIP). https://arxiv.org/abs/2202.06934.
He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. “Deep Residual Learning for Image Recognition.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1512.03385.
Lin, Tsung-Yi, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. “Feature Pyramid Networks for Object Detection.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1612.03144.
Lin, Tsung-Yi, Michael Maire, Serge Belongie, et al. 2014. “Microsoft COCO: Common Objects in Context.” European Conference on Computer Vision (ECCV). https://arxiv.org/abs/1405.0312.
Liu, Shu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. 2018. “Path Aggregation Network for Instance Segmentation.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1803.01534.
Liu, Ze, Yutong Lin, Yue Cao, et al. 2021. “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows.” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). https://arxiv.org/abs/2103.14030.
Liu, Zhuang, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. “A ConvNet for the 2020s.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/2201.03545.
Redmon, Joseph, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. “You Only Look Once: Unified, Real-Time Object Detection.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1506.02640.
Ren, Shaoqing, Kaiming He, Ross Girshick, and Jian Sun. 2015. “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/1506.01497.
Tan, Mingxing, Ruoming Pang, and Quoc V. Le. 2020. “EfficientDet: Scalable and Efficient Object Detection.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1911.09070.
Zhao, Yian, Wenyu Lv, Shangliang Xu, et al. 2024. “DETRs Beat YOLOs on Real-Time Object Detection.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/2304.08069.