Deep Diaries · · 5 min read
Exploring YOLOv9: PGI Integration and GELAN Architecture for Enhanced Object Detection
YOLOv9 .. Advancements in Efficiency, Accuracy, and Competitiveness
Originally published on Deep Diaries on Substack. Reproduced here as written.

Recently , YOLOv9 was published based on YOLOv7 and Dynamic YOLOv7 but with a superior performance on the other objects detectors and we are going to discuss the key ideas behind YOLOv9. but before we start to discuss YOLOv9 there are some points we need to understand first
Deep Supervision
Technique used in deep learning tasks. It involves adding intermediate supervision signals at multiple layers of a deep neural network.
In traditional deep learning architectures, such as convolutional neural networks (CNNs), supervision is typically applied only at the final layer, where the network's predictions are made. However, deep supervision introduces additional supervision signals at intermediate layers of the network.
With deep supervision, the loss function is applied not only at the output layer but also at intermediate layers. This provides additional gradient signals that can help alleviate the vanishing gradient problem and facilitate more effective training. Additionally, deep supervision can encourage the network to learn useful features at different levels of abstraction, leading to potentially better overall performance on the task.
Information Bottleneck Principle
(The main problem addressed by YOLOv9)
Simply information bottleneck states that each time we go deeper in a neural network (data transformation) we loss information.
Assuming that “I” indicates the mutual information between two quantities, “f, g” are transformation function (aka network layers) and “X” is the original data, then
which means that the last (output) layer contains less information about the original data than any of the layers, considering that the network depends on the output of last layer and the labels (target) to calculate loss and gradients to update network parameters so that may result in “unreliable gradients and poor convergence”
this problem can be solved by introducing the transformations as a series of a reversible functions that can be tracked back to the original data.
Reversible Functions
“When a function r has an inverse transformation function v we call this function reversible function”
for simplification lets consider the following function r
and its reverse function v
In this case
“When the network’s transformation function is composed of reversible functions, more reliable gradients can be obtained to update the model. almost all of today’s popular deep learning methods are architectures that conform to the reversible property, such as
where l indicates the l-th layer of a PreAct ResNet and f is the transformation function of the l-th layer. PreAct ResNet repeatedly passes the original data X to subsequent layers in an explicit way.”
But as you can imagine this method has two issues:
It works well with very deep networks and fail with shallow ones
For difficult problems it’s not easy for us to find a mapping between input data and target labels
Which led to YOLOv9 idea that’s called programmable Gradient Information, which is a modified architecture base on reversible functions and deep supervision but with some updates to help solving the problem of data bottleneck without adding inference costs.

Programmable Gradient Information (PGI)
An auxiliary supervision framework which includes 3 components:
Main branch (Main network - only one used in inference time)
Auxiliary reversible branch (Deal with the deepening problems, information bottleneck)
Multi-level auxiliary information (Handle deep supervision error accumulation)
As we said previously that network deepening will cause information bottleneck, which will make the loss function unable to generate reliable gradients, YOLOv9 introduced the following concepts to mitigate this problem.
Auxiliary Reversible Branch (4.1.1)
Which provides information that maps from data to target by adding a reversible architecture, but that’s also will result in increasing in inference costs.
And to target this problem YOLOv9 introduced the reversible architecture as an expansion of deep supervision branch that will enhance the main branch but will be ignored in inference time and will also make the reversible architecture works well with shallow networks.
“Moreover, the reversible architecture performs worse on shallow networks than on general networks because complex tasks require conversion in deeper networks. Our proposed method does not force the main branch to retain complete original information but updates it by generating useful gradient through the auxiliary supervision mechanism. The advantage of this design is that the proposed method can also be applied to shallower networks.”
Multi-level Auxiliary Information (4.1.2)
One issue of deep supervision architectures is that after connecting the deep supervision architecture to a feature pyramid the shallow features will be guided to learn the features required for small object detection, and at this time the system will regard the positions of objects of other sizes as the background which will case the feature pyramid to loss information about target objects.
And to mitigate this problem YOLOv9 introduced an idea that each feature pyramid needs to receive information about all target objects so that subsequent main branch can retain complete information to learn predictions for various targets by simply insert an integration network between the feature pyramid layers and the main branch which will result in combining the gradients from different prediction heads to update the main branch.
“At this time, the characteristics of the main branch’s feature pyramid hierarchy will not be dominated by some specific object’s information. As a result, our method can alleviate the broken information problem in deep supervision.”
Generalized Efficient Layer Aggregation Network
Also YOLOv9 introduced a new network architecture that’s called Generalized Efficient Layer Aggregation Network (GELAN) which is a combination between CSPNet and ELAN

“We designed generalized efficient layer aggregation network (GELAN) that takes into account lighweight, inference speed, and accuracy.
We generalized the capability of ELAN, which originally only used stacking of convolutional layers, to a new architecture that can use any computational blocks.”
Technically YOLOv9 paper presented the results of the GELAN architecture using Conv Blocks [ELAN], Res blocks, Dark blocks, and CSP blocks as you can see in the following table.

Results
After integrating all of the previous ideas we discussed YOLOv9 out-perform the state-of-art object detectors as you can see in the following table

Conclusions
YOLOv9 introduces PGI to address the information bottleneck problem and the challenge of deep supervision mechanisms in lightweight neural networks. It presents GELAN, a highly efficient and lightweight neural network designed specifically for object detection tasks. GELAN demonstrates robust performance across different computational blocks and depth configurations, making it adaptable to various inference devices.
By incorporating PGI, both lightweight and deep models witness significant accuracy enhancements. The fusion of PGI and GELAN in YOLOv9 enhances its competitiveness, showcasing notable improvements. YOLOv9 achieves a reduction of 49% in parameters and 43% in computational load compared to YOLOv8, while still exhibiting a 0.6% increase in Average Precision (AP) on the MS COCO dataset.