Uncertainty-aware 3D Autolabeler
Human 3D perception seems effortless as our biology enables us to capture 3D information effectively from the world; our physical senses collect information every time we interact with the world, and our brain processes those information intelligently so we could perceive the 3D world. Our ability to estimate scene depth, recognize and segment objects almost seems automatic. But a computer's understanding of the world is currently represented through images or 3D data. These data help the computer or the model to "see" the 3D world that we live in. Images provide meaningful visual representation of the world, but it's 2D nature limits computer to fully grasp 3D real-world understanding. On the other hand, 3D data captures richer information about the world and aids computers to better understand the 3D visual world.
Many machine learning or computer vision models learn through a supervised fashion. These models need 'supervision' from the true label during the training (or learning). This is akin to a kid trying to learn the difference of a cat or dog. One would instinctively show the kid tens or hundreds of photos of these animal while describing whether the photos are cats or dogs (the true label). After several examples, the kid has been 'trained'; it already learns how to differentiate the two. In order to teach a computer model to "see" in the real world, one way is to give it the true labels during the learning or training process. If we want a model to detect objects from 2D images, we could give it bounding boxes for cars or traffic lights. To help the model detect objects from the 3D data, we could give 3D bounding box surrounding the car in the 3D space.
Supervised learning works in many domains, but the challenge in 3D detection is the need for hugely labeled data. This requires the need for a set of images and 3D data with multiple 2D and 3D boxes. To annotate or create these 3D bounding boxes for numerous images or 3D data is costly. Imagine the amount of labor required to draw 2D boxes or 3D boxes. That's tedious!
That is the motivation of our work: we tackle 3D automatic annotation from few labeled data. Our proposed model, called MEDL-U, does not need to learn from a lot of annotated data (data with bounding box labels). Instead, we allow the 3D annotation model to learn from relatively fewer samples (like 10% of the entire dataset). A consequence of training from fewer samples is noisy pseudolabels. To make it effective, we require the model to not only produce bounding box parameters but also a measure of uncertainty or noise of the predictions. As an illustrative example, we say that the model predicts the bounding box height, width, and center for a car, but it also says "I'm 30% unsure of this." This is the key idea of our work.
Our paper Uncertainty-aware 3D Automatic Annotation based on Evidential Deep Learning, which was accepted at the International Conference on Robotics and Automation conference in 2024, tackles this problem. We work on the problem of 3D Automatic Annotation from few labels by incorporating an uncertainty estimation module to a Transformer-based 3D annotator.
Given the input 2D image with 2D bounding boxes of cars and LiDAR (point cloud) data, we train our autolabeler to learn to predict 3D boxes of cars. Using only a few labeled data, we train the model to predict the 3D box parameters such as the center, length, width, height, and rotation. In addition to that, the model is also trained to predict the noise or uncertainty associated with each predicted parameter. For example, when the model is uncertain of the predicted length due to some ambiguity in the LiDAR point cloud data, it would give a high uncertainty value for the length. When the car is occluded, the model can be uncertain about the predicted center, giving center a high uncertainty value.
We utilize a model called MTrans to be the backbone feature extraction and multitask learner for our framework. For the uncertainty estimation, we utilize a uncertainty estimation framework called the Evidential Deep Learning (EDL) framework. Basically, in EDL, we assume the parameters of the 3D bounding box follows a Normal distribution with unknown mean and variance. We further assume that the mean and the variance are from a Gaussian prior and inverse-gamma prior, respectively. The distribution of the parameters mean and variance then follows the normal inverse-gamma distribution. This allows us to train a uncertainty estimation module in the form of a feedforward neural network that will predict the parameters of the inverse-gamma distribution (we omit the math details here). These parameters are used to calculate our prediction and the associated uncertainty (please see our paper for the full details).
In the end, we come up with a trained automatic annotator that can be utilized to annotate the remaining 3D data in the dataset; we predict the 3D box parameters in addition to the uncertainty estimates. Now we have labels for our entire dataset and uncertainty estimates to quantify the noise with the models prediction.
How can we use these box parameters and uncertainty estimates to train downstream 3D detectors? Note that popularly 3D detectors are trained in a deteministic fashion using only the labels or pseudolabels without taking into account the uncertainty. In our case, we would transform any deterministic 3D detector to probabilistic one and use the pseudolabels and the uncertainty for the training.
In the experimental results, we demonstrate that the proposed semi-supervised approach MEDL-U can outperform other semi-supervised approaches of 3D object annotation. The uncertainty estimates proved to be effective in improving downstream training of 3D detectors. In addition, probabilistic detectors trained using outputs of the MEDL-U surpass deterministic detectors trained using output from previously proposed annotators. It also achieves state-of-the-art results on important 3D benchmark, the KITTI official test set.
Our work, among a vast literature of studies on uncertainty estimation of machine learning models, shows the effectiveness of using uncertainty as an crucial signal to improve vision models for 3D object detection. Labels, whether human or AI-generated, always contain noise; hence machine learning models must effectively learn and estimate this noise as well. It is our hope that more AI models for tasks such as depth estimation, image segmentation, or language modeling are equipped with effective and mathematically grounded uncertainty estimation module, providing a measure of trust and reliability to AI models. In the age of advanced LLM capabilities, generative AI, AI agents, and information abundance, trust is a becoming scarce. And humanity could be aided by models that estimate the trustworthiness of their outputs.