Below are the idea in plain words, the datasets, three scoring methods to compare, the metrics, the mistakes that invalidate results and the scope for B.Tech and M.Tech.
- The goal: a classifier that can answer unknown instead of guessing.
- Three scores to compare: maximum softmax probability, ODIN and the energy score. None needs the classifier retrained.
- Standard data: CIFAR-10 as the known classes, with CIFAR-100 or another image set as the unknown.
- Standard metrics: AUROC, and the false positive rate when 95% of known images are accepted.
- The main trap: tuning the method on the same unknown images you test on.
What is unknown class detection, in plain words?
A softmax classifier must spread 100 points of probability over the classes it knows. Show a cat-and-dog classifier a truck and it still answers cat or dog, often with high confidence. Unknown class detection adds a second check: a score that says how familiar the input looks, and a threshold below which the model answers unknown. The same idea protects a medical model from a scan it was never trained for, and a factory camera from a defect it has never seen.
Which datasets should you use?
| Role | Dataset | Detail |
|---|---|---|
| Known classes | CIFAR-10 | 60,000 colour images of 32 by 32 pixels in 10 classes, 6,000 per class, split into 50,000 for training and 10,000 for testing |
| Unknown, similar | CIFAR-100 | 100 classes with 600 images each. Remove classes that overlap with CIFAR-10 before you use it as unknown |
| Unknown, different | SVHN or a texture set | Images that look nothing like the training classes; the easier case |
| Alternative | Held-back classes | Train on six CIFAR-10 classes and treat the other four as unknown |
CIFAR figures are from the dataset’s own page.
How do you build it step by step?
- Train an ordinary classifier on the known classes, such as a ResNet-18, and record its test accuracy.
- Score with maximum softmax probability, the baseline of Hendrycks and Gimpel.
- Add ODIN. Liang, Li and Srikant use temperature scaling and a small perturbation of the input to pull known and unknown scores apart, without changing the trained network.
- Add the energy score of Liu and colleagues, computed from the same logits.
- Set the threshold on known validation images, at the point where 95% of them are accepted.
- Evaluate each score on each unknown set, in one table.
- Build the demo: upload any image and get a class with its confidence, or the answer unknown.
Which metrics should you report?
- AUROC: how well the score separates known from unknown images across all thresholds.
- False positive rate at 95% true positive rate: the share of unknown images still accepted when 95% of known ones are. Lower is better.
- Accuracy on the known classes, to show that detection did not harm classification.
The papers show the size of change to expect. The ODIN paper reports the false positive rate on a DenseNet trained on CIFAR-10 falling from 34.7% to 4.3% at a 95% true positive rate, and the energy paper reports the average false positive rate falling by 18.03% against softmax confidence on a WideResNet trained on CIFAR-10. Your own numbers will differ with the network and the unknown sets.
What mistakes invalidate the results?
- Tuning on the test unknowns. Choose temperature, perturbation size and threshold on known validation data or a separate unknown set.
- Overlapping classes. If an unknown set contains trucks and so does training, the result means nothing.
- One easy unknown set. Report a similar set and a different one.
- Saying it detects anything unknown. It detects the kinds you tested.
- Leaving out the baseline. Every method is judged against maximum softmax probability.
How does the scope change for B.Tech and M.Tech?
| Level | Scope | What to show |
|---|---|---|
| B.Tech mini | A softmax threshold on a small classifier, with held-back classes as unknown | Score histograms for known and unknown images |
| B.Tech major | Three scores compared on two unknown sets | A results table, ROC curves and an upload demo |
| M.Tech / M.E. | A base paper reproduced, plus one extension: training with outlier examples, unknown objects in detection, or a medical imaging dataset | An ablation study and a paper in IEEE format |
This topic is too abstract for most Diploma projects; a better fit is on our Diploma projects page. For M.Tech work, see how we build M.Tech projects from a base paper, and the trustworthy AI questions in our PhD topics in computer science.
What goes in the report, and which viva questions come up?
State the problem with one clear failure example, review the three methods, describe the setup exactly, then give the table and the curves. A paper for this topic follows the IEEE paper format. Prepare these:
- Why is a softmax classifier confident on inputs it has never seen?
- What does temperature scaling change?
- How is the energy score computed from the logits?
- What does a false positive rate at 95% true positive rate mean in practice?
- How is this different from anomaly detection?
Related guides: AI-generated image detection and the explainable AI project ideas in our machine learning list.
This guide is a plan to discuss with your guide, and the figures in it come from the linked dataset pages and papers, not from our own work. Our 8 delivered case studies are M.Tech and M.E. projects, each shown with its real paper pages, screens and results.
Frequently asked questions
What is unknown class detection in machine learning?
It is the task of recognising that an input does not belong to any class the model was trained on. It is also called out-of-distribution detection or open-set recognition. The model keeps its normal classes and gains one more possible answer: unknown.
What is the simplest method for out-of-distribution detection?
The maximum softmax probability baseline from Hendrycks and Gimpel. Take the classifier’s highest class probability as a confidence score, and call the input unknown when the score falls below a threshold. It needs no retraining, so it is the first row in every comparison table.
Which datasets are used for out-of-distribution detection?
A common setup trains on CIFAR-10, which has 60,000 colour images of 32 by 32 pixels in 10 classes, and uses other image sets such as CIFAR-100 or SVHN as the unknown inputs. You can also hold back some classes of one dataset and treat them as unknown.
Is unknown class detection a good M.Tech project?
Yes. It has clear base papers, standard metrics and plenty of room for one measured extension, such as a new scoring rule, a harder near-unknown test set or unknown objects in detection. For B.Tech, comparing three scores with a demo is enough.
Can The Ultimate Project World help with an out-of-distribution detection project?
Yes. Tell us your level, your base paper if you have one, and your review dates in the free consultation. We will tell you exactly which parts we can take on, such as the code, the experiments, the paper, the thesis and viva preparation.
