Case study · M.Tech · Medical imaging

Diabetic retinopathy detection using deep learning

This M.Tech project does diabetic retinopathy detection using deep learning: it grades retinal fundus photos on the five-grade clinical scale by combining two public Vision Transformer (ViT-B/16) models into one ensemble, without retraining either network.

Last updated

Results page from the project’s IEEE-format paper on diabetic retinopathy detection: per-sample prediction, ablation and latency tables, each with a matching chart.
IEEE-format paper
Home screen of the retinal disease classifier: ViT-B/16 model summary cards and a gallery of sample fundus images, one for each diabetic retinopathy grade.
Web app
Inference latency for each stage of the diabetic retinopathy detection pipeline, shown as a pie chart and a bar chart, with the ViT forward passes taking most of the time.
Results

From this delivered M.Tech project, shown with names and institute details removed. Select an image to zoom.

Fundus-specific preprocessing makes lesions easier to see, and attention rollout draws a heatmap of the retinal regions behind each grade. The student received a FastAPI web app, the Python code, an IEEE-format paper and the M.Tech report.

What is this project, at a glance?

Project at a glance
LevelM.Tech
DomainMedical imaging, explainable AI (XAI)
Problem typeFive-class image classification: diabetic retinopathy severity (No DR, Mild, Moderate and Severe NPDR, Proliferative DR) from one fundus photo
Core methodsEnsemble of two pretrained ViT-B/16 checkpoints with label remapping, circular crop and Ben Graham preprocessing, flip test-time augmentation, attention rollout
StackPython, PyTorch, Hugging Face Transformers, FastAPI, plain HTML/CSS/JavaScript
DeliverablesIEEE-format paper, M.Tech report, web app, Python codebase, presentation outline

What problem does this project solve?

Diabetic retinopathy, damage to the retina’s blood vessels caused by diabetes, is a leading preventable cause of blindness. Catching it early means grading fundus photos regularly, and there are far fewer eye specialists than screenings needed.

Deep learning can grade these photos, but the student faced three blocks:

  • Single public models were unreliable. Two openly shared ViT checkpoints had opposite biases: one over-called mild disease, the other missed advanced disease.
  • Grad-CAM does not fit transformers well. It assumes a CNN’s spatial feature maps, which a ViT does not have.
  • Retraining was out of reach. Large fundus datasets need access approval and a GPU.

So the question became: can diabetic retinopathy detection using Vision Transformers work with an ensemble of existing models plus better preprocessing, and give a more useful and explainable grader on an ordinary laptop?

How does the diabetic retinopathy detection pipeline work?

Every image goes through five steps. No network is trained; the gains come from how existing models are combined and what they are shown.

  1. Crop to the retina

    The black border is cut away, so more of the model’s 224 × 224 input is actual retina.

  2. Make lesions stand out

    A Ben Graham transform subtracts a blurred copy of the image, evening out the lighting so small lesions such as microaneurysms stand out.

  3. Run both models on two versions

    Each ViT-B/16 model scores the raw and the enhanced image, each together with its mirror image (test-time augmentation).

  4. Put both models on one scale

    The two checkpoints order their labels differently, so a mapping layer moves every output onto the same five grades before averaging. Without it, the ensemble would silently mix up classes.

  5. Explain the answer

    Attention rollout traces attention through all 12 transformer layers and overlays the image patches that drove the prediction as a heatmap.

Four-layer architecture diagram: HTML, CSS and JavaScript front end, FastAPI application layer, ensemble service with attention rollout, and PyTorch with Transformers model layer.
ArchitectureFour-layer system architecture: HTML/JS front end, FastAPI application layer, ensemble service (CCER + attention rollout) and PyTorch/Transformers model layer.

What did we build?

A five-tab single-page web app (Predict, Model, Diseases, Results, About) on a FastAPI backend. Drop in a fundus photo or pick a bundled, openly licensed sample, and the app returns the grade, a short clinical note, the probability of every grade and the attention heatmap.

The front end uses no JavaScript framework, so it is easy to explain. An optional ResNet50 training script for the RFMiD dataset is kept for later work.

Home screen of the retinal disease classifier with model summary cards and sample fundus images, one for each diabetic retinopathy grade.
Web app · home screenHome screen of the retinal disease classifier: ViT-B/16 model summary cards and a gallery of sample fundus images, one for each diabetic retinopathy grade.
Live prediction in the web app: a fundus image graded Moderate NPDR, with a probability bar for each of the five severity grades and an attention-rollout heatmap over the retina.
Web app · live predictionLive prediction in the web app: an uploaded fundus image is graded ‘Moderate NPDR’, with a probability bar for each of the 5 severity grades and an attention-rollout heatmap showing where the model looked.

Shown with student, guide and institute details removed.

What did the results show?

Because nothing was retrained, the evaluation was small and qualitative: openly licensed fundus photos covering every grade, a treated retina and one out-of-distribution case. It is a proof of concept, not a clinical benchmark.

  • Ensemble vs single models. The ensemble got more grades right than either model alone and did not call the out-of-distribution image diabetic retinopathy.
  • Ablation. Ben Graham preprocessing was the one change that moved an advanced-disease sample from “No DR” to a disease grade.
  • Weak spot. Mild and moderate grades were still under-called, likely because of imbalanced data behind the public checkpoints.
  • Speed. A full prediction with heatmap took 2.11 s on a laptop without a discrete GPU.
Pie chart and bar chart of inference latency per pipeline stage, totalling 2.11 seconds, dominated by the four ViT forward-pass stages.
Results · latencyInference latency for each pipeline stage (total 2.11 s), shown as a pie chart and a bar chart. The four ViT forward passes take most of the time.
Two-column results page from the IEEE-format paper with per-sample prediction, ablation and latency tables and matching charts.
IEEE-format paperResults page from the IEEE-format paper: per-sample prediction, ablation and latency tables, each with a matching chart.

Shown with student, guide and institute details removed.

What did the student receive?

Everything needed to run, submit and defend the project, with the paper written to the IEEE paper format for conferences.

What was delivered
DeliverableWhat it contains
IEEE-format paperPDF and LaTeX source: method, results, ablation, latency and limitations
M.Tech reportProject report (thesis) as PDF and LaTeX source
Web appFastAPI backend and a five-tab front end with sample images
Python codebaseEnsemble inference, preprocessing, attention rollout, a command-line predictor and a setup README
Presentation outlineSlide-by-slide outline: problem, contributions, results, limitations and future work

For your own report or slides, see dissertation, report & PPT support.

How could you adapt this project?

Keep the pipeline and change one thing, so your project has its own contribution. A lighter version (one model, the heatmap and a simple web page) can suit a B.Tech final year project; for more medical imaging topics, see our computer vision ideas among machine learning projects for final year.

Topic ideas, not delivered projects

Suggestions to discuss with your guide. The delivered work is the project above.

  • Topic ideaTrain your own ensemble memberFine-tune a model on APTOS 2019 with a class-balanced loss and measure the change on mild and moderate grades.
  • Topic ideaMulti-disease retinal screeningMove from five DR grades to multi-label diagnosis on RFMiD, including glaucoma and macular degeneration.
  • Topic ideaSharper transformer explanationsCompare attention rollout with gradient-weighted methods and check which heatmaps match visible lesions.
  • Topic ideaGrading on a phoneExport the models to ONNX, quantise them and grade fundus photos offline on a smartphone.

Frequently asked questions

Can I get a similar diabetic retinopathy detection project?

Yes. We build a new project around your own topic, dataset and guide’s requirements, such as another grading approach or another retinal disease. The scope is agreed in the free consultation.

Which datasets can a diabetic retinopathy detection project use?

Common public fundus datasets are APTOS 2019, EyePACS, Messidor and RFMiD. This project used two public pretrained checkpoints and openly licensed sample photos.

Do I need a GPU for a Vision Transformer project?

Not for this design: with no retraining, it ran on a laptop without a discrete GPU. Fine-tuning a ViT usually needs a cloud notebook GPU or a lab machine.

What will I need to explain in the viva?

Expect questions on how a ViT splits an image into patches, why two models with opposite biases make a useful ensemble, what Ben Graham preprocessing does, how attention rollout differs from Grad-CAM, and the limits of a small evaluation set.

Is this a medical diagnosis tool?

No. It is an academic prototype, and the app says so. Clinical use would need validation on graded patient images and regulatory approval.

Delivered work

Related case studies

More medical imaging projects we delivered, shown with names and institute details removed.

Planning a medical imaging project? Start with your topic.

Share your level, topic and deadline. You get a written plan and a fixed quote, and the consultation is free.

B.Tech · M.Tech · PhD · Projects · Papers · Thesis · Reports

WhatsApp Free consultation