Research statement

What I work on

  • FocusInterpretable, efficient deep vision for safety-critical decisions
  • DomainsDriver monitoring · medical image segmentation
  • MethodsAttention-enhanced CNNs · CNN segmentation · Grad-CAM attribution

My research is on attention-based deep learning for visual recognition, and on the methods that make such models auditable and efficient enough to be used outside a laboratory.

Two things interest me in particular.

What the model attends to

In our study of driver distraction, a multi-head attention CNN reached a mean accuracy of 99.48% over stratified five-fold cross-validation. I find the Grad-CAM attribution more informative than that figure: the maps concentrated on hands, face and mobile device, which are the regions a human observer would use. Attribution is one way to check whether a model has learned the task or the dataset.

What that costs

The same model runs at 122 frames per second with 42.92 million parameters. I would like to understand that trade-off better — when attention and attribution can be added cheaply and when they cannot — because a method that needs large compute is available to fewer people.

My undergraduate thesis addressed a related problem in medical imaging: segmenting tumor regions from brain MRI with convolutional networks. I would like to continue in this direction for doctoral study, on attention-based and hybrid CNN–Transformer architectures whose explanations are evaluated as carefully as their accuracy.

Alongside this, my graduate coursework has covered sequence models — RNN, LSTM, encoder–decoder and Transformer architectures. The mathematical notes and implementations are on GitHub.

Projects

In detail

2026 · Published in Algorithms (MDPI)

Explainable driver behavior detection with multi-head attention

Problem. Driver-monitoring systems that flag distraction need to be accurate and auditable at the same time: a system that cannot show why it raised an alert is hard to evaluate and hard to trust.

Approach. We combined a convolutional backbone with a multi-head attention module, so the network captures both local spatial detail and long-range dependencies in the driver's activities. Grad-CAM was then used to produce attribution maps over the input, and an ablation study with k-fold validation isolated what the attention mechanism actually contributed.

Result. On the State Farm Distracted Driver dataset, performance held up consistently across all ten distraction classes in precision, recall and F1, and the model stayed fast enough for real-time monitoring.

99.48%Mean accuracy
122Frames per second
42.92MParameters
5-foldStratified CV

Abdullah Al Mamun, Md Shahidul Islam Shabuz, Md Nahidur Rahaman, Khawja Imran Masud, Md. Biddut Hossain

2020–2021 · B.Sc. thesis, DUET Gazipur

Brain tumor MRI segmentation using convolutional neural networks

Problem. Manual delineation of tumor regions in brain MRI is slow and varies between readers, which is what makes automatic segmentation worth pursuing.

Approach. I built the full pipeline: image preprocessing, a CNN-based segmentation model, training, and evaluation against reference annotations, supervised by Prof. Dr. Mohammod Abul Kashem.

Continuity. Segmentation in MRI and distraction recognition in a vehicle raise the same question: whether the reader can see what the model used to reach its answer. Hybrid CNN–Transformer segmentation is the direction I would like to take this next.

Doctoral study

I am applying to PhD programs for Fall 2027. If anything above is close to what your group works on, I would be glad to hear from you — shahidul.shabuz@gmail.com.