Publications
Peer-reviewed papers, workshop papers, and work under review, in reverse chronological order. An asterisk (*) marks equal contribution.
2026
-
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings
In Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026
Main track. Co-first author.
Geo-temporal understanding, the ability to identify the location, time, and contextual features of an image from visual cues alone, is a fundamental aspect of human intelligence with wide-ranging applications, from disaster response to autonomous navigation and geography education. While recent vision–language models (VLMs) have shown progress in image geo-localization using conspicuous cues like landmarks or road signs, their ability to understand temporal signals and related spatial reasoning cues remains underexplored. To address this gap, we introduce TimeSpot, a comprehensive benchmark for evaluating real-world geo-temporal reasoning in VLMs. TimeSpot comprises 1,455 images spanning 80 countries, where models must infer temporal attributes (season, month, time of day, daylight phase) and geolocation attributes (continent, country, climate zone, environment type, latitude–longitude coordinates) directly from the visual input. In addition, it includes spatial reasoning tasks that require integrating geographical, spatial, and temporal cues to solve complex understanding problems. Unlike prior benchmarks that emphasize obvious cues or iconic imagery, TimeSpot prioritizes diverse and subtle settings, reflecting the difficulty of reasoning under real-world uncertainty. Our evaluation of state-of-the-art VLMs, including both open- and closed-source models, reveals consistently low performance across tasks, highlighting substantial challenges in achieving robust temporal and geographic reasoning. These findings underscore the pressing need for improved methods to enable reliable and trustworthy geo-temporal understanding in VLMs, paving the way for future research in this critical domain.
@inproceedings{ridoy2026timespot, title = {{TimeSpot}: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings}, author = {Wasi, Azmine Toushik and Ridoy, Shahriyar Zaman and Tonmoy, Koushik Ahamed and Tshering, Kinga and Hasan, S. M. Muhtasimul and Faisal, Wahid and Mohiuddin, Tasnim and Parvez, Md. Rizwan}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)}, year = {2026}, note = {Main track. Co-first author.}, url = {https://arxiv.org/abs/2603.06687} } -
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
In Proceedings of the 14th International Conference on Learning Representations (ICLR), 2026
Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision–language models (VLMs). Prior work largely relied on synthetic or LLM-generated environments with limited task designs and puzzle-like setups, failing to capture the real-world complexity, visual noise, and diverse spatial relationships that VLMs encounter. To address this, we introduce SpatiaLab, a comprehensive benchmark for evaluating VLMs' spatial reasoning in realistic, unconstrained contexts. SpatiaLab comprises 1,400 visual question–answer pairs across six major categories: Relative Positioning, Depth & Occlusion, Orientation, Size & Scale, Spatial Navigation, and 3D Geometry, each with five subcategories, yielding 30 distinct task types. Each subcategory contains at least 25 questions, and each main category includes at least 200 questions, supporting both multiple-choice and open-ended evaluation. Experiments across diverse state-of-the-art VLMs, including open- and closed-source models, reasoning-focused, and specialized spatial reasoning models, reveal a substantial gap in spatial reasoning capabilities compared with humans. In the multiple-choice setup, InternVL3.5-72B achieves 54.93% accuracy versus 87.57% for humans. In the open-ended setting, all models show a performance drop of around 10–25%, with GPT-5-mini scoring highest at 40.93% versus 64.93% for humans. These results highlight key limitations in handling complex spatial relationships, depth perception, navigation, and 3D geometry. By providing a diverse, real-world evaluation framework, SpatiaLab exposes critical challenges and opportunities for advancing VLMs' spatial reasoning, offering a benchmark to guide future research toward robust, human-aligned spatial understanding.
@inproceedings{ridoy2026spatialab, title = {{SpatiaLab}: Can Vision–Language Models Perform Spatial Reasoning in the Wild?}, author = {Wasi, Azmine Toushik and Faisal, Wahid and Rahman, Abdur and Anik, Mahfuz Ahmed and Shahriar, Munem and Topu, Mohsin Mahmud and Meem, Sadia Tasnim and Priti, Rahatun Nesa and Mitu, Sabrina Afroz and Hoque, Md. Iqramul and Ridoy, Shahriyar Zaman and Ali, Mohammed Eunus and Hawsaly, Majd and Raza, Mohammad and Parvez, Md. Rizwan}, booktitle = {Proceedings of the 14th International Conference on Learning Representations (ICLR)}, year = {2026}, url = {https://arxiv.org/abs/2602.03916} } -
SEISMOS: A Statistical Signal Detection Framework for Semantic Chunking
In Advances in Neural Information Processing Systems (NeurIPS), 2026
Main conference (poster).
Chunking is a hidden bottleneck in dense retrieval and retrieval-augmented generation: it determines the units that can be indexed, retrieved, and ultimately used as evidence. Yet most systems still segment documents using fixed token windows or globally thresholded semantic similarity drops, treating chunking as an engineering heuristic rather than a statistical decision. We propose SEISMOS, a statistical signal detection framework for semantic chunking. SEISMOS models consecutive sentence-embedding cosine similarities as a document-level signal and derives boundary decisions from the null hypothesis of no semantic transition. Across four development BEIR corpora, we establish three corpus-invariant properties: bimodal document structure, rapid survival decay of shallow local minima, and positive lag-1 autocorrelation. These properties show that semantic boundaries should be detected by variance-normalized deviations, not by document means or global similarity thresholds. Accounting for autocorrelation leads to a normalized discrete Laplacian detector that identifies significant semantic valleys through a single interpretable decision rule. Evaluated on five BEIR benchmarks, SEISMOS consistently improves over fixed-length, recursive, and production semantic chunking baselines under one fixed operating configuration. Without corpus-specific retuning, the same detector transfers to a held-out TREC-COVID corpus. Our results suggest that the boundaries needed for effective retrieval are already encoded in embedding-similarity signals, and that principled, efficient, LLM-free chunking can be obtained by detecting them statistically.
@inproceedings{saad2026seismos, title = {{SEISMOS}: A Statistical Signal Detection Framework for Semantic Chunking}, author = {Saad, Ishtiak Mahmud and Islam, Mominul and Ridoy, Shahriyar Zaman and Ahsan, Md Manjurul and Wasi, Azmine Toushik}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026} } -
From Language Specifications to Executable Turing Machines: Evaluating LLMs as Computational Machine Designers
In Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
First author.
@inproceedings{ridoy2026tmeval, title = {From Language Specifications to Executable Turing Machines: Evaluating {LLMs} as Computational Machine Designers}, author = {Ridoy, Shahriyar Zaman and Hasan, S. M. Muhtasimul and Wasi, Azmine Toushik and Tshering, Kinga and Tonmoy, Koushik Ahamed and Khan, Md Mosaddek}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, year = {2026}, note = {First author.} } -
Can LLMs Design Computational Machines? Pushdown Automaton Synthesis as a Test of Structured Computational Reasoning
In Findings of the Association for Computational Linguistics: EMNLP 2026, 2026
@inproceedings{mahmud2026pda, title = {Can {LLMs} Design Computational Machines? Pushdown Automaton Synthesis as a Test of Structured Computational Reasoning}, author = {Mahmud, Syed Mumtahin and Lina, Nazira Jesmin and Ridoy, Shahriyar Zaman and Hasan, S. M. Muhtasimul and Khan, Md Mosaddek}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, year = {2026} } -
BengaliMoralBench: A Benchmark for Auditing Ethical Reasoning in Large Language Models in Bengali Language and Culture
In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2026
Co-first author.
As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, which is spoken by over 285 million people and ranked 6th globally, remains underexplored. Existing ethics benchmarks are largely English-centric and shaped by Western frameworks, overlooking cultural nuances critical for real-world deployment. To address this, we introduce BengaliMoralBench, the first large-scale ethics benchmark for the Bengali language and socio-cultural contexts. It covers five moral domains, Daily Activities, Habits, Parenting, Family Relationships, and Religious Activities, subdivided into 50 culturally relevant subtopics. Each scenario is annotated via native-speaker consensus using three ethical lenses: Virtue, Commonsense, and Justice ethics. We conduct systematic zero-shot evaluation of prominent multilingual LLMs, including Llama, Gemma, Qwen, and DeepSeek, using a unified prompting protocol and standard metrics. Performance varies widely (50–91% accuracy), with qualitative analysis revealing consistent weaknesses in cultural grounding, commonsense reasoning, and moral fairness. BengaliMoralBench provides a foundation for responsible localization, enabling culturally aligned evaluation and supporting the deployment of ethically robust AI in diverse, low-resource multilingual settings such as Bangladesh.
@inproceedings{ridoy2026bengalimoralbench, title = {{BengaliMoralBench}: A Benchmark for Auditing Ethical Reasoning in Large Language Models in Bengali Language and Culture}, author = {Ridoy, Shahriyar Zaman and Wasi, Azmine Toushik and Tonmoy, Koushik Ahamed and Hasan Rafi, Taki and Chae, Dong-Kyu}, booktitle = {Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT)}, year = {2026}, note = {Co-first author.}, url = {https://arxiv.org/abs/2511.03180} } -
Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising
In Proceedings of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL-IJCNLP), 2026
Main conference.
-
Geo-spatial and Geo-temporal Reasoning in Vision–Language and Large Language Models: A Review
In Proceedings of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL-IJCNLP), 2026
Main conference. Co-first author.
-
Position: Real-World Clinical NLP Robustness Requires Messy, Multimodal, Longitudinal, Privacy-Preserving Corpora
In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2026, 2026
Findings. Co-first author. Also presented at the EurIPS 2025 Workshop on Multimodal Representation Learning for Healthcare (MMRL4H).
Real-world healthcare data is inherently complex, riddled with noise, incompleteness, heterogeneity, and temporal irregularities, making the development of clinically robust AI systems particularly challenging. Traditional AI models in healthcare are often developed on idealized, clean datasets that fail to reflect the operational realities of clinical settings. This limits their generalizability and effectiveness when deployed in real-world environments. There is a critical need for a paradigm shift toward methodologies that embrace the "messiness" of clinical data while ensuring patient privacy and ethical integrity. Current models and pipelines are ill-equipped to manage this complexity in a scalable, trustworthy manner. We introduce the concept of the "Messy Clinic" dataset and a corresponding blueprint designed to guide the development of healthcare AI systems grounded in real-world data characteristics. This framework supports seamless multimodal data integration, spanning EHRs, imaging, genomics, and wearables, while incorporating advanced privacy-preserving techniques such as federated learning, synthetic data generation, and differential privacy. The blueprint offers (1) a structured methodology for longitudinal, multimodal data management; (2) robust AI/ML techniques tailored to noisy and incomplete inputs; and (3) a comprehensive governance framework addressing ethical concerns like consent, accountability, and bias mitigation. Embracing the "Messy Clinic" paradigm represents a foundational shift in healthcare AI development, from artificial idealism to real-world robustness. This approach promises to accelerate medical research, enable personalized care, and ultimately transform healthcare delivery by aligning AI systems with the authentic complexity of clinical practice.
@inproceedings{wasi2025position_clinical_nlp, title = {Position: Real-World Clinical {NLP} Robustness Requires Messy, Multimodal, Longitudinal, Privacy-Preserving Corpora}, author = {Wasi, Azmine Toushik and Ridoy, Shahriyar Zaman}, booktitle = {Findings of the Association for Computational Linguistics: AACL-IJCNLP 2026}, year = {2026}, note = {Co-first author. Also presented at the EurIPS 2025 Workshop on MMRL4H.} } -
ICML Workshop
When Machines Decide If a Human Wrote It: Creativity in the Age of AI Detectors
In ICML 2026 Workshop on Human-AI Co-Creativity, 2026
AI detectors, originally developed to flag machine-generated content, are now widely used to evaluate human writing in educational and creative contexts. By relying on rigid linguistic metrics and statistical norms, these systems subtly reshape human expression through algorithmic conformity. Acting as gatekeepers of authentic expression, they privilege conformity over originality and disproportionately penalize non-native English speakers and marginalized voices, causing psychological harm, academic penalties, and cultural erasure. This paper argues that such systems impose an algorithmic aesthetic that suppresses rebellion, hinders discovery, and diminishes the joy of creation. Reframing the issue as one of civil rights and human flourishing, we propose three interventions: restricting AI detectors in creative and educational spaces, promoting glitch aesthetics to valorize imperfection, and protecting creative anonymity to foster free experimentation. Drawing on legal and cultural policy precedents, we contend that these measures are essential to safeguarding human imagination. Ultimately, we advocate for institutional safeguards that champion risk, surprise, and dissent as vital components of a thriving creative future.
@inproceedings{azmine2026machines, title = {When Machines Decide If a Human Wrote It: Creativity in the Age of {AI} Detectors}, author = {Wasi, Azmine Toushik and Ridoy, Shahriyar Zaman and Ahsan, Md Manjurul}, booktitle = {ICML 2026 Workshop on Human-AI Co-Creativity}, year = {2026} } -
NeurIPS Workshop
Compliance–Correctness Decomposition for Post-Training Evaluation of Small Language Models
In NeurIPS 2026 Muslims in ML (MusIML) Workshop, 2026
@inproceedings{tshering2026compliance, title = {Compliance--Correctness Decomposition for Post-Training Evaluation of Small Language Models}, author = {Tshering, Kinga and Das, Sumon and Ridoy, Shahriyar Zaman and Alim, Md. Samiul and Wasi, Azmine Toushik and Chae, Dong-Kyu and Moni, Mohammad Ali}, booktitle = {NeurIPS 2026 Muslims in ML (MusIML) Workshop}, year = {2026} } -
NeurIPS Workshop
Harness Recession: When Small Models Outgrow Their Scaffolding
In NeurIPS 2026 SLMs for Agentic Systems (SLM-Agents) Workshop, 2026
@inproceedings{alim2026harness, title = {Harness Recession: When Small Models Outgrow Their Scaffolding}, author = {Alim, Md. Samiul and Tamim, Mahir Shahriar and Ridoy, Shahriyar Zaman and Khan, Sharjil and Wasi, Azmine Toushik}, booktitle = {NeurIPS 2026 SLMs for Agentic Systems (SLM-Agents) Workshop}, year = {2026} } -
Under Review
Retrieval Geometry Shapes Cache-Based CLIP Adaptation
2026
Under review at ICLR 2027.
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, ImageNet-A cache gain ranges from at most +0.44 points for CLIP and MAE to +19.7 ± 0.4 for DINOv2-L, while label-free retrieval-space selection retains 98% of oracle gain on ImageNet-V2. These results show that memory quality depends not only on which examples are stored, but also on how they are retrieved. Motivated by this finding, we propose MARC (Memory Augmented Retrieval for CLIP), a training-free system that uses frozen CLIP for prediction and DINOv2-B for retrieval with a single fusion weight. A single-view cache repairs 1074 ± 21 baseline errors, compared with 878 ± 4 for a 64-view ensemble, at roughly one seventh of the cost. Across four ImageNet distribution shifts, MARC reaches a 67.91% OOD average and, at matched DINOv2-B scale and eight views, achieves 64.17 ± 0.31% versus 62.75 ± 0.15% for a graph-based cache system while running 2.6 times faster. Overall, our results establish retrieval space as a first-order design choice for robust cache-based adaptation in remote sensing, scientific imaging, and changing visual environments.
@article{tamim2026marc, title = {Retrieval Geometry Shapes Cache-Based {CLIP} Adaptation}, author = {Tamim, Mahir Shahriar and Alim, Md. Samiul and Wasi, Azmine Toushik and Ridoy, Shahriyar Zaman and Nesa, Meharun and Yousuf, Mohammad Abu and Lamb, Alex and Moni, Mohammad Ali}, journal = {arXiv preprint arXiv:2609.23409}, year = {2026}, url = {https://arxiv.org/abs/2609.23409} } -
Under Review
Beyond Build Validity: Evaluating Behavioral Correctness in LLM-Generated Context-Free Grammars
2026
Under review at ACL Rolling Review (May 2026). Co-first author.
-
Under Review
InvariantBench: Can Large Language Models Exhibit Inherent Reasoning Across Equivalent Transformations?
2026
Under review at ICLR 2027. A workshop version appeared at the AI for Math Workshop at ICML 2026.
@inproceedings{wasi2026invariantbench, title = {{InvariantBench}: Can Large Language Models Exhibit Inherent Reasoning Across Equivalent Transformations?}, author = {Wasi, Azmine Toushik and Khan, Mahir Absar and Rahman, Abdur and Faisal, Wahid and Saha, Sukanta and Bhuyan, Saimon and Sukanya, Mahdiya Rahman and Priti, Rahatun Nesa and Hoque, Md. Iqramul and Shahriar, Munem and Islam, Raima and Sushil, Sanatan and Ridoy, Shahriyar Zaman and Sultan, Kazi Raiwan and Islam, MD Shafikul and Ahsan, Md Manjurul and Parvez, Md Rizwan}, booktitle = {AI for Math Workshop at the International Conference on Machine Learning (ICML)}, year = {2026}, note = {Workshop version. Extended version under review at ICLR 2027.} }
2025
-
CAPSTONE: Composable Attribute-Prompted Scene Translation for Zero-Shot Vision–Language Reasoning
In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, 2025
Co-first author.
Interpreting visual scenes with high-level reasoning is essential for many real-world applications, such as autonomous systems and content moderation, but training and maintaining Vision–Language Models (VLMs) remains resource-intensive and opaque. In this work, we present CAPSTONE, a lightweight, modular framework designed for industrial settings. Instead of relying on multimodal training or fine-tuning large models, CAPSTONE transforms outputs from off-the-shelf vision models into structured text prompts that can be interpreted by a frozen Large Language Model (LLM). This plug-and-play architecture enables reasoning over visual input without access to raw pixels, dramatically reducing computational cost and complexity. On the POPE dataset, our system, using a 7B LLM, outperforms several fully trained VLMs in zero-shot evaluations, while on the VSR benchmark, the 4B model achieves competitive results, together demonstrating strong generalization without retraining. CAPSTONE offers a scalable and interpretable alternative for companies looking to integrate visual reasoning capabilities without the burden of full-scale VLM pipelines.
@inproceedings{hossain2025capstone, title = {{CAPSTONE}: Composable Attribute-Prompted Scene Translation for Zero-Shot Vision–Language Reasoning}, author = {Hossain, Md. Ismail and Ridoy, Shahriyar Zaman and Farazi, Moshiur and Mohammed, Nabeel and Rahman, Shafin}, booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track}, year = {2025}, note = {Co-first author.} } -
Context-Aware Data Cleaning: Optimizing Bengali Text for Contextual Text Classification
SN Computer Science, 6(5), 422, 2025
In Natural Language Processing (NLP), textual data is foundational, yet it presents substantial challenges, especially for under-resourced languages like Bengali. The complexity and volume of Bengali textual data require sophisticated data cleaning techniques. Traditional methods often neglect critical contextual information essential for effective textual analysis. This study highlights the need for context-aware data cleaning, a methodology that maintains linguistic context while removing noise. The study compares context-aware and traditional data cleaning approaches tailored for Bengali text to improve the performance of contextual transformer-based models. Conventional techniques in this study include symbol and punctuation removal, stop-word elimination, stemming, and removing HTML tags or URLs. In contrast, context-aware techniques involve spelling correction, tagging HTML and URLs, preserving punctuation and emojis, and selectively removing less important words using TF-IDF. The current initiative assesses the impact of these strategies through rigorous dataset curation and extensive training in machine learning, deep learning, and transformer-based models on four prominent Bengali datasets: BEmoC, SentNoB, UBMEC, and EmoNoBa. Results show context-aware data cleaning significantly outperforms traditional methods, particularly in enhancing transformer-based model performance. The developed context-aware data cleaning pipeline integrates various techniques, achieving a baseline accuracy improvement of up to 4% across three of the four datasets. These findings underscore the importance of preserving sentence-level context in Bengali for optimal NLP performance while minimizing noise. Additionally, the research introduces a novel context-aware data cleaning pipeline and provides detailed algorithms for its implementation, advancing NLP research and applications in Bengali and similar linguistic contexts.
@article{faisal2025context, title = {Context-Aware Data Cleaning: Optimizing Bengali Text for Contextual Text Classification}, author = {Faisal, Moshiur Rahman and Fahad, Abdur Rahman and Ridoy, Shahriyar Zaman and Sultana, Jannat and Ria, Zinnat Fowzia and Rahman, Md. Hasibur and Uddin, Mohammed Arif and Rahman, Rashedur M.}, journal = {SN Computer Science}, volume = {6}, number = {5}, pages = {422}, publisher = {Springer Nature}, year = {2025}, doi = {10.1007/s42979-025-03891-9} } -
Under Review
PANCHINI: Predictive Analytics for Child Marriage in Bangladesh Using Machine Learning Insights
2025
Under review at Array (Elsevier).
Child marriage remains a significant social issue, particularly in developing countries such as Bangladesh, where deep-rooted cultural traditions and socio-economic disparities perpetuate the practice despite legal restrictions and awareness campaigns. This study explores the underlying factors contributing to child marriage among Bangladeshi women, especially in rural and underdeveloped regions where traditional intervention strategies have struggled. By leveraging machine learning (ML) and natural language processing (NLP), we develop a predictive model to assess the likelihood of child marriage from individual demographic profiles. Our analysis focuses on key socio-economic, educational, and cultural factors, offering insights into regional variations across Bangladesh. We construct a novel dataset that combines macro- and micro-level socio-economic variables tailored to this problem. Model performance is evaluated across multiple algorithms, with ensemble Blending, particularly the EMNet configuration, achieving the strongest results, including accuracy (94.80%) and AUC (97.58%). We also evaluate large language models (LLMs) as prompt-based classifiers on relevant textual fields and as assistants for schema normalization and feature enrichment, benchmarking their outputs against supervised ML baselines. This data- and language-driven approach improves precision in identifying high-risk individuals and communities and yields actionable insights for policymakers. By integrating predictive analytics and LLM-based evaluation into intervention design, the work supports more targeted, effective policies to mitigate child marriage and promote long-term social empowerment.
2024
-
EnStack: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source Code
In Proceedings of the IEEE International Conference on Big Data (IEEE BigData), 2024
First author.
Automated detection of software vulnerabilities is critical for enhancing security, yet existing methods often struggle with the complexity and diversity of modern codebases. In this paper, we introduce EnStack, a novel ensemble stacking framework that enhances vulnerability detection using natural language processing (NLP) techniques. Our approach synergizes multiple pre-trained large language models (LLMs) specialized in code understanding: CodeBERT for semantic analysis, GraphCodeBERT for structural representation, and UniXcoder for cross-modal capabilities. By fine-tuning these models on the Draper VDISC dataset and integrating their outputs through meta-classifiers such as Logistic Regression, Support Vector Machines (SVM), Random Forest, and XGBoost, EnStack effectively captures intricate code patterns and vulnerabilities that individual models may overlook. The meta-classifiers consolidate the strengths of each LLM, resulting in a comprehensive model that excels in detecting subtle and complex vulnerabilities across diverse programming contexts. Experimental results demonstrate that EnStack significantly outperforms existing methods, achieving notable improvements in accuracy, precision, recall and F1-score. This work highlights the potential of ensemble LLM approaches in code analysis tasks and offers valuable insights into applying NLP techniques for advancing automated vulnerability detection.
@inproceedings{ridoy2024enstack, title = {{EnStack}: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source Code}, author = {Ridoy, Shahriyar Zaman and Shaon, Md. Shazzad Hossain and Cuzzocrea, Alfredo and Akter, Mst Shapna}, booktitle = {Proceedings of the IEEE International Conference on Big Data (IEEE BigData)}, pages = {6356--6364}, address = {Washington, DC, USA}, year = {2024}, organization = {IEEE}, note = {First author.}, url = {https://arxiv.org/abs/2411.16561} } -
IEEE IS
An Efficient Text Cleaning Pipeline for Clinical Text for Transformer Encoder Models
In Proceedings of the IEEE 12th International Conference on Intelligent Systems (IS), 2024
First author.
Best Paper Award at the IEEE 12th International Conference on Intelligent Systems (IS) 2024, Varna, Bulgaria. View certificate
It might be challenging to choose the best text preprocessing strategy in the field of natural language processing (NLP) due to the variety of techniques available. Given the popularity of transformer models, we wondered if preprocessing was necessary and, if so, what methods would improve the models' performance. Especially when working with clinical text data, accuracy is crucial. Our goal was to find an appropriate preprocessing pipeline for clinical texts that maintains or improves model performance. We experienced four common preprocessing techniques and their groupings on two datasets from MIMIC-3 and PubMed. We used four models: BERT base, BioBERT, BioClinicalBERT, and RoBERTa. The varied accuracy results from existing techniques inspired us to develop a new pipeline to improve accuracy. Our pipeline starts with removing repeated punctuation, normalizing the text with a CleanText function, and filtering less important words using TF-IDF scores to keep clinically applicable terms and moderate noise. Our results presented that our pipeline outperformed the base models. For the MIMIC-3 dataset, the BERT base model achieved 90.16% accuracy, and for the PubMed dataset, BioBERT achieved 64.20% accuracy. We also found that removing stop words decreased accuracy, while using TF-IDF either maintained or improved it up to 3%. Additionally, as we removed less important words from the documents our pipeline considerably reduced training time up to 17%.
@inproceedings{ridoy2024efficient, title = {An Efficient Text Cleaning Pipeline for Clinical Text for Transformer Encoder Models}, author = {Ridoy, Shahriyar Zaman and Sultana, Jannat and Ria, Zinnat Fowzia and Uddin, Mohammed Arif and Rahman, Md. Hasibur and Rahman, Rashedur M.}, booktitle = {Proceedings of the IEEE 12th International Conference on Intelligent Systems (IS)}, pages = {1--9}, address = {Varna, Bulgaria}, year = {2024}, organization = {IEEE}, note = {First author.} }
No publications match that filter.