About
I am an AI researcher working on reasoning and reliability in language and vision-language models. Currently, I am a Research Intern at Vanderbilt University, a Research Assistant at Cognitive Agents and Interaction Lab (CAIL), University of Dhaka, working with Dr. Md Mosaddek Khan, and a Research Collaborator with Qatar Computing Research Institute (QCRI), working with Dr. Md Rizwan Parvez. I also collaborate with Prof. Alex Lamb (College of AI, Tsinghua University; formerly Microsoft Research).
My research spans three connected directions: (1) multimodal reasoning, including spatial, temporal, and compositional understanding; (2) formal and symbolic reasoning, drawing on automata theory to connect linguistic specification, symbolic representation, and executable computation, so that reasoning can be studied through behaviorally verifiable mechanisms; and (3) emotion understanding and trustworthy evaluation, grounded in psychological theories and human affective factors. I develop evaluation frameworks to assess emotional competence, psychological alignment, bias, and reliability. More broadly, I study when models reason and why they fail, verify their capabilities through behavioral evaluation, and develop mitigation strategies to improve their reliability.
My publications include work at NeurIPS, ICML, ICLR, EMNLP, ACM FAccT, and AACL-IJCNLP, among other venues. I serve as a reviewer for ICLR, WACV, the NeurIPS, and ICML, NeurIPS, and ACL workshops.
I graduated summa cum laude in CSE from North South University (CGPA 3.88/4.00, top 2%), and my thesis research received the Best Paper Award at IEEE IS 2024. Previously, I interned at the Artificial Intelligence Institute, University of South Carolina (AIISC), working with Dr. Amitava Das, Aman Chadha (Google DeepMind), and Vinija Jain (Google).
I am seeking PhD (or MS) positions starting Fall 2027 in reasoning, evaluation, and trustworthy multimodal NLP. If my work aligns with your group, I would be glad to hear from you at shahriyar.zaman01@gmail.com.
News and Updates
All news-
🎉 Our paper SEISMOS: A Statistical Signal Detection Framework for Semantic Chunking is accepted to the NeurIPS 2026 main conference (poster)!
-
Serving as a reviewer for ICLR 2027.
-
Invited to serve as a reviewer for three NeurIPS 2026 workshops: AI4MetaScience, AIWILD, and SLM-Agents.
-
Three of our papers on geo-spatial and geo-temporal reasoning, crisis sentiment analysis, and clinical NLP robustness were accepted to AACL-IJCNLP 2026 (2 Main, 1 Findings)—one as first author, one as second author, and one as a co-author.
-
Invited to serve as a reviewer for WACV 2027.
-
Two papers on formal and structured computational reasoning are accepted to Findings of EMNLP 2026, one as first author and one as co-author.
-
Attending ICML 2026 at COEX in Seoul, South Korea, to present our main-track paper, TimeSpot, along with several workshop papers.
-
Invited to serve as a reviewer for the NeurIPS 2026 Creative AI Track.
-
Three papers are accepted to ICML 2026 workshops: two at the AI for Math Workshop (AI4Math) and one at Muslims in ML (MusIML). One AI4Math paper also received a reviewer-level Best Paper nomination recommendation.
-
Serving as a reviewer for the AI for Math Workshop at ICML 2026.
-
When Machines Decide If a Human Wrote It: Creativity in the Age of AI Detectors is accepted to the ICML 2026 Workshop on Human-AI Co-Creativity.
-
TimeSpot is accepted to the ICML 2026 main track, with me as co-first author, in collaboration with QCRI and CIOL. Read more
-
BengaliMoralBench, the first large-scale ethics benchmark for the Bengali language and culture, is accepted to ACM FAccT 2026, in collaboration with CIOL and Hanyang University.
-
SpatiaLab is accepted to ICLR 2026, in collaboration with QCRI, Monash University, and CIOL. Read more
Selected Publications
All publications-
Findings of EMNLP 2026
From Language Specifications to Executable Turing Machines: Evaluating LLMs as Computational Machine Designers
Evaluates whether LLMs can translate natural-language specifications into executable Turing machines, judged by behavioral correctness on held-out inputs rather than surface plausibility.
-
ICML 2026
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings
A benchmark of 1,455 real-world images from 80 countries that asks vision-language models to infer season, time of day, climate zone, and coordinates from pixels alone. State-of-the-art VLMs perform consistently poorly across tasks.
-
NeurIPS 2026
SEISMOS: A Statistical Signal Detection Framework for Semantic Chunking
Reframes semantic chunking as statistical boundary detection over sentence-embedding similarity signals. A single LLM-free detector improves dense retrieval over fixed-length, recursive, and production chunkers on five BEIR benchmarks and transfers to TREC-COVID without retuning.
-
EMNLP 2025 Industry
CAPSTONE: Composable Attribute-Prompted Scene Translation for Zero-Shot Vision–Language Reasoning
Turns off-the-shelf vision model outputs into structured prompts for a frozen LLM, enabling zero-shot visual reasoning with no multimodal training. A 7B pipeline outperforms fully trained VLMs on POPE.
-
Findings of EMNLP 2026
Can LLMs Design Computational Machines? Pushdown Automaton Synthesis as a Test of Structured Computational Reasoning
Introduces pushdown automaton synthesis as a controlled test of structured symbolic reasoning, verifying generated automata through structural buildability checks and execution against labeled context-free language test cases. State-of-the-art LLMs struggle to construct reliably executable PDAs.
-
ICLR 2026
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
1,400 visual question-answer pairs across 30 spatial task types in unconstrained real-world scenes. The best model reaches 54.9% accuracy in the multiple-choice setting against 87.6% for humans.
-
ACM FAccT 2026
BengaliMoralBench: A Benchmark for Auditing Ethical Reasoning in Large Language Models in Bengali Language and Culture
The first large-scale ethics benchmark for the Bengali language and culture: five moral domains and 50 subtopics, annotated by native-speaker consensus through virtue, commonsense, and justice lenses.
* denotes equal contribution. See all 17 papers with abstracts and BibTeX.
Research Experience
Full experience- Research Intern, Vanderbilt University (Remote)2026 – Present
- Research Assistant, Cognitive Agents and Interaction Lab, University of Dhaka (Remote)Dec 2025 – Present
- Research Collaborator, Qatar Computing Research Institute, QCRI (Remote)Jun 2025 – Present
- Research Collaborator, Apurba-NSU R&D Lab, North South UniversityNov 2024 – Jun 2025
- Research Intern, Artificial Intelligence Institute, University of South Carolina, AIISC (Remote)Sep 2024 – Jul 2025
- Research Assistant (NLP), North South UniversityDec 2023 – Feb 2024
Technical Skills
- Programming
- Python, C++, C, Java, SQL
- ML frameworks
- PyTorch, TensorFlow / Keras, Hugging Face Transformers and Datasets, scikit-learn, NumPy, Pandas, Matplotlib
- Methods
- Benchmark design and large-scale evaluation of LLMs and vision-language models, zero-shot and prompt-based pipelines, knowledge distillation and model compression, ensemble stacking, hallucination detection, adversarial robustness, transformer fine-tuning, behavioral evaluation of generated formal systems (grammars, automata, Turing machines)
- Tooling
- Git, Docker, LaTeX
- Languages
- English, Bangla (native)
Academic Service
All servicePeer Review
- ICLR 2027
- WACV 2027
- NeurIPS 2026, Creative AI Track
- NeurIPS 2026 AI4MetaScience Workshop
- NeurIPS 2026 AIWILD Workshop
- NeurIPS 2026 SLM-Agents Workshop
- AI for Math Workshop, ICML 2026
- ACL 2025 Student Research Workshop
Teaching
- Lab Instructor, Department of Electrical and Computer Engineering, North South University Feb 2025 – Dec 2025
- Teaching Assistant, North South University Feb 2023 – Jan 2025
Talks and Community
- Presented TimeSpot and workshop papers at ICML 2026, Seoul Jul 2026
- Lab Instructor Coordinator for CSE 225L, standardizing labs across sections
- Mentored peers in data structures and algorithms with NSU Problem Solvers 2021 – 2023
I am always happy to talk about research, collaborations, and PhD opportunities. Email is the best way to reach me. Contact details.