Latest EMNLP 2026 Main
DSAgentBench
Can agents automate end-to-end data-science workflows in real computer environments?
275 tasks 15 models deterministic evaluators 56.70% best result
Applied machine learning
At the Vector Institute, I turn applied ML research into production-ready systems, from benchmark design and model training to deployment and evaluation.
Latest EMNLP 2026 Main
Can agents automate end-to-end data-science workflows in real computer environments?
275 tasks 15 models deterministic evaluators 56.70% best result
Selected work
Projects across agent evaluation, multimodal learning, post-training, and visualization.
A benchmark for evaluating whether AI agents can complete end-to-end data-science workflows in real computer environments using notebooks, IDEs, terminals, browsers, and databases.
Evidence: 275 diverse tasks across the data-science lifecycle, 15 evaluated models, and deterministic evaluators. The best reported agent reached 56.70% task success.
A reference-grounded multi-agent pipeline for generating, verifying, repairing, and deduplicating technically demanding evaluation tasks.
Contribution: Built OCR ingestion and a multi-agent designer and verifier workflow with iterative repair, trace-aware validation, and semantic deduplication.
Evidence: Trace-aware quality control connects source ingestion to final benchmark validation.
An open multimodal dataset and scalable pipeline linking biomedical figures with captions and subcaptions for medical vision-language learning.
Contribution: Owned preprocessing and batched vLLM inference for Qwen2.5-VL and Qwen2.5 summary generation.
Evidence: Scales compound-figure extraction, summarization, and OpenCLIP training across 18 million biomedical image-text pairs with optimized batched VLM inference.
A GRPO framework that evaluates text-to-visualization outputs after code execution using textual, executable, and visual feedback.
Contribution: Co-developed post-execution multimodal feedback for aligning generated text, code, and rendered visualizations.
Evidence: Improves code executability and chart quality over strong prompting and supervised baselines.
A benchmark for agents that answer questions through real interactive dashboards.
Contribution: Co-created the benchmark to expose failures in grounding, planning, GUI interaction, and multi-view reasoning.
Evidence: Tests filters, tabs, view switching, visual grounding, and cross-view reasoning in realistic dashboard environments.
A challenging chart question-answering benchmark built around diverse real-world visualizations.
Contribution: Co-created the benchmark and evaluation protocol for open and closed vision-language models.
Evidence: 1,341 charts and 1,948 questions evaluated 21 VLMs.
Research
Recent work on AI agents, multimodal evaluation, data storytelling, and reliable generation.
EMNLP 2026 Main
Career
Roles in applied ML research, engineering, and teaching.
2025-Present
Associate Applied Machine Learning Specialist · Toronto
2023-2025
Graduate Research Assistant · Toronto
2021-2023
Lecturer · Bangladesh
Now
Recent publication and career updates.
| Aug 2026 | OpenPMC-18M was accepted at MICCAI 2026. |
|---|---|
| Aug 2026 | DSAgentBench, a benchmark for end-to-end data-science workflows in real computer environments, was accepted to EMNLP 2026 Main. |
| May 2026 | Released the preprint Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models, describing grounded task generation with multi-agent design and verification. |
Service
President of the Bangladeshi Graduate Student Association at York and reviewer for NeurIPS, ACL, EMNLP, COLM, NAACL, and ACL Rolling Review.