Applied machine learning

Building reliable AI agents, multimodal systems, and model evaluation infrastructure.

At the Vector Institute, I turn applied ML research into production-ready systems, from benchmark design and model training to deployment and evaluation.

  • Efficient Inference
  • LLM Evaluation
  • AI Agents
  • Post-Training
  • Multimodal AI
Portrait of Mohammed Saidul Islam

Latest EMNLP 2026 Main

DSAgentBench

Can agents automate end-to-end data-science workflows in real computer environments?

275 tasks 15 models deterministic evaluators 56.70% best result

Selected work

Projects across agent evaluation, multimodal learning, post-training, and visualization.

EMNLP 2026 Main

DSAgentBench

A benchmark for evaluating whether AI agents can complete end-to-end data-science workflows in real computer environments using notebooks, IDEs, terminals, browsers, and databases.

Evidence: 275 diverse tasks across the data-science lifecycle, 15 evaluated models, and deterministic evaluators. The best reported agent reached 56.70% task success.

  • Agents
  • Computer Use
  • Evaluation
  • Data Science
  • Tool Orchestration

Vector Institute · Foundation-model evaluation · 2026 preprint

Fine-Grained Benchmark Generation

A reference-grounded multi-agent pipeline for generating, verifying, repairing, and deduplicating technically demanding evaluation tasks.

Contribution: Built OCR ingestion and a multi-agent designer and verifier workflow with iterative repair, trace-aware validation, and semantic deduplication.

Evidence: Trace-aware quality control connects source ingestion to final benchmark validation.

  • OCR
  • Multi-agent Systems
  • Benchmark Generation
  • Evaluation
  • Deduplication

MICCAI 2026 · Multimodal medical AI

Open-PMC-18M

An open multimodal dataset and scalable pipeline linking biomedical figures with captions and subcaptions for medical vision-language learning.

Contribution: Owned preprocessing and batched vLLM inference for Qwen2.5-VL and Qwen2.5 summary generation.

Evidence: Scales compound-figure extraction, summarization, and OpenCLIP training across 18 million biomedical image-text pairs with optimized batched VLM inference.

  • VLMs
  • vLLM
  • Data Pipelines
  • Medical AI
  • HPC

EACL 2026 Main · Post-training

RL-Text2Vis

A GRPO framework that evaluates text-to-visualization outputs after code execution using textual, executable, and visual feedback.

Contribution: Co-developed post-execution multimodal feedback for aligning generated text, code, and rendered visualizations.

Evidence: Improves code executability and chart quality over strong prompting and supervised baselines.

  • GRPO
  • Post-training
  • Code Generation
  • Visualization
  • VLM Evaluation

EACL 2026 Findings · Multimodal agents

DashboardQA

A benchmark for agents that answer questions through real interactive dashboards.

Contribution: Co-created the benchmark to expose failures in grounding, planning, GUI interaction, and multi-view reasoning.

Evidence: Tests filters, tabs, view switching, visual grounding, and cross-view reasoning in realistic dashboard environments.

  • Multimodal Agents
  • GUI Interaction
  • Benchmarks
  • Planning
  • Visual Reasoning

ACL 2025 Findings · Benchmarking

ChartQAPro

A challenging chart question-answering benchmark built around diverse real-world visualizations.

Contribution: Co-created the benchmark and evaluation protocol for open and closed vision-language models.

Evidence: 1,341 charts and 1,948 questions evaluated 21 VLMs.

  • Chart QA
  • VLMs
  • Benchmarking
  • Multimodal Reasoning

Research

Selected publications

Recent work on AI agents, multimodal evaluation, data storytelling, and reliable generation.

  1. EMNLP Accepted

    DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Mizanur Rahman , Mohammed Saidul Islam , Ridwan Mahbub , Md Tahmid Rahman Laskar , Shafiq Joty and Enamul Hoque
    arXiv preprint arXiv:2608.103662026

    EMNLP 2026 Main

  1. EACL Conference

    Aligning Text, Code, and Vision: A Multi-Objective Reinforcement Learning Framework for Text-to-Visualization

    Mizanur Rahman , Mohammed Saidul Islam , Md Tahmid Rahman Laskar , Shafiq Joty and Enamul Hoque
    Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics2026
  1. EACL Findings

    DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

    Aaryaman Kartha , Ahmed Masry , Mohammed Saidul Islam , Thinh Lang , Shadikur Rahman , Ridwan Mahbub , Mizanur Rahman , Mahir Ahmed , Md Rizwan Parvez , Enamul Hoque and Shafiq Joty
    Findings of the Association for Computational Linguistics: EACL 20262026
  1. ACL Findings

    ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

    Ahmed Masry , Mohammed Saidul Islam , Mahir Ahmed , Aayush Bajaj , Firoz Kabir , Aaryaman Kartha , Md Tahmid Rahman Laskar , Mizanur Rahman , Shadikur Rahman , Mehrad Shahmohammadi , Megh Thakkar , Md Rizwan Parvez , Enamul Hoque and Shafiq Joty
    Findings of the Association for Computational Linguistics: ACL 20252025
  1. EMNLP Conference

    DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts

    Mohammed Saidul Islam , Md Tahmid Rahman Laskar , Md Rizwan Parvez , Enamul Hoque and Shafiq Joty
    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing2024
  1. EMNLP Industry

    Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices

    Md Tahmid Rahman Laskar , Mohammed Saidul Islam , Ridwan Mahbub , Mizanur Rahman , Amran Bhuiyan , Israt Jahan , Mir Tafseer Nayeem , Shafiq Joty , Enamul Hoque and Jimmy Xiangji Huang
    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track2025

Career

Experience and education

Roles in applied ML research, engineering, and teaching.

2025-Present

Vector Institute

Associate Applied Machine Learning Specialist · Toronto

2023-2025

Intelligent Visualization Lab, York University

Graduate Research Assistant · Toronto

2021-2023

Islamic University of Technology

Lecturer · Bangladesh

MSc Computer Science - York UniversityBSc Computer Science and Engineering - Islamic University of Technology

Now

Recent milestones

Recent publication and career updates.

Aug 2026 OpenPMC-18M was accepted at MICCAI 2026.
Aug 2026 DSAgentBench, a benchmark for end-to-end data-science workflows in real computer environments, was accepted to EMNLP 2026 Main.
May 2026 Released the preprint Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models, describing grounded task generation with multi-agent design and verification.

Service

Leadership and reviewing

President of the Bangladeshi Graduate Student Association at York and reviewer for NeurIPS, ACL, EMNLP, COLM, NAACL, and ACL Rolling Review.

Contact

Get in touch about applied AI, agents, evaluation, or multimodal systems.