ML systems engineering
Scalable Python pipelines, automated data quality, profiling, reproducibility, and deployment-oriented GPU inference.
Mohammed Saidul Islam · Vector Institute
I design scalable data pipelines, automated evaluation systems, multimodal workflows, and efficient model-inference infrastructure.
Expertise
I am an Associate Applied Machine Learning Specialist at the Vector Institute, working across modular Python systems, large-scale preprocessing, agentic LLM/VLM workflows, post-training, evaluation, and reproducible inference on Linux and HPC.
My research interests include multimodal reasoning, visualization intelligence, trustworthy model evaluation, and agentic AI. I am currently exploring mechanistic interpretability in vision-language models.
Scalable Python pipelines, automated data quality, profiling, reproducibility, and deployment-oriented GPU inference.
LLM/VLM workflows, multimodal data processing, planning and reflection, post-training, and tool-augmented generation.
Benchmark construction, robustness testing, model-as-judge protocols, error analysis, and failure-mode tracking.
Selected work
Applied ML and research systems spanning foundation-model evaluation, multimodal agents, and visualization intelligence.
A reference-grounded multi-agent pipeline for generating, verifying, repairing, and deduplicating technically demanding evaluation tasks.
Trace-aware quality control from source ingestion through final benchmark validation.
A GRPO post-training framework using post-execution feedback across textual correctness, code executability, and visualization quality.
Improves executable code generation and rendered chart quality over strong baselines.
A benchmark for multimodal agents that must ground questions, plan interactions, operate real dashboards, and reason across views.
Surfaces practical failures in grounding, planning, interaction, and visual reasoning.
Career
A path from teaching software fundamentals to building applied multimodal ML systems.
Sep 2025 — present
Associate Applied Machine Learning Specialist · Toronto
Sep 2023 — Aug 2025
Graduate Research Assistant · Toronto
Jul 2021 — Aug 2023
Lecturer · Bangladesh
Milestones
Selected career and publication milestones.
| May 2026 | Released the preprint Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models, describing grounded task generation with multi-agent design and verification. |
|---|---|
| Mar 2026 | Two papers appeared at EACL 2026: RL-Text2Vis in the main proceedings and DashboardQA in Findings. |
| Nov 2025 | New EMNLP 2025 work on geo-economic bias in chart-to-text and deploying tiny LVLM judges. |
Contact
Email is the best way to reach me. You can also explore my work on GitHub and Google Scholar.