Released the preprint Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models, describing grounded task generation with multi-agent design and verification.