LamaLab × M3RG · Benchmark
MaCBench
Can vision-language models assist in automated scientific discovery?
Task columns
Time axis
Show
Each dot is a model at its release date; the stepped line is the best overall
score achieved up to that point (state of the art). Hover a dot for details.
Best score per task over time. Click any task to expand.
Each model family's score over time. A red ▼ marks a release that scored lower than the family's previous one. Click any task below to expand.
Overall score per model family over time. Lines connect a family's successive releases; a red ▼ marks a release that scored lower than the family's previous one.
Diagnostic views grounded in the MaCBench paper: how models perform across the three pillars of the scientific workflow, where perception outruns reasoning, and how performance falls as tasks need more steps.
How MaCBench's 34 tasks are grouped for the analysis panels — derived from the paper's framework (pillars & categories) and results text (perception vs reasoning, step count).