EvalRouter now charts every benchmark run.
Two models can post the same score and fail completely different tasks. A single number cannot show that.
Each results page now opens with:
- the outcome of every task, model by model
- each score with its interval
- score against