UnifyBench helps people compare language models with an experimental Overall ranking built from sourced benchmark comparisons and a fixed reference panel. Customize optional capability weights to reflect your priorities; when performance data is missing, it stays unknown instead of being guessed.
Model comparisons can be difficult to interpret when benchmarks, scales, and priorities differ. We built UnifyBench to make the sources visible, keep unknowns honest, and let each person decide which capabilities matter most.