Fund performance is almost never reported as a bare number. It is reported against a benchmark, and the benchmark chosen determines what the comparison actually tests.
A raw return answers no useful question
Knowing that a fund gained ground over a period says nothing about whether the manager added anything. Broad markets move, and a fund holding that market moves with it.
The benchmark separates the part of the result that came from being invested at all from the part that came from the specific decisions the fund made.
Without that separation, a rising market flatters every fund inside it and a falling one condemns every fund equally, regardless of how they were run.
The benchmark has to match the mandate
A fund that buys small domestic companies cannot be usefully compared to an index of large ones, because the two populations behave differently for reasons unrelated to skill.
Regulators and prospectuses require a fund to name a benchmark that reflects how it actually invests. The stated mandate and the stated index are supposed to describe the same territory.
Where they diverge, the comparison stops measuring management and starts measuring the gap between two different markets, which is not what anyone reading it assumes.
Index construction rules are not neutral
Two indexes covering the same nominal market can differ in how they weight holdings, how often they reconstitute, and where they draw the size boundary.
Those rules change the index's behavior. A weighting method that concentrates in the largest companies produces a different path from one that spreads holdings evenly.
So a fund can look better or worse purely because of construction choices in the yardstick, without anything changing in the portfolio being measured.
Tracking error describes the distance
The variability of the difference between a fund and its benchmark is reported as tracking error. It measures how tightly the fund follows the index rather than whether it beat it.
An index fund aims to keep that figure small, because its purpose is replication. An active fund necessarily runs a larger one, since deviating from the index is the entire point.
Reading the two together is more informative than reading either alone, because a small margin produced with large deviations is a different result from the same margin produced with small ones.
Benchmarks are selected before the fact for a reason
Naming the index in advance prevents a fund from choosing, after the period ends, whichever comparison happens to look most favorable.
Changing benchmarks is permitted when a strategy genuinely changes, but each change breaks the continuity of the record and makes long histories harder to interpret.
The fixed reference is what turns a performance number into a claim that can be checked, which is why the choice receives as much scrutiny as the result.