The annual performance review is close to universal and close to universally disliked. Surveys of both managers and employees consistently find low confidence that it measures anything accurately.

The specific failure modes are well documented, which means they can be addressed. Most organisations don't.

Recency bias

The largest single problem. A review covering twelve months is written from memory, and memory is dominated by the recent.

The practical effect is that a strong final quarter substantially outweighs a mediocre first three, and a difficult month before the review outweighs eleven good ones.

Employees learn this quickly and behave accordingly, which is entirely rational and produces effort distributed by review timing rather than by need.

The fix is straightforward and requires discipline: contemporaneous notes throughout the year. A manager who writes a few lines after anything notable has evidence rather than impressions. Almost nobody does this, and it's the single highest-return practice available.

The halo effect

A general impression of somebody contaminating assessment of specific dimensions.

If a manager rates somebody highly overall, they'll tend to rate them highly on every criterion, including ones where performance is actually mixed. The ratings correlate far more strongly with each other than genuine independent assessment would produce.

This makes multi-dimensional rating scales largely decorative. They look rigorous and produce one judgement expressed several times.

Forced distribution

Where organisations require ratings to fit a predetermined curve — a fixed proportion in each band.

The rationale is preventing rating inflation, which is a genuine problem. The cost is that ratings become relative to a team rather than absolute, so an excellent performer in a strong team can receive a worse rating than a mediocre one in a weak team.

It also produces predictable behaviours: managers rotating who gets the low rating, avoiding hiring strong performers who'd compete for the top slots, and quiet negotiation between managers about whose people get which bands.

Several large organisations that used forced distribution for years have abandoned it, generally citing damage to collaboration.

The two-purpose problem

The most fundamental design flaw. Reviews typically serve two purposes at once: determining pay and supporting development.

These are incompatible in one conversation.

Development requires honesty about weaknesses, which requires psychological safety. Compensation makes the conversation adversarial — anything you admit may cost you money.

So employees defend rather than reflect, and managers soften feedback to avoid a difficult conversation about pay. Both purposes are served badly.

The fix is separation: conduct development conversations at a different time, in a different format, with no compensation implications. Organisations that do this report considerably more useful development discussions.

Self-assessment

Widely used and it introduces its own distortion.

Research on self-assessment consistently finds systematic differences in how confidently people rate themselves, and those differences correlate with characteristics unrelated to performance. Where self-assessment feeds into the final rating, it imports that variation directly.

It has value as a prompt — asking somebody what they're proud of and what they'd do differently produces useful material. As an input to a rating it's introducing bias.

What works better

Some practices with reasonable support behind them.

Frequent lightweight feedback. Specific, timely, close to the event. Feedback given six months later is nearly useless because the person can't reconstruct what they did.

Documented evidence rather than recalled impressions. The contemporaneous notes point again.

Focus on future behaviour rather than past rating. "Here's what would make the biggest difference next quarter" is actionable. "You were a three out of five" is not.

Calibration across managers. Discussion between managers about standards, before ratings are finalised, reduces the variation between individual managers' generosity. Different from forced distribution — it aligns interpretation of the scale rather than imposing a curve.

Separating pay from development. As above, and it's the change with the largest effect.

Whether to have them at all

Some organisations have abolished formal reviews and reported improvement. Others have tried and reinstated them, discovering that the review served functions nobody had articulated.

Those functions are real: a documented record for legal purposes, a defensible basis for pay decisions, a forced conversation that would otherwise never happen, and a consistency check across an organisation.

The honest position is that reviews are a poor instrument serving several necessary purposes, and abolishing them without replacing those functions creates different problems.

Which argues for fixing them rather than removing them — and the fixes are known, cheap, and mostly a matter of somebody deciding to do them.

Peer input, carefully

Multi-source feedback gets added to review processes regularly and it introduces problems that are worth anticipating.

Where peer input affects ratings and pay, it changes behaviour between colleagues in predictable ways. Reciprocal arrangements form. People become reluctant to challenge those who will assess them. Collaboration becomes political.

Where it is used developmentally, without feeding into ratings, it works considerably better. People give more honest input when nothing hangs on it, and the recipient can hear it without defending.

The distinction is the same one running through this whole subject: assessment and development are different activities, and mixing them degrades both. Almost every problem with review processes traces back to that single conflation, and almost every improvement starts with separating them.