My first calibration meeting was a shock. I had walked in with ratings I believed in, backed by a year of notes. I walked out with two of them lowered, not because the evidence was wrong but because I could not explain it in the ninety seconds I got before the next manager started talking.

Nobody had told me that calibration is a persuasion exercise with a hard time limit, and that the quality of the case matters as much as the quality of the work.

What the meeting is for

Calibration exists because managers grade differently. One manager's "exceeds" is another's "meets," and without a room where those get reconciled, the generous manager's team gets promoted faster for the same work. The room is there to make the scale mean the same thing everywhere.

That is a good purpose, and it is also why the room is adversarial in a way most managers are not prepared for. Everyone in it is defending their people, and the ratings at the top of the scale are scarce by design. Your job is to make sure the scarcity lands where it should.

The case is built all year

The rating I bring to calibration is the last step of a process that started in January. Every quarter I write, for each person, three or four lines: what they delivered, what it changed, and one thing that is holding them back. By review time I have a year of those and the rating writes itself.

More important, so does the argument. When another manager pushes back, I can say "she led the migration that cut the incident count by half, she did it with two people instead of the four we planned, and here is the design doc." That is a ninety second case. "She is great, everyone loves working with her" is not.

No surprises, in either direction

The rule I hold hardest is that nobody on my team should learn anything new from their written review. If the review says they are behind, they heard it in a 1:1 months ago and we have been working on it. If the review says they exceeded, they heard that too, and they know what the promotion case looks like.

This is harder than it sounds because it means having the uncomfortable conversation in March instead of December. It is also the only version of a review process that is fair. A rating someone could not have predicted is a rating they could not have changed.

What I argue for in the room

I argue for specifics over adjectives, every time. When someone describes an engineer as "a strong performer," I ask what they shipped. When someone is being marked down for "communication," I ask what happened. Half the time the answer moves the rating, in one direction or the other.

I argue against the halo. The engineer on the visible project gets remembered. The engineer who kept the platform stable so the visible project could ship gets forgotten, unless her manager brings the receipts. I bring the receipts.

I argue for consistency across the room, including against my own interest. If I am pushing for a rating for my engineer that I would challenge if another manager proposed it, I should not be pushing.

The written review

After calibration, the review itself is short. What they did, what it changed, what is next. I write it for the person, not for the file. I have read reviews that were three pages of hedged corporate language, and I have watched the engineer look for the one sentence that told them where they actually stood. Put that sentence first.

The year the scale moved

Reviewing performance in the year we moved to an AI first way of working was the hardest cycle I have run, because the definition of good output changed in the middle of the year. An engineer who wrote a lot of code by hand in the first half and a lot of generated code in the second half was not necessarily more productive. An engineer who shipped less code and caught more bad generated changes in review might have been the most valuable person on the team.

I went into calibration with that argument written down, because I knew every other manager would be measuring the old way. Some of them had not thought about it. The rating that mattered most that year was for an engineer whose visible output had dropped and whose judgment had become the thing the whole team relied on. I got the rating I asked for, because I had the examples, and because I had told her in July exactly what case I would be making.