Why we don't tie OKR outcomes to performance reviews

Ask a team to set an ambitious goal, then make it clear their bonus depends on hitting it, and watch what happens next. The goal gets smaller. Not because anyone stopped caring, but because they just did the maths. An 80% chance of a comfortable target beats a 40% chance of a bold one, every time real money is on the line.

That’s the trap a lot of performance management systems build into OKRs without realising it. And it’s the trap we’ve deliberately avoided at Tability.

What matters Comfortable target Bold target
Odds of hitting it ~80% ~40%
Bonus if you hit it Full bonus Same full bonus
Expected bonus (odds × payout) Higher Lower
Actual business value if hit Modest Meaningful
What a comp-motivated rational person picks This one Rarely

The core problem: incentives beat intentions

OKRs (Objectives and Key Results) were built to do one thing: help teams set and track ambitious goals. Andy Grove designed them at Intel as a learning tool, a way to figure out what’s actually achievable by regularly aiming a bit past it. Google popularised a scoring model where hitting 70% of a key result counts as a solid result, because 100% every quarter usually just means the targets were never ambitious to begin with.

Then somewhere between Silicon Valley and the rest of the corporate world, a lot of companies bolted OKR scores directly onto performance reviews and bonus calculations. It seems logical on paper: you set goals, you measure progress, you reward the people who deliver. In practice, it breaks the entire mechanism OKRs are supposed to run on.

The problem isn’t that people become lazy. It’s that they become rational. If your key result score determines your rating, your bonus, or your next promotion cycle, you’re not going to gamble on a stretch goal. You’re going to negotiate the target down to something you’re confident you can hit, then quietly celebrate when you clear it with room to spare. That’s not a failure of effort. That’s the system working exactly as designed, just not the design anyone intended.

This isn’t a new failure mode either. Management by objectives, the direct predecessor to OKRs, ran into the same wall decades earlier: tie ratings to whether you hit the objective, and objective-setting quietly turns into an exercise in setting objectives you already know you can hit.

Betterworks bet on the opposite approach. We don’t think it holds up.

This isn’t a hypothetical debate. Betterworks, one of the bigger platforms in this space, builds its product around linking goals directly to performance ratings and compensation decisions. Their pitch is that connecting the two keeps people accountable and gives managers a defensible, data-backed way to differentiate performance.

We get the appeal. It’s tidy. One system, one dataset, one source of truth for “did you do the job.” But tidy isn’t the same as correct. Every study on goal-setting and incentive design points the same direction: attach hard consequences to a specific number, and people will manage the number, not the outcome. Sales teams sandbag pipeline forecasts. Engineers pad estimates. Managers set easy key results dressed up as ambitious ones. None of this is a character flaw. It’s what happens when you tell smart people that the honest, ambitious answer costs them money.

What OKRs are actually for

OKRs work best as a learning system, not a scorecard. The point of a 70% average score isn’t to measure how hard someone worked. It’s to tell you whether the team is calibrating its ambition correctly. Consistently hitting 100%? Targets are too soft. Consistently landing at 30%? Something’s broken in planning, resourcing, or both.

None of that signal survives contact with a bonus structure. Once a score determines pay, the score stops being a measurement tool and becomes a negotiation. And once it’s a negotiation, you’ve lost the thing that made OKRs useful in the first place: an honest read on how ambitious the team is willing to be.

So how do we actually separate the two?

Separating OKRs from performance reviews doesn’t mean OKRs have no consequences. It means the consequences live somewhere else. This is standard practice for any function running a real operating cadence. StratOps teams treat check-ins, scoring, and performance conversations as separate systems that feed each other, not one dial that’s asked to do both jobs at once.

  1. Keep two different conversations, on two different cadences. OKR check-ins happen weekly or fortnightly and are about progress, blockers, and confidence, not judgement. Performance reviews happen quarterly or annually and are about the person’s overall contribution, growth, and behaviour.
  2. Judge people on inputs a manager can actually observe, not the raw key result number. Did they take on hard problems? Did they communicate blockers early? Did they help the team recalibrate when a target turned out to be wrong?
  3. Make the scoring model explicit, and repeat it often. If 70% is a good outcome, say that in every check-in, every planning cycle, every all-hands.
  4. Protect the people who miss ambitious goals honestly. If someone sets a genuine stretch goal, works hard, communicates clearly, and lands at 60%, that has to be a better outcome, reputationally and materially, than someone who set a safe goal and hit 100%. It’s the same pressure that pushes people toward fake positive OKR updates when the truth costs them something. Remove the cost, and the honesty comes back on its own.

Here’s what that split looks like in practice:

What happens OKRs tied to comp OKRs kept separate
Targets Get negotiated down before the quarter starts Stay ambitious, because there’s nothing to lose
Check-ins Feel like an audit Stay focused on progress and blockers
Honest misses Feel like punishment Get judged on effort, adaptability and honesty
The score Becomes a number to manage Stays a true signal of calibration

“But then what motivates people?”

The most common objection: if OKRs don’t touch pay, why would anyone push hard on them?

Two answers. First, intrinsic motivation research is pretty clear that autonomy, mastery, and purpose drive sustained effort more reliably than a bonus tied to one specific number, especially for knowledge work where the “right” target is genuinely uncertain going in. Second, and more practically: people are still motivated by review outcomes, promotion, and reputation. None of that goes away when you remove OKRs from the equation. It just gets judged on the fuller picture instead of a single lagging score.

A quick example

Picture two product managers, given roughly the same resourcing and the same market. One sets a target to grow activation by 8%, hits 9%, and gets a glowing review. The other sets a target to grow activation by 25% (a genuine bet that a new onboarding flow could move the number), lands at 16%, and ships a version of the product that quietly becomes the foundation for the next two quarters of growth.

Score the first PM 100% and the second 64%, and a comp-linked system rewards the wrong one every time. That’s not a hypothetical edge case. It’s the default outcome of tying pay to a raw key result percentage, and it’s exactly why we keep the two systems apart.

The trade-off, and why we still think it’s worth it

This approach asks more of managers. Rating “did you do the job well” without a single clean number attached is genuinely harder than pointing at a dashboard and saying “70%, so that’s your rating.” It requires actual judgement, actual conversations, and a level of trust that a lot of performance management software is specifically designed to avoid needing.

We think that trade-off is worth it. The alternative is a system that looks rigorous on a slide deck and quietly trains your best people to stop being ambitious. If OKRs only ever get used to set the safest goal that still looks good, you’ve built an elaborate reporting exercise, not a growth tool.

This is where Tability comes in, not as the thing that solves the judgement problem for you (nothing does), but as the layer that makes the separation easy to run in practice. Check-ins stay lightweight and focused on progress and blockers. Scoring stays visible without being wired into anyone’s compensation. Managers get a clear history of how someone approached their goals, ambition, adaptability, honesty about misses, without a single number pretending to summarise a whole quarter of work.

The bottom line: OKRs only work as a learning tool for as long as they stay separate from the pay conversation. The moment they merge, the incentives quietly rewrite the goals, no matter how good the framework looks on paper.

Keep your OKRs ambitious

If you’re trying to keep OKRs ambitious instead of watching them shrink to whatever’s safe, sign up free or book 30 minutes with us and we’ll show you how teams run this in practice.

Author photo

Bryan Schuldt

Co-Founder & designer, Tability

Share
Weekly insights for outcome-driven teams
Subscribe to our newsletter to get actionable insights in your inbox.
Related articles
Read more →