For the first twenty-five years I ran a company, we were great at putting out fires. Nobody owned making sure there was never a fire. Everyone owned their job, and the fires started between the jobs. When it was time for a performance review, I had my opinion about the person's work and they had theirs. Nobody had written down what winning looked like for their role, so I told them what to do. The moment I did, I owned the outcome, not them. If the fix failed, that was on me. If it worked, they were following orders. Either way, the person doing the job stopped owning their work.
A scorecard is where this ends up, not where it starts. It is a one-page definition of what success looks like for every role in the company. It names the function the person owns, their mission, the number that says whether the function is winning, the behaviors the company expects them to live, and the skills required to do the work. Every person accountable for a function inside the company has one, from the CEO down to an individual contributor. The scorecard is how they know if they are doing their job.
Throughout I use the word "seat" for the role a person occupies. A seat may cover one function or several. When a founder owns marketing, sales, and delivery, their seat covers all three. When a Head of Sales is hired, the founder's seat narrows and a new seat opens. The scorecard follows the seat, not the function.
Building a scorecard for the first time is mostly assembly work. You pull content from three documents that should have already been developed prior to getting to this point:
- The Functional Accountability Chart (FAC), which holds the mission, critical number, and supporting indicators for every function
- The company Values, which define the behaviors the team expects of each other
- A Competencies library that defines the skills required at each level of responsibility.
What This Looks Like in Practice
Take the Head of Marketing at a SaaS company finishing Q2. Her critical number is qualified pipeline generation. The company set green at $2M, red below $1M, for her cycle. This quarter she came in at $1.7M, yellow. She scored Always on all four values and Always on seven of the eight competencies for her tier, with a Sometimes on Vision and Strategy, declining trajectory. The review surfaced that her weekly one-to-ones have drifted into status checks, and the team is not hearing the strategic story often enough.
Her classification is B. Yellow on the critical number alone disqualifies her from A under my formula, regardless of values and competencies; the Sometimes on Vision and Strategy confirms B and points the development commitment at where to focus.
In the conversation, the Head of Marketing went first, as the process requires. She walked her manager through her own read: pipeline at 85% of target, a Sometimes on Vision and Strategy because her weekly one-to-ones had drifted into status checks and the team wasn't hearing the strategic story often enough. The manager listened, asked questions, and then shared his own assessment, which lined up with hers. Two beats also went on the table that the scorecard didn't capture: she had killed a brand-damaging campaign mid-quarter (a judgment call her manager weighed in her favor) and her two best team members had started covering Vision and Strategy for her in standups (a structural workaround that explained the rating). The development commitment came out of that conversation, not the rating itself.
The development commitment reads: Situation, the team cannot articulate the quarterly strategy. Cause, her communication cadence has compressed around operational firefighting. Correction, publish a one-page quarterly strategy note at the start of Q3 and hold a 30-minute team Q&A every month of the cycle. Follow-up, at the next scorecard review the manager confirms the strategy note was published at Q3 kickoff and that all three monthly Q&As happened. Green if both delivered on time. Yellow if delivered late or partially. Red if either was missed. The ratings and the commitment go on the scorecard, signed off by her and her manager. Next quarter's review opens with that follow-up.
That is one quarter, one seat, one page. The rest of the article walks through what each piece is, where it comes from, and how the conversation that produces it actually runs.
What the Scorecard Contains
Five elements, drawn from the three documents above: the mission, the critical number with its supporting indicators, the values, and the competencies for the seat's tier.
The mission answers why the role exists. Without it, the person in the seat optimizes for tasks rather than outcomes. "Run marketing" is not a mission. "Generate qualified pipeline that enables the sales team to hit revenue targets" is a mission. It names the outcome the function has to produce.
The critical number is the one metric that says quantitatively whether the function is winning or not. Each critical number carries a green threshold and a red threshold. Green for this cycle means the function is working right now. Red means it is not. Yellow is not set separately; it is the middle zone that emerges once green and red are fixed. Without a critical number and its thresholds, performance is felt, not seen.
The supporting indicators are the leading and lagging measures that sit around the critical number. A leading indicator is what you can measure this week to know where your critical number is headed, while there is still time to save it from going red. Without it, you find out you missed your number only after you missed it. A lagging indicator is the outcome that proves the leading indicator translated into real business results. Without it, you can celebrate good inputs that never produced output. Both use the same two-threshold model as the critical number: green and red are set, yellow is the resulting middle zone.
Those three come from the FAC row. The scorecard adds two more.
The values alignment is a quarterly read on whether the person is actually living each company value. Always, Sometimes, or Never. This is the behavioral side of the scorecard. It answers whether the person belongs in the company you are building. Without values alignment on the page, performance and values pull against each other, and people who hit the numbers but poison the culture keep getting rewarded.
The role competencies are the specific skills required to do this job at the level required. Same three-point rating: Always, Sometimes, or Never. This is the capability side. It answers whether the person has the skills to do the job. Without named competencies, development conversations are about personality or they go in circles.
To pick the right competencies for a seat, look at the Function Organization Chart. The seat's position in the hierarchy tells you which tier of the competency library applies: individual contributor, people manager, department head, or executive. Each tier inherits the one below and adds more.

I chose not to put technical skill competencies on my scorecards. My reasoning was that the stated mission, the critical number, and its supporting indicators already measure whether the function is producing the outcome it is responsible for. That is the practical test of whether the person has the technical chops to do the job. Naming each technical skill separately felt redundant, and it pulled development conversations toward skill checklists instead of outcomes. If you cannot decide, start without technical competencies. Add them only after a full quarter of reviews shows a gap the current scorecard cannot explain.
Together the five elements answer three questions on one page. Is this function producing the outcome the business needs? Is the person living the values the company expects? Do they have the skills to do the job?
Scoring values and competencies is data work, not gut work. Scored consistently across quarters, these ratings become a record you can read. A competency that stays at Sometimes for three quarters is not a personality problem. It is a signal that coaching is not moving it. At that point the question shifts. It is no longer how to coach the skill, but whether this is the right seat for this person or the right scorecard for this seat.

How A, B, and C Work
The classification is the founder's call about how the company defines A, B, and C. Different founders draw the lines differently. The discipline is written-down and consistent: write your formula, apply the same one to every seat in a tier, every quarter.
Here is how I do it. If you already have a written formula that is running, reconcile yours against mine and keep what is tighter. If you have never written one down, use mine for one cycle, see what it tells you across every seat, then write your own.
My A Player. Four conditions, all true at review:
- Every critical number is 🟢 Green
- All values are rated Always
- All competencies are rated Always
- If a development commitment was active for the cycle, the Follow-up confirms it closed
A single Sometimes on any value or competency drops the seat to B for the cycle. I keep it strict because A has to be rare. If everyone on a tier ends up rated A, the rating has stopped distinguishing performance, which was the whole point.
My C Player. Three triggers. Any one makes the seat C for the cycle:
- Any critical number is 🔴 Red
- Any value is rated Never
- A competency has been rated Never for two cycles in a row
Notice the asymmetry between values and competencies. A Never on a value triggers C the same cycle. A Never on a competency only triggers C after two consecutive cycles. The difference is deliberate. Values are character. You don't coach someone into integrity or care; they either show up that way or they don't. So a Never on a value is a same-cycle conversation about whether this person belongs in this seat at all. Competencies are skill, and skill takes time to coach. One Never starts a watch; two consecutive Nevers make it a C trigger.
A high performer who violates values is still a C player, often called a Toxic-A. Strong numbers make them visible and influential. Their values violation makes that visibility dangerous to the culture. The Toxic-A typically lands at C through the value-Never gate.
The same-cycle conversation has specific words: "I rated you Never on [value]. The moment was [specific moment]. The classification this cycle is C. The next 90 days are an improve-or-transition plan, not a coaching cycle." A value-Never should rarely be a surprise at the review. If it is, that is the manager's coaching gap, not the seat owner's failure to read minds, and it gets named in the manager's own next review.
My B Player. Whatever isn't A and isn't C. Most people most of the time. B is not a verdict on the person; it is a flag that one or two coachable gaps are real this cycle. The development commitment is the lever that moves them.
Yellow is the watch zone. A Yellow critical number drops the seat to B for the cycle (the A bar requires every critical number Green). It does not by itself trigger C; that takes Red. But Yellow that persists has consequences: two consecutive Yellow cycles, or three Yellows in any rolling four cycles, escalates to a same-cycle Red conversation. The manager and seat owner sit down inside the cycle, not at the next review, and decide whether the cause is inside the seat's span of control or outside it. If inside, the cycle classifies as Red and the C trigger applies. If outside (a supplier issue, a market move, a capital decision the seat doesn't own), the rating stays Yellow and the system constraint gets named on the scorecard so it doesn't come up again next cycle. The rule is the forcing function. The judgment is still the manager's.
Where the Scorecard Comes From
The scorecard is compiled, not designed. No part of it is invented on the fly. The mission, critical number, and supporting indicators come from the FAC row. The values come from the company values list, whether those were recently named or have been lived for years. The competencies come from the tier of the library defined by the seat's position on the Function Organization Chart. All of the pieces must reconcile.
Shannon Susko's Metronomics framework teaches the scorecard as its own design document. I build the mission, critical number, and supporting indicators into the FAC row itself. The scorecard then compiles from the FAC row, the values, and the competencies. Same destination, different path.
The rollout is gradual and sequential. Four stages, each with a clear way to know when you're done before moving to the next.
Stage 1. Share the critical numbers. Put them on a scoreboard, in the weekly meeting, in whatever system the team already uses. Let people see their function's number. Let them use it to make decisions. No scorecards, no ratings, no pay tied to the numbers yet. Spend at least one quarter here. Some critical numbers will be wrong. Some thresholds will be too loose, others too tight. Use the time to tune. Done when: the team is used to looking at the numbers, discussing them, and adjusting behavior in response, and the thresholds have stopped churning.
The number stays the seat owner's, even when it's bad. The most common way Stage 1 fails is the manager carrying the number for them. Three tells:You retune the threshold quietly between meetings instead of asking the seat owner to propose the new number in the weekly meeting.You hunt the data yourself because the seat owner's read isn't ready by the meeting.You explain a poor result to others on their behalf so they don't have to.
Make them do it badly for a quarter rather than do it well for them. Each tell trains the team to think the number is yours, not theirs, and the scorecard fails silently before it is even built.
Stage 2. Map the full FAC across the whole company. Every function, every owner, every critical number with its thresholds, every leading and lagging indicator. Not row by row as you roll out scorecards. The FAC is the shared picture of how the company makes money, and a partial picture is worse than no picture. Done when: every seat the company plans to scorecard has a FAC row it can inherit from.
Stage 3. The CEO writes and signs off on the first scorecard. If you are the CEO, you cannot ask anyone else to have a scorecard before yours is written. Your scorecard models the vulnerability you are asking others to show. Do not skip this step. Share the result with the full leadership team. Done when: the CEO has one full quarterly review on their own scorecard before cascading.
Stage 4. Each leader builds theirs next, then each person they manage, and so on. One scorecard per seat. The scorecard follows the seat, not the function. Done when: every seat on the FOC has a scorecard that has been through at least one review.
Across all four stages, do not tie compensation to the critical number. Pay tied to a metric you have not yet validated guarantees you pay for the wrong behavior. You will then tune the threshold, because the first draft is always wrong. Tuning for improvement is healthy. Tuning because you tied pay to a number before it earned trust is a self-inflicted credibility wound, and the team can tell the difference. Compensation covers how to wire pay to the system once the numbers have earned trust.
The scorecard adds the values column and the competencies column for the person in the seat, but only once the FAC row it sits on has been set in the context of every other row. As the company evolves, FAC rows get updated: new functions, splits, restructures. That is maintenance, not creation.
One person, one scorecard. When a seat covers multiple FAC rows, those rows combine: multiple missions, multiple critical numbers, one values assessment, one competencies assessment. When the seat is split, the scorecard is rebuilt for the narrower role.

Values can be in draft form. The competencies library cannot; you cannot rate against skills that have not been named. Build it first using the Competencies article and its Builder prompt. The CEO's FAC row should be written and tested against real performance.
Scorecard Builder Tools
Compile scorecards for any seat, or every seat, from your FAC, Values, and Competencies. Pure compilation, no invention during the build. 30-60 min per seat.
The Quarterly Review Conversation
Scorecards exist to be reviewed. A 90-day cycle for each seat, staggered across the team so reviews don't all land in the same week. A leader with six direct reports doing all six reviews in the same week burns out and shortcuts; spread them across the quarter. The review is where the scorecard stops being paper and starts driving behavior.
The week before the review, the person in the seat rates themselves. They do it alone. They score themselves against their critical number, against each value, against each competency. If there was a development commitment from last quarter's review, they assess whether they achieved its measure of success. They write down what they are seeing and what they are not.
The manager does the same assessment independently. Same scale. Same content. No conversation yet.
The review meeting has a deliberate order. The person being reviewed speaks first, walking through their self-assessment in full: their reading of the critical number, each value, each competency, and the evidence behind each rating. The manager listens and asks questions to fully understand the perspective. Only then does the manager share their own assessment, which may have shifted based on what they heard. Telling closes the conversation; asking opens it. When the ratings match, move on; alignment is confirmed. Where they diverge, the gap is the whole point of the review. If a person rates themselves Always on accountability and the manager rates them Sometimes, that is where the conversation lives. Spend the time on the friction, not on what you both already agree about. The two ratings get walked together with the evidence each side brings, and the manager has the final call on the rating that goes into the Review Record. The seat owner's self-rating stays visible (recorded as their perspective) so the gap isn't erased; it is named, and that is what makes the next conversation different.
The rating mechanic that keeps reviews honest: Always is the default, no evidence required. Sometimes and Never need a specific observed moment since the last review, named by the manager, not asked of the seat owner. The seat owner does not have to prove they were Always; the manager has to name the moment that says they weren't. No moment, no rating change; the rating stays Always. The one exception is when both sides land on a contest the manager can name but neither can produce a proving moment for; in that case the default tilts to Sometimes and the conversation opens the gap. Without these two rules together, review meetings drift in two directions: into "I think she's bad at this" without specifics, or into quiet down-rating of the operator who doesn't self-promote.
The scoring scale is three points: Always, Sometimes, Never. It is intentionally blunt. Five-point scales invite hedging. Everyone clusters around three and a half and calls it done. Three points force a binary on each line item: is this consistent, or not? You either saw the behavior since the last review or you didn't. What you compare across quarters is each rating on its own line, not a number rolled up across many. Sometimes is not a safe middle. It is a flag that development is needed. The ratings are not a verdict on the person. They are a read on a specific behavior in a specific quarter. A person who is rated Sometimes on accountability this quarter is not a Sometimes person. They are someone whose accountability has a specific, coachable gap right now.
Do not embellish the scale. Three points stays three points. Not five, not seven, not weighted composites. Every addition here dilutes the conversation rather than sharpens it.
Never sum the scores. Never average them. Someone who scores Sometimes on one competency does not become a B Player overall. A weighted composite hides the exact coordinate you need. If a value is rated Never and the critical number is Green, the problem is not a 2.4 average. The problem is that this person is a Toxic-A, and the composite makes that invisible. You are not grading a test. You are coaching a person. Read each score as its own observation. If nuance is needed, add a one-word note: improving, stable, or declining. That tells you whether to invest more coaching or start a different conversation.
The meeting covers five areas. First, a check-in on the prior development commitment (if any) and whether it met its measure of success. Then how the critical number performed over the quarter, the ratings comparison on values and competencies, and (if the company is running skip-level reviews) the synthesis from the skip-level review for anyone who manages others. It closes with one new development commitment for the next quarter.
Skip-level reviews are an advanced practice that makes most sense once scorecards are standard and the seat has enough direct reports for the survey to be anonymized meaningfully. When the company is running them, the synthesis is shared input across every party in the review, including the seat owner. It informs each party's perspective; it is not a separate input that only the coach brings.

The review ends with one development commitment. Not three. Not five. One. It has four parts: Situation (what is happening), Cause (why), Correction (the specific action the person will take), and Follow-up (what will be measured, what success looks like, and when it will be assessed). "Work on communication" is not a development commitment; it is a restatement of the gap. A real commitment names all four parts. The action has to be something the person will do, visibly, in the next 90 days, with a way to tell whether it worked.
The four parts have a discipline that keeps them useful:
- Situation describes the observable behavior. No adjectives ("timely," "consistent"), no shoulds ("she should be raising"), no consequences ("too late to course-correct"). The standard already names what good looks like; the situation just describes the gap.
- Cause names the mechanism behind the situation. One sentence. Not the downstream effect, and not blame.
- Correction states the new behavior positively. What does the seat owner do now. Don't editorialize the prior process.
- Follow-up names what will be measured, what success looks like, and when the assessment happens. The measure tells you what Green looks like (the commitment closed), what Yellow looks like (partial), and what Red looks like (missed), so the next review has something concrete to verify, not a subjective read. The cadence matches the corrected behavior: weekly behaviors get weekly Follow-up in the seat owner's one-on-one; per-cycle behaviors (don't pivot the roadmap mid-quarter) get Follow-up at the next quarterly review. Mismatched cadence reproduces the lag the commitment was supposed to fix.
Adjectives invite argument; facts don't. The discipline turns the development commitment into a contract rather than a performance review.
Stop here if the four-part write-up isn't crisp. A vague Situation, a fuzzy Cause, a Correction that names a feeling instead of an action, or a Follow-up without a measurable check means the development commitment is not ready to sign. Send it back for another pass. A vague commitment is worse than no commitment, because it gives both sides something to argue about instead of something to do.
The development commitment does not have to come from a rated gap on the scorecard. The most common source is a Sometimes or Never that surfaced in the review, but the commitment can equally be a critical-number issue, a skill the company needs the seat owner to develop (AI fluency, financial modeling, public speaking), a career-growth area the seat owner wants for a future role, a theme from the skip-level review, or anything else the manager and seat owner agree to work on for the next 90 days. The four-part structure applies regardless of source. The closure check at the next review is the Follow-up's measurement, not a rating on the scorecard.
Each review produces its own development commitment, or none. There is nothing to carry between reviews. No list of open commitments, no formula for whether last quarter's work counts toward this quarter's classification. If the same gap shows up cycle after cycle, you write a fresh commitment for it. The manager and the seat owner decide each time whether the prior approach worked or needs sharpening.
The CEO review follows the same pattern with one wrinkle. The CEO has no manager. Instead, whoever plays that role for the CEO (a coach, the board, an owner) takes the manager role for the review. The result is shared with the full leadership team. Sharing with the leadership team is what makes the CEO's accountability visible rather than theoretical. It models the transparency the CEO expects from everyone else.
The review conversation itself is where the change happens, so I have built a second prompt called Scorecard Review. Each party (the seat owner, the manager, and any third-party reviewer like a coach or board chair) runs it independently before the meeting. It produces their perspective draft: ratings, evidence on the gaps, and a development commitment recommendation. Neither side sees the other's ratings before the meeting.
After the meeting, the perspectives are reconciled, with the manager having the final call where ratings differ. The manager (or coach) writes up the final Review Record using the agreed conclusions: the critical number result, each value, each competency, the improving/stable/declining note, the classification, and the next cycle's development commitment. The prompt produces the inputs; the meeting produces the verdict. Once signed off, the scorecard entry is final and sits alongside prior quarters in the scorecard history.
Scorecard Review Tools
Walk through the quarterly self-rating, manager rating, and development commitment. Independent ratings before the meeting, one structured conversation in it. 30-45 min per party.
What Comes Next
The classification is then stress-tested in the quarterly A-Player Team Assessment meeting, where the leadership team reviews the full roster together on a scatter plot that places every person by their values and performance ratings. Borderlines get resolved in the room. That is where "A player" stops being abstract and starts carrying the same meaning across every manager in the company.
The classification is not about labels. It is about where to invest coaching energy. It is also where compensation decisions get anchored, which is a separate conversation covered in Compensation. A players need to be recognized and stretched. B players need focused development so they can move up. C players need a clear 90-day improvement plan or a respectful transition.
The 90-day improvement plan for a C player is a development commitment with higher stakes: improve or transition. Same four-part structure, with the Follow-up defining both what success looks like and what happens if success isn't reached.
Why This Works
People stop going home wondering whether they are doing a good job. They can see their critical number, their supporting indicators, and their own values and competency ratings. The conversation moves from "how do you feel about your performance?" to "here are the numbers, here are the gaps, here is what we are going to do about it."
For managers, the shift is from giving answers to asking questions. Once the seat owner can already see what's working and what isn't, the manager stops being the person who provides solutions and becomes the person who asks what the seat owner is going to do about what they can already see. Ownership stays with the person in the seat, where it belongs.
Over time, the team gets stronger by itself. A players attract A players. B players either become A players, transition to different roles, or leave. C players transition or leave. Accountability becomes real instead of theoretical, and you stop being surprised by performance problems because you are catching them quarterly, not annually.
For my first twenty-five years running a company, my team and I were great at putting out fires. None of us owned making sure no fires started. The scorecard is what changed that. Every function got an owner who saw the smoke before I did. We stopped fighting fires and started building a company that didn't need them. That is the shift the scorecard was designed to create.