Scorecards, Leveling & Decision Rules
If maybe, then no. • Leveling is settled before the offer, because level sets the band.
Overview
A loop without a scorecard is a collection of opinions. This part is the scoring machinery: how each round is written down, how the team converges on a level, and the small set of hard rules that keep the bar from drifting. It is deliberately mechanical, because mechanical is what makes it fair and repeatable.
Two layers, applied every round
Every round produces two things. First, a per-pillar score on a 1–4 scale, each with an evidence note and a counter note — the counter forces the interviewer to argue against their own read. A 4 is the top of the scale. Second, a level on one shared ladder used everywhere: L3 through L7, expressed as rows so that red flags live in the low cells and strong signals in the high cells.
| Level | What it reads like |
|---|---|
| L3 | Executes scoped, well-defined tasks with guidance; reasons by analogy to prior jobs; generic, "we" answers |
| L4 | Owns a scoped problem independently; defensible when pushed; one real domain of depth, thin elsewhere |
| L5 | Sets direction on a scoped area; makes reversible calls without asking; depth in several domains, consequences named |
| L6 | Shapes strategy across teams; names non-goals and tradeoffs first; deep and broad; says where each domain is overrated |
| L7 | Drives multi-year direction; reframes the problem and names the binding constraint; could teach the specialist |
The gate rule and the "4" rule
Not all pillars are equal in a given round. Each round has one or two load-bearing pillars — the ones that round exists to test. The rule: a load-bearing pillar must reach the top of the scale to advance. A 2-out-of-4 on a gate pillar is a fail even when adjacent pillars are strong, because the round was designed to measure exactly that thing. A brilliant system designer who cannot reason about blast radius is not a pass on the round whose job is blast radius.
If maybe, then no
The most important decision rule in the system is the tie-breaker: ambiguity on a load-bearing pillar resolves to no-hire. A "maybe" is a "no." This feels harsh and is correct — the cost of a wrong yes (an eventual managed exit, months of drag, a demoralized team) dwarfs the cost of a wrong no (you miss one good person who has other options). Teams that let maybes through end up running the firing half of this system far more often than they should.
A companion rule catches shallow confidence: if a candidate answers a mentality probe in under ninety seconds and never says what actually goes wrong, cap them at L5. Speed without failure-mode awareness is not seniority.
Hard gates, checked early and re-checked
A few requirements are pass/fail and are screened at the recruiter stage, then re-confirmed: genuine production experience in the core technology, a real zero-to-one build (built, not merely adopted or migrated), the arc into the domain, and — increasingly non-negotiable — AI actually in the daily workflow. A gate that is technically met but only at trivial scale is recorded as "pass-weak," which is a flag, not a clearance.
The debrief and leveling before the offer
When a candidate clears every round, the panel meets and agrees on a single level — "high L5," say — before anything goes to offer, because the level determines the compensation band. Do not let leveling happen implicitly during offer negotiation; decide it in the debrief where the evidence is fresh and the whole panel is present.
Calibrate the interviewers, not just the candidates
The scorecard only works if interviewers apply it consistently, so interviewer calibration is itself gated. A new interviewer shadows a round once and reverse-shadows once (they lead, an experienced interviewer observes) before running it solo. Until a round is calibrated across the team, staff two interviewers on it; once it's stable, one runs it. Counter-intuitively, calibrate against a mid-level engineer, not your most senior one — a very senior person "would ace the loop regardless" and can't tell you where the bar actually sits. A simple health check: have two interviewers score the same candidate independently and look at how often their yes/no agrees. Agreement far above chance means the bar is real; near chance means you're measuring noise.