GTM engineering · WorkflowReference design

Calibrate account scores against outcomes

Compare fit scores with mature deal outcomes. Change one rule at a time, keep a holdout, and retain the old version so every account movement can be explained.

Calibrate account scores against outcomesDesign map
  1. 01
    Build a mature cohort

    Export wins, losses and disqualified accounts with known outcomes.

  2. 02
    Compare the current model

    Freeze features as they were before the outcome.

  3. DecisionDoes the change improve a mature holdout without unacceptable exclusions?
    Alternate path

    No: retain the current model and investigate the missing evidence.

    Continue

    Yes: review movements and release a versioned change.

  4. 03
    Challenge the proposed rule

    Check holdout performance and commercial costs.

  5. 04
    Version and release

    Review tier movements, preserve rollback and monitor.

The score is a decision rule

A monthly review or a recurring pattern of bad-fit meetings triggers calibration. The output is an approved rule change with evidence, affected accounts and a rollback plan. It is not an AI-generated weight table that immediately reorders the sales queue.

Keep fit separate from timing. Company size, business model and ability to benefit describe fit. A relevant recent change describes timing. A familiar brand or a busy LinkedIn account is not automatically a better prospect.

Build a cohort you can trust

Use closed-won, closed-lost and disqualified records. Record whether the loss was fit, timing, commercial terms, competition or an unknown reason. Treat open deals as unfinished observations. Segment by market and motion where the economics differ.

Recover each feature as it existed before the sales outcome. Using today's customer status, newly discovered budget or post-sale product usage leaks the answer into the model. If historical features cannot be recovered, name that limitation and use a prospective test instead.

Run the review

  1. Freeze the current score version and the cohort cutoff date.
  2. Reconcile duplicate opportunities to accounts. Define whether the test predicts qualified opportunity, win or another specific outcome.
  3. Examine exclusions before weights. A highly ranked account that cannot buy should not be rescued by engagement points.
  4. Compare outcome rates by score band and segment. Show counts and denominators, not only a blended percentage.
  5. Investigate false positives and false negatives with sales. A low-scoring win may reveal an omitted segment rather than a bad coefficient.
  6. Propose a small, readable change. State the expected queue effect and the cost of mistakenly excluding an account.
  7. Test on a time-based holdout that was not used to choose the rule. Where the sample is small, retain the rule as a hypothesis.
  8. Review accounts that change tier. Release the version, alert owners and set a check date before rescoring open work.

A calibration receipt

Version proposalIllustrative example
Question
Does the new size rule improve qualified-opportunity selection?
Training cohort
Mature outcomes before the cutoff; counts recorded in the register
Holdout
Later outcomes not used to pick the rule
Risk
Excluding a small, high-value segment
Release
Draft until the owner reviews tier changes and rollback

When to hold the change

Too few mature outcomes, inconsistent loss reasons and a change in sales coverage can all distort apparent performance. Do not compensate with more elaborate math. Repair the inputs, or run the rule in shadow mode while the current routing stays live.

Measure selection precision, missed qualified accounts and workload by tier. Review whether sales worked the cohorts comparably. A higher win rate caused by giving one tier all the senior attention does not prove the score predicts fit independently.

Write the learning back

Store the rule, evidence, reviewer, version date and outcome definition in the company brain. The CRM receives score, tier, reason and version. Keep the previous version available for rollback and historical comparison. Use the decision register to decide whether the test warrants a permanent change.

Publication and source record

Published on GTMhub: . Last reviewed: .

How we run this

Put it to work in your company.

Our GTM engineering lane plugs into your company’s GTM brain. Start with one lane; services can run in parallel.

Book a call
  1. Week 1Your GTM brain

    Monday kickoff. A 45-minute review Friday to confirm the context.

  2. Weeks 2–3Connect the lane plugin

    Wire it into your existing tools and test the work with your team.

  3. Week 4Live, then improve

    About three hours of your time in month one, then 15 minutes a week.

The first term is three months. The brain and lane plugin are handed over in full after it. Your accounts and data stay yours.