Rank-conditioned empirical study · 14

What 5d, 9d, and professional policies see in the same Go position

Across the same 50 professional games and 500 positions, a 9d Human SL policy covered 82.7% of the AI one-point good set at the median, versus 77.5% for 5d. The difference is real, but it is not a ladder that rises in every position.

  1. Give four human-policy observers the identical frozen positions

    The sample fixes 50 professional games and ten middlegame positions per game, for 500 decisions. A strong b11 Transformer teacher supplies candidates at 128 and 512 visits. Human SL then returns rank_5d, rank_9d, era-matched professional, and 2023 professional policies for every position. Because the positions are identical, the comparison measures policy-distribution shifts rather than giving each rank a different test.

    One frozen sample flows through 5d, 9d, era-matched professional, and 2023 professional policy channels.
  2. Do not begin by asking who guessed one exact coordinate

    Candidates retained by the 512-visit teacher and lying within one point of its leader form the good-move set. The set respects a continuous value surface: if six moves are nearly tied, entering five of them is not a failure to see the idea. Rank comparisons therefore use policy mass on the good set while retaining exact-top probability as a secondary measure, instead of turning search noise into one required answer.

    The exact top move is one point; the one-point good set is the more stable teaching region.
  3. The 9d policy sees more on average, by about five points

    Median good-set coverage is 77.5% for the 5d policy and 82.7% for 9d. Paired by position, 9d is higher by 5.1 percentage points on average, with a whole-game cluster-bootstrap 95% interval of 4.3–5.9 points. The 9d policy covers more in 76.2% of positions. This is a stable movement of probability mass, not a claim that the higher profile wins every item.

    Paired across 500 positions: 5d 77.5%, 9d 82.7%, and a 5.1-point mean difference.
  4. A higher rank is not more AI-like in every position

    Nearly one quarter of positions do not give the 9d policy higher coverage, and a small tail gives 5d materially more mass on the current AI-good region. Frequencies of familiar techniques, position type, sparse training strata, or an unstable teacher can all contribute. The defensible claim is that the overall distribution shifts with rank—not that one Human SL probability reads a particular player's mind or certifies a rank.

    The aggregate trend preserves counterexamples instead of erasing position type and model uncertainty.
  5. A teaching boundary is more useful than a rank label

    If a move is already common under 5d, review may only need to confirm the problem it solves. A sharp rise from 5d to 9d suggests a near-term learning target. If even 9d and professional policies assign very little mass, first audit budget, model, and symmetry stability before collecting it as a deep move. A product should show this reachability band instead of declaring that every 5d player must know one coordinate.

    Confirm, teach, collect, or defer according to reachability and label stability together.
  6. Personalized review needs receipts, not a hidden total score

    Each proposed lesson should retain played-move loss, the one-point candidate set, target Human SL profile, set probability, search budget, budget agreement, and model identity. If only a final label such as “for 5d” survives, a model upgrade becomes indistinguishable from human improvement. Keeping the components visible lets a player or coach explain exactly why a ranking is unhelpful and allows the system to recompute it later.

    Loss, discoverability, and stability remain separate so future models can rerun the ranking.
  7. The next experiment must test transfer, not another average

    This study finds a stable rank-conditioned distribution difference; it does not show that reading an explanation improves play. The causal next step randomizes review ordered by maximum point loss against review ordered jointly by loss, reachability, and stability. Days later, structurally similar positions from different games should test checking order and candidate sets. Rank-aware teaching becomes a learning fact only when judgment changes in unfamiliar positions.

    After offline probability differences come delayed recall and transfer to unfamiliar positions.

Data, model, and record sources

  1. KataGo · Human SL Analysis Guide
  2. KataGo public networks and rating notes
  3. AEB/CWI public-domain professional records

Before generating rank-aware teaching

  • The same frozen positions are compared
  • A good set replaces one required coordinate
  • Target profile and search budget are recorded
  • Counterexamples and unstable labels remain visible
  • Learning impact is still marked untested
Open the discoverability study