How we rank D2 volleyball
Every ranking is a claim about who's better. We test ours the only fair way: replay the season one day at a time, predict each match using only what happened before it, and keep score. Then we run every other kind of system through the same test, on the same matches.
The short version
- Points beat wins. How many rallies a team wins says more about how good it is than whether it won. Systems built on wins alone (the RPI, the NCAA's NPI, Colley) predicted worst in every season we tested: log loss 0.500 against our 0.430 over 9,254 matches.
- Conference strength has to be measured. D2 conferences mostly play each other, so a 17–1 record in one conference and a 17–1 record in another can mean very different things. We judge each team against its own conference, and each conference by how its teams do against everyone else.
- Recent form counts more, but conference strength keeps its evidence. The non-conference matches that tell conferences apart are played early, so they never fade.
- Last season is the starting point, then fades as this season's matches arrive. That's why our early-season numbers are steadier than most.
- Two questions, two rankings. Résumé ranks who has earned the most; Power ranks who'd win tomorrow. We show both, plus the NCAA's NPI, because they answer different questions.
How we test
- Walk forward, no hindsight. For every match day, each system is fitted only on matches played before that day, then predicts that day's matches. Nothing it predicts was in the data it learned from.
- Same matches for everyone. Every system predicts every match between two D2 teams that have each played at least once that season.
- Score the confidence, not just the pick. Our main score is log loss: calling a team a 90% favorite and losing costs far more than calling it 55% and losing, so a system can't look good by being overconfident. Lower is better; flipping a coin scores 0.693. We also track the share of winners picked and the subsets that matter most: cross-conference matches and the NCAA tournament.
- Is the difference real? We test it two ways. For win chances, a paired bootstrap: resample the matches thousands of times and count how often the other system comes out ahead. For rankings, McNemar's test on the matches where the two rankings disagree. "p < .001" means a gap that size would happen by luck less than one time in a thousand.
- Tuned on one slice, checked on others. The model's settings were chosen on the first weeks of 2026, and the win-chance calibration on 2024 alone; both were then confirmed on full seasons they never saw, so the results aren't fitted to the test.
The results
All 3 seasons together (2024, 2025, 2026): 9,254 matches
| System | Log loss | Better than a coin flip | Picked the winner | Cross-conference | Seasons we beat it (p < .01) |
|---|---|---|---|---|---|
| D2VB Power (ours) Points, conference hierarchy, recency, last season | 0.430 | 38% | 79.3% | 0.460 · 78% | — |
| Pablo-style Points (capped), recency | 0.459 | 34% | 77.9% | 0.516 · 75% | 3 of 3 |
| Our old Power Points, shrink to the D2 average | 0.459 | 34% | 78.2% | 0.524 · 75% | 3 of 3 |
| Massey (1997 least squares) Point margin | 0.482 | 31% | 78.1% | 0.572 · 76% | 3 of 3 |
| Sets only Sets won and lost | 0.482 | 30% | 76.9% | 0.548 · 74% | 3 of 3 |
| Wins only Wins and losses (the RPI/NPI/Colley family) | 0.500 | 28% | 75.1% | 0.551 · 71% | 3 of 3 |
| Elo with margin (538-style) Wins, scaled by margin, game by game | 0.506 | 27% | 75.0% | 0.554 · 72% | 3 of 3 |
| Ranking | Matches | Its higher-ranked team won | Ours, same matches | Where they disagreed | Significance |
|---|---|---|---|---|---|
| RPI (classic 25/50/25) | 9,254 | 73.5% | 78.4% | ours right 955, theirs right 504 | p < .001 |
| NCAA Power Index (our calculation) | 9,221 | 74.9% | 78.5% | ours right 846, theirs right 523 | p < .001 |
| AVCA coaches poll | 1,332 | 84.5% | 85.3% | ours right 54, theirs right 43 | p = .310 (not yet significant) |
Averages are weighted by matches. The ranking test pools every match where the two rankings disagreed. The coaches' poll only ranks 25 teams and covers far fewer matches, so it takes several seasons of disagreements to separate it from ours; this test grows every week.
Season by season
Every system, predicting this season
| System | Uses | Log loss | Better than a coin flip | Picked the winner | Cross-conference | vs. ours |
|---|---|---|---|---|---|---|
| D2VB Power (ours) | Points, conference hierarchy, recency, last season | 0.436 | 37% | 78.9% | 0.464 · 79% | — |
| Our old Power | Points, shrink to the D2 average | 0.486 | 30% | 77.2% | 0.552 · 75% | ours better, p < .001 |
| Pablo-style | Points (capped), recency | 0.489 | 30% | 76.5% | 0.542 · 75% | ours better, p < .001 |
| Sets only | Sets won and lost | 0.517 | 25% | 75.8% | 0.576 · 73% | ours better, p < .001 |
| Massey (1997 least squares) | Point margin | 0.523 | 24% | 77.1% | 0.617 · 75% | ours better, p < .001 |
| Wins only | Wins and losses (the RPI/NPI/Colley family) | 0.530 | 24% | 73.1% | 0.571 · 70% | ours better, p < .001 |
| Elo with margin (538-style) | Wins, scaled by margin, game by game | 0.535 | 23% | 73.3% | 0.563 · 72% | ours better, p < .001 |
2,085 matches between D2 teams, August 20, 2026 to October 7, 2026, each predicted from only the matches before it. Lower log loss is better; a coin flip scores 0.693. Cross-conference and NCAA tournament columns: log loss · % picked. "vs. ours": a paired test of the log losses on the same matches.
Rankings without win chances
RPI, the NCAA's NPI and the coaches' poll only rank teams, so the test is simpler: did the better-ranked team win? Ours is scored on exactly the same matches.
| Ranking | Matches | Its higher-ranked team won | Ours, same matches | Where they disagreed | Significance |
|---|---|---|---|---|---|
| RPI (classic 25/50/25) Every match between D2 teams, ranked from results before that day. | 2,085 | 71.6% | 78.7% | ours right 249, theirs right 101 | p < .001 |
| NCAA Power Index (our calculation) Every match between D2 teams, ranked from results before that day. | 2,078 | 72.8% | 78.7% | ours right 229, theirs right 106 | p < .001 |
| AVCA coaches poll Matches with at least one ranked team, using the poll out before the match (unranked counts below every ranked team). | 270 | 85.2% | 87.4% | ours right 17, theirs right 11 | p = .345 (not yet significant) |
What each ingredient adds
Starting from our old model and adding one piece at a time.
| Model | Log loss | Picked | Cross-conference log loss |
|---|---|---|---|
| Our old Power | 0.486 | 77.2% | 0.552 |
| + conference hierarchy | 0.466 | 77.5% | 0.520 |
| + recency (within conference) | 0.464 | 77.6% | 0.518 |
| + robust weights | 0.463 | 77.6% | 0.516 |
| + last season as a starting point | 0.444 | 78.5% | 0.481 |
| + calibrated win chances | 0.439 | 78.5% | 0.469 |
| + serve and tempo adjustment | 0.436 | 78.9% | 0.464 |
Are the win chances honest?
When we call a team a 75% favorite, it should win about 75% of the time. The top bar is what we said, the bottom bar what happened.
| We gave the favorite | Matches | Average chance | Favorite won | |
|---|---|---|---|---|
| 50–60% | 317 | 55.1% | 53.0% | |
| 60–70% | 293 | 65.0% | 62.8% | |
| 70–80% | 309 | 75.0% | 76.1% | |
| 80–90% | 439 | 85.1% | 83.4% | |
| 90–100% | 727 | 96.2% | 95.2% |
Month by month
| Month | Matches | Ours | Our old Power | Wins only |
|---|---|---|---|---|
| August | 240 | 0.514 · 77% | 0.575 · 75% | 0.594 · 70% |
| September | 1,513 | 0.424 · 80% | 0.480 · 78% | 0.528 · 73% |
| October | 332 | 0.435 · 76% | 0.447 · 76% | 0.496 · 76% |
Every system, predicting the 2025 season
| System | Uses | Log loss | Better than a coin flip | Picked the winner | Cross-conference | NCAA tournament | vs. ours |
|---|---|---|---|---|---|---|---|
| D2VB Power (ours) | Points, conference hierarchy, recency, last season | 0.415 | 40% | 80.2% | 0.432 · 79% | 0.490 · 75% | — |
| Our old Power | Points, shrink to the D2 average | 0.441 | 36% | 79.0% | 0.496 · 76% | 0.489 · 75% | ours better, p < .001 |
| Pablo-style | Points (capped), recency | 0.442 | 36% | 79.1% | 0.487 · 76% | 0.489 · 75% | ours better, p < .001 |
| Massey (1997 least squares) | Point margin | 0.460 | 34% | 79.1% | 0.540 · 76% | 0.489 · 75% | ours better, p < .001 |
| Sets only | Sets won and lost | 0.460 | 34% | 78.0% | 0.506 · 76% | 0.514 · 75% | ours better, p < .001 |
| Wins only | Wins and losses (the RPI/NPI/Colley family) | 0.483 | 30% | 76.2% | 0.522 · 74% | 0.542 · 71% | ours better, p < .001 |
| Elo with margin (538-style) | Wins, scaled by margin, game by game | 0.493 | 29% | 76.4% | 0.533 · 75% | 0.537 · 71% | ours better, p < .001 |
3,668 matches between D2 teams, August 28, 2025 to December 12, 2025, each predicted from only the matches before it. Lower log loss is better; a coin flip scores 0.693. Cross-conference and NCAA tournament columns: log loss · % picked. "vs. ours": a paired test of the log losses on the same matches.
Rankings without win chances
RPI, the NCAA's NPI and the coaches' poll only rank teams, so the test is simpler: did the better-ranked team win? Ours is scored on exactly the same matches.
| Ranking | Matches | Its higher-ranked team won | Ours, same matches | Where they disagreed | Significance |
|---|---|---|---|---|---|
| RPI (classic 25/50/25) Every match between D2 teams, ranked from results before that day. | 3,668 | 74.0% | 79.0% | ours right 389, theirs right 205 | p < .001 |
| NCAA Power Index (our calculation) Every match between D2 teams, ranked from results before that day. | 3,658 | 75.5% | 79.0% | ours right 348, theirs right 219 | p < .001 |
| AVCA coaches poll Matches with at least one ranked team, using the poll out before the match (unranked counts below every ranked team). | 544 | 83.3% | 84.2% | ours right 18, theirs right 13 | p = .473 (not yet significant) |
What each ingredient adds
Starting from our old model and adding one piece at a time.
| Model | Log loss | Picked | Cross-conference log loss |
|---|---|---|---|
| Our old Power | 0.441 | 79.0% | 0.496 |
| + conference hierarchy | 0.427 | 79.7% | 0.459 |
| + recency (within conference) | 0.425 | 79.7% | 0.460 |
| + robust weights | 0.425 | 79.7% | 0.460 |
| + last season as a starting point | 0.418 | 80.1% | 0.443 |
| + calibrated win chances | 0.417 | 80.1% | 0.437 |
| + serve and tempo adjustment | 0.415 | 80.2% | 0.432 |
Are the win chances honest?
When we call a team a 75% favorite, it should win about 75% of the time. The top bar is what we said, the bottom bar what happened.
| We gave the favorite | Matches | Average chance | Favorite won | |
|---|---|---|---|---|
| 50–60% | 533 | 55.0% | 56.7% | |
| 60–70% | 561 | 65.0% | 64.9% | |
| 70–80% | 610 | 75.1% | 74.1% | |
| 80–90% | 700 | 84.8% | 84.7% | |
| 90–100% | 1,264 | 96.1% | 97.2% |
Month by month
| Month | Matches | Ours | Our old Power | Wins only |
|---|---|---|---|---|
| September | 1,408 | 0.423 · 79% | 0.481 · 77% | 0.522 · 73% |
| October | 1,385 | 0.403 · 82% | 0.412 · 81% | 0.461 · 78% |
| November | 810 | 0.419 · 80% | 0.420 · 79% | 0.444 · 78% |
| December | 60 | 0.477 · 77% | 0.472 · 77% | 0.563 · 73% |
Every system, predicting the 2024 season
| System | Uses | Log loss | Better than a coin flip | Picked the winner | Cross-conference | NCAA tournament | vs. ours |
|---|---|---|---|---|---|---|---|
| D2VB Power (ours) | Points, conference hierarchy, recency, last season | 0.442 | 36% | 78.7% | 0.486 · 77% | 0.494 · 72% | — |
| Pablo-style | Points (capped), recency | 0.460 | 34% | 77.6% | 0.522 · 74% | 0.492 · 73% | ours better, p < .001 |
| Our old Power | Points, shrink to the D2 average | 0.462 | 33% | 77.9% | 0.527 · 75% | 0.491 · 75% | ours better, p < .001 |
| Massey (1997 least squares) | Point margin | 0.479 | 31% | 77.7% | 0.566 · 75% | 0.494 · 74% | ours better, p < .001 |
| Sets only | Sets won and lost | 0.484 | 30% | 76.5% | 0.568 · 72% | 0.521 · 72% | ours better, p < .001 |
| Wins only | Wins and losses (the RPI/NPI/Colley family) | 0.499 | 28% | 75.1% | 0.562 · 70% | 0.516 · 73% | ours better, p < .001 |
| Elo with margin (538-style) | Wins, scaled by margin, game by game | 0.504 | 27% | 74.4% | 0.567 · 71% | 0.509 · 74% | ours better, p < .001 |
3,501 matches between D2 teams, August 25, 2024 to December 12, 2024, each predicted from only the matches before it. Lower log loss is better; a coin flip scores 0.693. Cross-conference and NCAA tournament columns: log loss · % picked. "vs. ours": a paired test of the log losses on the same matches.
Rankings without win chances
RPI, the NCAA's NPI and the coaches' poll only rank teams, so the test is simpler: did the better-ranked team win? Ours is scored on exactly the same matches.
| Ranking | Matches | Its higher-ranked team won | Ours, same matches | Where they disagreed | Significance |
|---|---|---|---|---|---|
| RPI (classic 25/50/25) Every match between D2 teams, ranked from results before that day. | 3,501 | 74.2% | 77.6% | ours right 317, theirs right 198 | p < .001 |
| NCAA Power Index (our calculation) Every match between D2 teams, ranked from results before that day. | 3,485 | 75.6% | 77.7% | ours right 269, theirs right 198 | p = .001 |
| AVCA coaches poll Matches with at least one ranked team, using the poll out before the match (unranked counts below every ranked team). | 518 | 85.3% | 85.3% | ours right 19, theirs right 19 | p = 1.000 (not yet significant) |
What each ingredient adds
Starting from our old model and adding one piece at a time. (Without a previous season on file, the last step can't be tested here.)
| Model | Log loss | Picked | Cross-conference log loss |
|---|---|---|---|
| Our old Power | 0.462 | 77.9% | 0.527 |
| + conference hierarchy | 0.449 | 78.7% | 0.502 |
| + recency (within conference) | 0.447 | 78.7% | 0.503 |
| + robust weights | 0.447 | 78.5% | 0.503 |
| + last season as a starting point | Not tested: no previous season | ||
| + calibrated win chances | 0.444 | 78.5% | 0.491 |
| + serve and tempo adjustment | 0.442 | 78.7% | 0.486 |
Are the win chances honest?
When we call a team a 75% favorite, it should win about 75% of the time. The top bar is what we said, the bottom bar what happened.
| We gave the favorite | Matches | Average chance | Favorite won | |
|---|---|---|---|---|
| 50–60% | 577 | 54.9% | 53.6% | |
| 60–70% | 561 | 65.1% | 64.2% | |
| 70–80% | 600 | 75.2% | 78.0% | |
| 80–90% | 674 | 85.3% | 85.6% | |
| 90–100% | 1,089 | 96.1% | 95.5% |
Month by month
| Month | Matches | Ours | Our old Power | Wins only |
|---|---|---|---|---|
| September | 1,290 | 0.486 · 77% | 0.536 · 75% | 0.558 · 71% |
| October | 1,278 | 0.425 · 79% | 0.432 · 78% | 0.487 · 77% |
| November | 874 | 0.396 · 81% | 0.393 · 82% | 0.429 · 79% |
| December | 59 | 0.502 · 80% | 0.502 · 80% | 0.524 · 78% |
This season's results update every day as matches are played (last run Oct 8, 4:56 AM CT).
The systems
| System | What it uses | Strength of schedule | Blowouts | How we tested it |
|---|---|---|---|---|
| D2VB Power (ours) | Rallies won, venue, date | Solved jointly, with conference strength measured from cross-conference play | Surprising results count a little less | Directly |
| RPI | Wins and losses | Opponents' win % (50%) and their opponents' (25%) | Ignored | Directly (classic 25/50/25) |
| NCAA Power Index (NPI) | Wins, losses, venue | Average opponent NPI (75%), bonus for good wins | Ignored | Our calculation of the published formula (used for projections) |
| AVCA coaches poll | Coaches' votes | Voters' judgment | Voters' judgment | Directly, from each week's poll |
| Pablo (RichKern.com) | Point share, venue, date | Solved jointly (least squares) | Capped at about 59% of points | A replica built from its published method; Pablo's site doesn't allow automated access, so we can't test the live ratings |
| Massey | Point margin | Solved jointly (least squares) | Counted in full (1997 method) | Directly (his original method) |
| Elo (FiveThirtyEight style) | Wins, scaled by margin | Game by game | Diminishing (log of margin) | Directly |
| Sets only | Sets won and lost | Solved jointly | A 3–0 is a 3–0 | Directly |
| Wins only (Colley/RPI information) | Wins and losses | Solved jointly | Ignored | Directly |
| KenPom / T-Rank | Points per possession (basketball) | Solved jointly | Expected mismatches count less | Their volleyball equivalent, a separate sideout and serving rating per team, predicted worse than one rating (below); their mismatch weighting is in ours |
How the Power rating works
- The evidence. For every match, the log of the ratio of rallies won (49 to 38 is ln(49/38) = 0.25). Weighted by how many rallies were played.
- The model. That number ≈ rating(home) − rating(away) + a home edge, fitted across all matches at once so every result is judged against the strength of who it came against.
- The hierarchy. Rating = conference strength + standing within the conference. A team with little evidence is assumed to be like its conference, not like the average D2 team. Conference strength comes from cross-conference matches, which hold full weight all season.
- Recency. A team's standing within its conference weighs a match half as much after 30 days.
- Robust weights. A result far from what was expected (a surprise blowout) counts a little less, so one strange night can't swing a team.
- Last season. 90% of each conference's strength and 70% of each team's standing in it carry over, worth about one match of evidence, and fade as this season's matches come in.
- Serve and tempo adjustment. Not every point repeats. Teams that win points with service aces, or by forcing passing errors, do a little worse next time than their point totals suggest; teams that get aced a lot also do a little worse; teams with more attacks per set (long rallies, defense into transition) and more setter assists per set do a little better. Each team's season-to-date numbers adjust its rating slightly. We found it in 2025 and 2026 and confirmed it on 2024, which played no part in finding it; each season below is scored with weights trained without it.
- Win chances. Rating differences give the chance of winning a rally; the scoring rules (sets to 25, a fifth to 15, win by 2) turn that into set and match chances. The ± on the rankings page is one standard deviation: teams whose ranges overlap are close to even.
Résumé
For each team: how likely would the 25th-best team be to win at least as many matches against the same opponents, at the same sites? The less likely, the more impressive the record, and the higher the Résumé rank. Only wins and losses count; the power ratings judge how hard each match was. It's the same idea as ESPN's Strength of Record and Bart Torvik's Wins Above Bubble, which the NCAA's basketball committee now uses.
What we tried that didn't help
- Sideout and serving as separate ratings (the volleyball version of KenPom's offense and defense). It predicted worse than one rating: two numbers per team means twice as much to estimate from the same matches, and sets are decided almost entirely by the overall share of rallies won.
- Counting wins or sets on top of points. Predicted worse in every version we tried.
- Capping blowouts outright. Worse than down-weighting surprises.
- Recency on everything. Fading old matches across the board weakens conference strength, because the matches that measure it are played early.
- Travel, time zones, start times and rest. Over three seasons, how far a team traveled, which way, how many time zones, its body clock and days of rest made no difference beyond what the rating already expects.
- "Clutch." Winning rallies late in close sets, close-set records and recent form didn't carry over from match to match.
- Box-score stats in bulk. Of 31 team and position stats, only the serve and tempo ones above added anything; the rest are already in the points.
Keeping it honest
- This season's test reruns every day. If the numbers slip, you'll see it here.
- Every change to the model is tested the same way before it ships, and logged.
- Our match predictions, saved before each match, are graded on the accuracy page.
Sources: Pablo FAQ · Massey (1997) · NCAA D2 NPI overview · KenPom methodology · LRMC (Kvam & Sokol) · Ranking and schedule connectivity (Osting et al.) · RPI