Research log
Every question we've tested, how we tested it, what we found and what we decided, including the ideas that didn't hold up.
October 7, 2026
Advanced stats and outside factors: do they add to the rating?
Question. vballr.com shows deeper D1 stats (first-ball vs transition, pass-to-kill, clutch). Do any stats, or travel, time zones, start time, rest and schedule, predict results beyond our odds?
How we tested it. Surveyed vballr (D1; contact-by-contact play-by-play from school live stats; composite rankings of AVCA, RPI, Evollve, Massey, KPI and its own). D2 play-by-play records only how each rally ended, so first-ball/transition and pass- or dig-to-kill are not computable for D2. Built 9,205 D2 matches (2024-2026) with walk-forward odds, venue cities geocoded, time zones, local start times, rest and schedule density; and 31 season-to-date team and position stats for 5,177 matches (2025-2026). Each factor tested as a logistic regression with our odds as a fixed offset; combined models trained on one season and tested on the other.
What we found. No effect beyond our odds: travel distance, direction, time zones, body clock, rest, back-to-back days. Small, real context effects: more upsets than expected in non-conference matches, at neutral sites and in a team's second match of the day; fewer in conference play; correcting for them improved held-out log loss only slightly (2024 -.0018, 2025 -.0001, 2026 -.0012). Our model gives the listed home team a home edge at neutral sites that is not there (about 4 points). Of 31 stats, four were consistent in both seasons: attacks per set (+), aces per set (-), opponents' reception error % (-), own reception error % (-) and setters' assists per set (+). All 31 together did not help out of sample; the four together cut log loss by about .002 (2026 to 2025 p = .001; 2025 to 2026 p = .12), but they were chosen with both seasons, so a clean 2024 holdout is pending. Clutch play, close-set record, form, inconsistency and depth of attack showed nothing.
Decision. Rating unchanged for now. Next: 2024 holdout for the four-stat adjustment; test a zero home edge at detected neutral sites inside the rating. Also fixed the bootstrap random-number generator (uneven draws); every published comparison still holds.
October 6, 2026
Three seasons of testing, calibration, and a public methodology page
Question. Does the new power rating hold up across seasons, against the RPI, NPI and the coaches poll, and are its win chances honest?
How we tested it. Backfilled 2024 (4,107 matches) and the 2024 and 2025 AVCA polls. scripts/backtest.mjs replays a season day by day for seven win-chance systems (ours, old, Pablo replica, Massey 1997, Elo, sets-only, wins-only) and three rankings (RPI, NPI, AVCA poll), with ingredient-by-ingredient tests, calibration, paired bootstrap and McNemar tests.
What we found. Over 9,205 matches (2024 to 2026 so far): ours log loss .432 and 79.2% picked; Pablo-style .459, old .459, Massey .482, sets-only .482, wins-only .500, Elo .507; ours significantly better than each in every season. Picking winners: RPI 71.6 to 74.2% and NPI 72.9 to 75.6% against ours 78.0 to 79.5% (p < .001 each season); AVCA poll behind ours every season (85.3 vs 86.1, 83.3 vs 84.4, 84.9 vs 87.5) but not yet significant. Win chances ran slightly overconfident; tempering rating edges by 1.15, chosen on 2024, improved 2025 and 2026.
Decision. Shipped the calibration (RALLY_TEMP). methodology.html shows every season and the pooled results; the current season reruns daily on the server.
October 6, 2026
A better power rating: conference hierarchy, recency, last season
Question. Is points-for-points the right basis, and can we beat it? Fans questioned rankings like a 17-1 Angelo State at 12th.
How we tested it. Surveyed RPI, NPI, Massey, Pablo, KenPom/T-Rank, Sagarin, Elo/Glicko, Colley, Keener, LRMC and Bradley-Terry with shrinkage (research/power-rankings.md). Backfilled 2025 set scores (4,125 matches). Replayed 2025 and 2026 day by day, predicting each day only from earlier matches, scored by log loss with paired bootstrap tests; also cross-conference, late-season and NCAA-tournament subsets, and how often each ranking agrees with results already played.
What we found. Margins beat wins: on 2025, log loss .4246 for our points model vs .4504 sets-only, .4749 wins-only (the RPI/NPI family) and .4858 Elo. Shrinking teams toward their own conference instead of the D2 average, recency within conference (30-day half-life) with conference strength kept at full weight, and Huber weights cut log loss to .4159 (p < .001), cross-conference .4407 to .4257. Last season as a prior (90% of conference strength, 70% of standing within it) cut opening-week-on log loss on 2026 from .476 to .437. A sideout/serve split (KenPom-style offense and defense) predicted worse than one rating. Agreement with past results is unchanged (83.9% vs 84.0%).
Decision. Power now runs on the new model (src/rating.mjs, prior in src/power-prior.json, regenerated each new season with scripts/power-prior.mjs). Résumé is now strength of record (how hard the record would be for a top-25 team against the same schedule); NPI keeps its own column. Ranks 5 to 25 sit within each other's margin of error, now shown as ± on Rally %.
October 3, 2026
Can betting markets help?
Question. Do sportsbooks' lines or methods offer anything for D2 volleyball?
How we tested it. Searched for D2 women's volleyball lines and for how volleyball is modeled for betting.
What we found. No sportsbook posts lines on D2 women's volleyball, and books' terms forbid scraping. Their method, rally-by-rally probability built up to sets and matches, is what our power rating already does.
Decision. Borrow methods and test each one (above), never lines.
October 3, 2026
Do service errors cost weak-passing teams more?
Question. A team that passes poorly is in serve receive more often after every missed serve. Does that show in results?
How we tested it. Split teams by season sideout % (bottom vs top third) and compared fewer vs more service errors within the same aces band.
What we found. Not in NCAA data: weak-receiving teams with more errors won about as often (or slightly more). Likely hidden, because tough serving forces bad passes that NCAA stats don't record.
Decision. Not used as a lever with NCAA data alone. A question for coaches' own charting (pass ratings).
October 3, 2026
Serving as a combination: aces and errors together
Question. Is a team that misses serves AND rarely aces the real problem?
How we tested it. 4,206 team-matches grouped by aces and errors against the D2 middle (1.6 aces, 2.0 errors per set); aces per error; point-scoring % (points won on own serve).
What we found. More aces and more errors won 69%; more aces and fewer errors 64%; fewer aces and more errors 35%; fewer aces and fewer errors 29%. Low aces is the problem either way. Aces per error under 0.4 won 22%, over 1.5 won 69%. Point-scoring % of 50%+ won 84%, under 35% won 12%: the best single serving number.
Decision. Added point-scoring %, aces per error and the serving profile as levers and to the writer's guidance.
October 3, 2026
Are fewer service errors better?
Question. Reports told teams to cut service errors. Does that help win?
How we tested it. Same 2,103 matches: win rate for the team with fewer service errors; win rate by errors per set; effect with aces held equal.
What we found. The team with fewer service errors won only 43%. Teams that miss more also ace more: aggressive serving wins. With aces held equal, extra errors cost a little, much less than aces gain.
Decision. Reports never advise simply cutting service errors; serving targets are set on aces, aces per error or point-scoring %.
October 3, 2026
What decides D2 matches?
Question. Which in-match numbers separate winners from losers?
How we tested it. 2,103 D2-vs-D2 matches: how often the team that wins each stat wins the match, win rate by a team's own level, and a logistic regression on the differences so each stat is measured with the others held equal.
What we found. The team that hits better wins 91% of matches; hitting .250+ wins 88%, .300+ 97%, under .150 under 20%. More aces wins 78%. With everything held equal, hitting % dominates, then aces; blocks and passing add little once those are known. The six stats together call the winner about 93% of the time.
Decision. Reports now build their game plans and targets around hitting %, the hitting % allowed, serving and passing, with the evidence given to the writer every build.
October 3, 2026
Model experiments: caution, recency, travel
Question. Do changes oddsmakers use improve our win chances on matches the model hasn't seen?
How we tested it. Walk-forward test on 1,806 D2 matches: fit each day on earlier results, predict that day's matches, score accuracy and Brier.
What we found. Trusting new results faster (rating prior 75 instead of 150) improved every period: 77.7% right vs 76.9%, Brier 0.1579 vs 0.1590, 81.6% since Sept. 15. More early caution (prior 300–1,200) made it worse. Weighting recent matches more (21–60 day half-life) made no difference. A penalty for the second day of a road back-to-back helped by a hair (within noise).
Decision. Adopted prior 75. Rejected the rest.
October 3, 2026
Should we shade heavy favorites, like a sportsbook?
Question. Books pull extreme odds toward the middle. Would that fix our top-end overconfidence?
How we tested it. Fitted a shading factor on the first 60% of the season and tested it on the last 40% it hadn't seen (Brier score, lower is better).
What we found. Worse: Brier 0.1370 as is vs 0.1393 shaded. Since mid-September our 94% favorites win 94%; the overconfidence was an early-season effect, not a constant one.
Decision. Rejected. The real fix is a better starting point for each season (last season's strength), to be tested.
October 3, 2026
The season backtest: are our predictions honest?
Question. Across every D2 match this season, using only data from before each match, how often does the favorite win, and do our ranges hold as often as they claim?
How we tested it. For 1,844 D2-vs-D2 matches with box scores: the power rating was refit for each day on earlier results only; each team's ranges came from its own earlier matches (70% prediction intervals, five or more matches).
What we found. The favorite won 77% of matches (81% since Sept. 15). 67.5% of about 21,000 70% ranges held, close to the 70% target. Win chances were well calibrated in the middle and slightly overconfident at the top (said 96%, won 93%), almost entirely in the first weeks of the season.
Decision. Publish the scorecard and pre-register predictions for every upcoming match from Oct. 3 on, so results can't be adjusted after the fact.
October 3, 2026
Does the scouting picture hold up? A report against its match
Question. We wrote a scouting report before Angelo State at UT Dallas (Oct. 2). How much of it held when the match was played?
How we tested it. Every checkable claim in the report was compared with the box score and play-by-play: the prediction, the opponent's tendencies, our own team's tendencies and the three key targets.
What we found. Angelo State won 3–1 (87% favorite; 3–1 was the second most likely score). Expected rally share 54.4%, actual 55.5%. About 12 of 17 checkable tendencies held. The ones about who gets the ball, team passing rates and sideout rates held; single-match efficiency and serve-receive targets were the misses (the 'steady' passers gave up 5 of 6 aces).
Decision. Built the automatic accuracy model so this check runs on every match, not one.
October 3, 2026
A second report against its match: Angelo State at UT Tyler
Question. Our family and coach reports before Angelo State at UT Tyler (Oct. 3): how much held when the match was played?
How we tested it. Every checkable claim in both reports was compared with the box score: the prediction, the three targets in each edition, team tendencies and the player-level calls.
What we found. Angelo State won 3–1 (72% favorite; 3–1 was the most likely score at 28%). Expected rally share 52.3%, actual 55.0%. Of the three targets, the hitting % target held (.286, against .163 for UT Tyler) and decided the match; the service-error target missed (11 errors to 5) and Angelo State won anyway; the aces target narrowly missed (1.75 per set, still 7 aces to 3). Team-level calls held: Angelo State won the sideout battle (69.2% to 57.0%), UT Tyler's block was big (2.75 per set), Angelo State's passing neutralized UT Tyler's serving (3 aces), and the hitter we said to let swing hit .107 on a team-high 28 swings. Player-level calls missed most: the hitter we called the one to respect took 4 swings, the middle we called a weak spot hit .529, and a bench hitter took 22 swings. As against UT Dallas, the 'steady' libero was aced at a higher rate than the passers we targeted.
Decision. Confirms hitting % as the deciding lever and dropping service-error targets. Player-level calls are the weakest part of a report, mostly because NCAA data can't see lineup changes or injuries; reports should lean on team-level levers and hedge calls about individual players.
October 5, 2026
Should wins count in the power rating?
Question. The power rating had 17–1 Nebraska-Kearney (89th-toughest schedule) above 16–0 Ferris State (20th), and 13–3 Colorado Mesa above 17–1 Angelo State. Would counting match and set wins, or capping blowouts, make it better?
How we tested it. Replayed 1,196 D2 matches from Sept. 10 on, predicting each one using only results from before it: the points-only rating, the same blended 25–50% with a win-based or set-based rating, and versions that cap how lopsided one match can count.
What we found. Points only predicted best: 80.5% right, Brier .1343. Adding match wins: .1378; set wins: .1368; both: .1397. Capping blowouts put Ferris State first but also predicted worse at every level (.1356 to .1443). Teams that beat weak opponents badly really are the better bet next time; teams that win close matches are a little less so.
Decision. The power rating stays points-based, for predictions and simulations. The rankings page now also shows a Résumé rank, our NCAA Power Index (wins weighted by opponent and site, the formula that picks the tournament), with a switch between the two views: Ferris State is 1st on résumé, Angelo State 3rd.
October 6, 2026
CORA in D1 and D3: the same test, the same result
Question. Does the rating built for D2 hold up in Division I and Division III, against each division's own NCAA formula (the RPI in D1, the NPI in D3) and the AVCA polls?
How we tested it. Downloaded 2024, 2025 and 2026-to-date D1 and D3 results and set scores from the NCAA feed (about 32,000 matches) and every weekly AVCA poll. Membership: 10+ matches in the division's feed. RPI from division matches only (checks against the NCAA's published 2026 D1 RPI at rank correlation .999). D3 NPI with the committee's published settings (20/80, home/away 1.0/1.0, QWB 55.5 x .6, 10 wins), within 0.2 of the NCAA's published 2025 values. Same walk-forward replay as D2; 2025 and 2026 start from the previous season. Found and fixed along the way: conference strengths stopped after six rounds and had not settled where conferences rarely meet (up to 0.044 off in D1, 0.086 in D3); they now iterate with a block shift per conference until settled (within 0.002 of a tight solution). D2 results unchanged to the third decimal.
What we found. D1, 2024 to 2026 (11,992 matches): CORA log loss .462 and 77.8% picked; Pablo-style .487, Massey .510, sets-only .513, wins-only .532, Elo .540; CORA significantly better than each in every season. RPI: 71.9% vs CORA 77.2% on the same matches; where they disagreed CORA was right 1,310 times to 673 (66%), p < .001. AVCA poll: 85.3% vs 86.0%, not significant. D3 (13,318 matches): CORA .396 and 81.7%; Pablo-style .435, Massey .449, wins-only .489; NPI 75.9% vs CORA 81.4%, disagreements 1,361 to 630 (68%), p < .001; AVCA 83.8% vs 84.7%, not significant. Across divisions and seasons CORA makes 16 to 21% fewer wrong calls than the NCAA formula.
Decision. Built /d1/ and /d3/ (CORA top 50 plus the public test), refreshed by scripts/divisions.mjs on the server. Off in production until DIVISIONS=d1,d3 is set on Railway (Bo's approval). Public claims: decisively better than the RPI and NPI; level with or slightly ahead of the coaches poll, not significantly.