Using Graph Theory To Predict NCAA Tournament Outcomes

← Back to Stories (view on slashdot.org)

Using Graph Theory To Predict NCAA Tournament Outcomes

Posted by timothy on Tuesday March 13, 2012 @12:49AM from the more-interesting-than-what-is-at-heart-portrayed dept.

New submitter SocratesJedi writes "Like many technically-minded people, I don't have a lot of time to keep up with sports. Nevertheless, trying to predict the outcome of the NCAA men's basketball tournament is a fun activity to share with friends, family and colleagues. This year, I abandoned my usual strategy of quasi-randomly choosing teams and instead modeled the win-loss history of all Division I teams as a weighted network. The network included information from 5242 games played during the 2011-2012 season. From this, teams came be ranked using tools from graph theory and those rankings can be used to predict tournament outcomes. Without any a priori information, this method accurately identified all the #1 seeds in the top 5 best teams. It also predicts that at least one underdog, Belmont (#14 seed), will reach the Elite Eight. Although the ultimate test will be how well it predicts tournament outcomes, initial benchmarks suggest 70-80% accuracy would not be unreasonable."

7 of 91 comments (clear)

Min score:

Reason:

Sort:

past history by Collin · 2012-03-13 01:01 · Score: 5, Insightful

wouldn't running the algorithm against past years' records and testing against past tournament results be the best possible test to tune the algorithm?
1. Re:past history by PatDev · 2012-03-13 01:35 · Score: 5, Insightful
  
  I worked in a research group in college that worked on exactly this problem - predicting NCAA tournaments with a graph-theoretic approach. That is exactly how you test the algorithm. And the cited estimate of 70-80% accuracy seems made up. People who research the field know that there is far less certainty than that. At something like 20% confidence, your prediction should be something like 20%-90%.
  
  The problem stems from the fact that we traditionally predict a team will win if it is a stronger or better team, and we use our graph theory to produce relative team ratings. And if each game of the tournament were played over and over again with the winner of the majority going to the next round, then our methods would work even better. As it stands though, we are trying to predict a single sampling from a probability distribution - which will necessarily have error. Informally, the real tournament has upsets (when a weaker team beats a stronger one). Our algorithms can't predict these, the best they can do is gain a better understanding than humans as to which team is better.
  
  Add to that the fact that the tournament is structured hierarchically - a mis-prediction in the first round prevents you from even attempting to predict later games (and by NCAA bracket scoring, that counts the same as mis-predicting those later games). So early upsets can potentially have large negative outcomes on brackets.
Predicting the top is easy by elrous0 · 2012-03-13 01:01 · Score: 4, Insightful

Everyone knows who the big names are who are likely to make it to the final four. It's predicting how things will go at the middle and bottom, where teams are much more likely to be evenly matched, that's really hard.

--
SJW: Someone who has run out of real oppression, and has to fake it.
Re:70-80%? by MyLongNickName · 2012-03-13 01:20 · Score: 4, Informative

And my numbers are off. In 2011, 43 times out of 63, the lower seed won for about a 68% win rate.

--
See my journal for slashdot ID's by year. Mine created in 2005. http://slashdot.org/journal/289875/slashdot-ids-by-year
Re:Just take last years results by JayBean · 2012-03-13 01:22 · Score: 5, Insightful

That may work for pro sports, but not for college sports. In fact, because teams usually lose their nucleus after winning it all (players declare for the draft), it is rare for a team to make it to the final game two or more years in a row.
As a sports fan by jayhawk88 · 2012-03-13 01:23 · Score: 3, Interesting

Some problems I see. Disclaimer: I know there's a margin of error here as the author said, and I know my observations will be based largely on anecdotal evidence, making it inferior. But if sports were so easy to predict there would be no sports gambling.
- That's probably too far for Belmont; a #14 has only ever gotten as far as the Sweet 16, twice (Cleveland State '86, Chattanooga '97). Lowest seed to make an Elite 8 is Missouri in 2002 as a #12 . Belmont is actually going to be one of the more popular upset picks, but they would have to upset two far superior teams twice in 3 days.
- It's a bit too "chalk". #1 seeds generally survive the first two games (undefeated against #16's, 55-14 v. #8's, 59-6 v. #9's), but the #2's have it worse (only four losses v. #15's, but 58-21 v. #7's and 29-21 v. #10's). I know two #12's, a #13 and a #14 doesn't seem like "chalk" but historically it's much more likely that we'll see more #5-7 or #10-11's. To have only one #2 not make the Elite 8 and all the #1's would be almost unheard of.
- A #12 always beats a #5, but three of them doing so in one year would seem unlikely, as they're only 39-89 overall.
- Some of the other first round matchups seem a bit improbably. It has every #6 and every #7 winning, for example.
Not enough time? by babyrat · 2012-03-13 02:36 · Score: 3, Insightful

You don't have time to follow sports, but you have time to model "information from 5242 games played during the 2011-2012 season".
You could be honest and just say you don't really care, but get involved in the playoffs because everyone else is talking about it.
I'm guessing your level 80 warlock probably doesn't 'have time' either. :)