OKC Thunder are facing a historically difficult playoff schedule

Has the league ever been as top heavy as 2016?

The top team in the league, the Golden State Warriors, became the greatest regular season team of all time by winning a record 73 games. The San Antonio Spurs won 67 games but were actually a shade above the Warriors in net rating. By SRS (net rating adjusted for opponents), these were the two best teams since Jordan’s Bulls. In the East you can’t forget about the Cleveland Cavaliers, who rolled through the Eastern Conference and are led by a guy named LeBron James, who just reached the NBA finals for the 6th consecutive time.

Then you have the Oklahoma City Thunder. Despite sporting a robust-but-not-historic SRS of 7.09 and having two of the top 5 players in the league, it was easy to overlook the Thunder this season due to the caliber of their opponents. And yet as of writing this, Oklahoma City is just one game away from winning the West and facing the Cavaliers in the NBA Finals. To say that the Thunder have been dealt a difficult hand in these playoffs is an understatement; if they indeed go on to win the championship this year, they would have knocked off the three aforementioned teams in addition to the 6th seeded Dallas Mavericks. But do the Thunder have the most difficult road to a championship of all time? I scraped all team SRS ratings from basketball-reference.com over the past 22 years to find out.

As it turns out, if OKC can knock off the Warriors tonight in Oakland, their playoff slate would be the toughest of any team since at least 1994, among teams that faced 4 playoff opponents (i.e. reached the Finals):

Most difficult playoff opponents, past 22 years
Team Season Seed Result Team SRS Average Opponent SRS
OKC 2016 3 ? 7.09 6.52
LAL 2008 1 L, Finals 7.34 6.25
HOU 1995 6 W, Finals 2.32 5.99
ORL 2009 3 L, Finals 6.48 5.85
LAL 2002 3 W, Finals 7.15 5.70

If the Thunder end up defeating two historically great teams in the Spurs and Warriors, and then face the Cavaliers, their average opponent SRS in the playoffs would be the highest in at least 22 years. The 2008 Lakers would come close after traversing through a tough Western Conference and then falling to an elite Celtics team in the finals, and the 1995 Houston Rockets faced an incredibly difficult slate winning the title as a 6 seed.

Since the Thunder are a great team themselves, their run to a title would not be the most unlikely title run in history. By my calculations – using Team and Opponent SRS to predict the likelihood of a series win – a Thunder championship in 2016 would be the 6th most surprising postseason run over the past 22 seasons:

Most unlikely playoff runs, past 22 years
Team Season Seed Team SRS Average Opponent
SRS
Expected Series
Wins
Series Wins Series Wins
Over Expected
HOU 1995 6 2.32 5.99 0.15 4 3.85
LAL 2001 2 3.38 5.54 0.50 4 3.50
HOU 1994 2 4.19 4.46 0.81 4 3.19
LAL 2010 1 4.72 4.23 0.94 4 3.06
DAL 2011 2 4.23 4.61 1.02 4 2.98
OKC** 2016 3 7.09 6.52 1.18 4 2.82

**If the Thunder go on to win the title

After a solid but underwhelming regular season capped off by one of the most difficult postseason runs of all time, the ’95 Rockets top this list. In fact, since they faced the far superior Utah Jazz in the first round, the Rockets would only be expected to win 0.15 series’ in 1995. That they went on to win the title is one of the greatest underdog stories in NBA history. The 2001 Lakers, despite being in the midst of a 3-peat, weren’t incredibly dominant in the regular season and actually had the lower SRS in all of their Western Conference matchups. If the Thunder join this list they’d be the best team, but also have the most difficult slate.

It’s unlikely the Thunder will defeat Golden State tonight – most projections give the Warriors about a 70% chance of winning – but if they do, they’d go down as having arguably the most difficult playoff road in history. And if Oklahoma City can pull off the upset, Kevin Durant and Russell Westbrook will be favored to defeat the Cavaliers and bring home their first NBA title, cementing their legacies as one of the best duos ever to play the game.

Does Playoff Experience Matter?

With two games left in the 2016 regular season, the Utah Jazz sat one game ahead of the Houston Rockets for the 8th and final seed in the Western Conference Playoffs. Utah hadn’t been to the postseason since 2012, and after amassing an impressive collection of young talent during its rebuilding phase, was eager to return. The Jazz ended up losing its final 2 games while the Rockets won both of theirs, sending Utah to the lottery for the 4th straight year. Due to the surprisingly strong middle-of-the-pack in the Eastern Conference the Jazz landed the 12th pick in the 2016 draft, while the Rockets landed at 15 and had the honor of getting pummeled by the Golden State Warriors; however, many fans argued that the difference in draft slot, plus the slight chance of landing a top 3 pick, is much less valuable than the playoff experience Utah’s young players could have gotten by reaching the postseason.

A major media narrative each year in the playoffs is that experience matters. For many fans and analysts, experience is justification for picking the older, wiser team – those who have “been there before” – to advance farther than the young up-and-comer. A team like Utah, even if they had made it, would certainly be overmatched due to their limited experience in a playoff setting. But does experience really matter, or is it just that the best teams tend to also be the most experienced?

I’ve come to find that in sports, many media narratives can be disproven as a myth using data, which is why I set out to test the theory on playoff experience. To do so I gathered data from basketball-reference.com from the past 15 playoffs (2001-2015). For each team I found the average prior playoff minutes played by each player, weighted by how often they play in a current season. To normalize for the fact that the best teams also tend to be the most experienced, I created a model which used each team’s SRS (net rating adjusted for opponents) and the SRS of their opponents (and presumptive opponents) to compute their expected series wins. As it turns out, the teams with the most playoff experience tend to fare better than expected:

Team Season Weighted Cumulative
Minutes
Expected Series
Wins
Series Wins Series Wins
Over Expected
LAL 2004 3605.31 0.83 3 2.17
SAS 2015 3534.23 0.79 0 -0.79
MIA 2014 3453.57 2.0 3 1.0
LAL 2011 3295.7 1.85 1 -0.85
DAL 2012 3076.75 0.16 0 -0.16
SAS 2008 3059.16 0.66 2 1.34
SAS 2014 3005.93 2.15 4 1.85
MIA 2013 2888.08 2.96 4 1.04
UTA 2001 2867.97 0.69 0 -0.69
LAL 2010 2865.95 0.94 4 3.06

Of the 10 most experienced playoff teams over the past 15 postseasons, three won the title, two made the finals and six outperformed their expected series win total. The 2004 Lakers are a great example of how having experience can help a team; though the Lakers weren’t as dominant that season as the teams that 3-peated earlier in the decade, they reached the finals by outlasting the Spurs and Kevin Garnett-led Timberwolves with an extremely playoff-tested group that included Kobe, Shaq and newly acquired Karl Malone.

On the contrary, the most inexperienced teams heading into the playoffs tend to underwhelm relative to expectations:

Team Season Weighted Cumulative
Minutes
Expected Series
Wins
Series Wins Series Wins
Over Expected
POR 2009 36.45 0.85 0 -0.85
WAS 2005 74.96 0.36 1 0.64
OKC 2010 78.87 0.5 0 -0.5
ORL 2001 102.55 0.32 0 -0.32
TOR 2007 107.42 0.78 0 -0.78
IND 2011 108.02 0.05 0 -0.05
TOR 2014 116.65 1.16 0 -1.16
GSW 2013 124.95 0.17 1 0.83
CHI 2006 143.21 0.29 0 -0.29
BOS 2002 143.89 0.82 2 1.18

Of this group, only three teams made it out of the first round and those teams were the only ones to outperform their baseline series win projections. A few teams on this list, like the 2010 Thunder, 2011 Pacers, 2014 Raptors and 2013 Warriors, went on to experience more meaningful playoff success in later years (The 2009 Portland Trail Blazers, with the young core of Greg Oden, Brandon Roy, LaMarcus Aldridge and Nicolas Batum, seemed destined for greatness had injuries not struck…one of the greatest “what-if’s” in NBA history, but I digress).

If we look at all teams over the past 15 years we can see that there is a small, but statistically significant relationship between playoff experience and playoff success:

playoff_experience

Let’s say the Jazz had snuck into the postseason and lost in 5 games to the Warriors, adding roughly 200 minutes to their weighted average playoff experience to their 2017 squad (if the majority of their players return). Based on the regression line above, I’d add  .04 series wins to Utah’s expected total next year due to this experience. Maybe that isn’t worth having a 2.5% chance at a top 3 pick, but evidence shows Utah would be slightly better off in the 2017 playoffs had they made it this year.

So even though experience isn’t the most crucial predictor of playoff success – actual team quality matters far more – it is true that teams with more experience have been more likely to outperform expectations in the poststason. The next time you hear “the more experienced team will win”, remember that it’s not just media hype.

Demographics in remaining primaries favor Bernie Sanders, not Hillary Clinton

Let me first say this: Nate Silver is one of my idols. His work on FiveThirtyEight pioneered statistical writing as a popular medium, particularly in sports; without it this blog wouldn’t exist and I might not have even been inspired to get my masters degree in Analytics.

However, I was struck by Silver’s most recent article on FiveThirtyEight, and wanted to offer a rebuttal. In the article, Silver makes the argument that the states Clinton won are more demographically similar to democrats in states which have yet to vote in primaries. Silver suggests that her success in these states favors Clinton moving forward, particularly in stats whose demographics conform closer to average. In his assertion, Silver calculated the root mean squared error between a state’s demographics and average Democratic demographics, where a low RMSE indicates that a state’s demographics are more representative of all democrats. Silver noted that of the states with the 9 lowest RMSE’s, Hillary Clinton won 8 of them.

To be clear: nothing Silver wrote in this piece is incorrect. It’s true that Clinton has outperformed Sanders in states with more “Democratic” demographics. However, implying that Clinton is in better shape because demographic favor her is disingenuous, and ignores more substantial reasons why Clinton performed the way she did. As you can see from the plot below, demographic similarity to average explains only 13% of the variance in voting outcome and the relationship is, in fact, not statistically significant1:

rmse_clinton

Even if we only look at wins and losses while ignoring margin of victory, a p-test would also suggest a non-statistically significant relationship between demographic similarity and support for Clinton2.

Really, using demographics to explain state polling behaviors is (literally) black and white. States with a large percentage of black voters tend to support Clinton and states with a large percentage of white voters tend to support Sanders, both to a statistically significant degree3:

demographics_all.png

To say that Sanders only does well with white voters or Clinton only does well with black voters isn’t quite true, as both sides are eager to point out: They perform roughly equally among the Hispanic/Latino and Asian/Other demographics, although Asian voters tend to lean slightly towards Sanders. But if we focus on only white and black voters, it’s clear Sanders has an advantage going forward. States in upcoming primaries have a higher percentage of white voters and a lower percentage of black voters than states who have already voted:

white_black.png

To take this analysis a step further, I used a linear regression model to predict voting outcome based solely on state demographics in states that haven’t voted yet. I then took each state’s pledged delegate counts to estimate each candidate’s delegates share. As it turns out, Sanders has more projected pledged delegates moving forward:

State Predicted Outcome Delegates Projected Clinton Delegates Projected Sanders Delegates
New Jersey Clinton +12.85 126.0 71.1 54.9
New York Clinton +11.48 247.0 137.68 109.32
Maryland Clinton +29.65 95.0 61.58 33.42
Pennsylvania Clinton +3.51 189.0 97.82 91.18
California Sanders +12.23 475.0 208.45 266.55
Delaware Sanders +9.56 21.0 9.5 11.5
Kentucky Clinton +18.49 55.0 32.58 22.42
Connecticut Sanders +19.63 55.0 22.1 32.9
Indiana Sanders +5.37 83.0 39.27 43.73
District of Columbia Clinton +45.79 20.0 14.58 5.42
Rhode Island Sanders +12.43 24.0 10.51 13.49
New Mexico Sanders +6.5 34.0 15.9 18.1
Montana Sanders +10.78 21.0 9.37 11.63
South Dakota Sanders +13.41 20.0 8.66 11.34
North Dakota Sanders +25.37 18.0 6.72 11.28
West Virginia Sanders +3.22 29.0 14.03 14.97
Oregon Sanders +47.37 61.0 16.05 44.95
TOTAL 776 797

Based purely on demographics, my model projects Sanders to win 11 upcoming states, compared to just 6 for Clinton, leading to a projected delegate margin of 797 to 776. Silver does have a point – of the four states whose demographics most closely resemble all democrats, my model projects Clinton as the favorite in all 4. Clinton surely has an advantage in both New York and New Jersey, both important primary states. However, his analysis swept aside one huge factor: California, which has a massive 475 pledged delegates, has voter demographics that are substantially more friendly to Sanders than Clinton. FiveThirtyEight projects Democratic voters in the Golden State to be only 11% African American while boasting the 2nd largest Asian American voting proportion, boding well for Sanders in the most crucial state for Democrats.

Now, a few things here. First, there are many, many more confounding factors that go into predicting a state’s voting behavior (caucus vs primary, closed vs open, etc). But for the sake of this post and for my rebuttal on Nate Silver, I’m focusing only on demographics here. Second, I’m not a political scientist by any means and didn’t research how each state determines how it distributes it’s delegates (if you know, leave a comment!). For simplicity’s sake I simply took the number of each state’s pledged delegates and multiplied by my model’s predicted voting percentage to calculate a rough projected delegate count for each candidate.

So keep in mind that these results are a relatively quick-and-dirty look at the relationship between state demographics and voting tendencies. And I think it’s safe to say that at the very least, demographics are not something Sanders has to worry about moving forward.

 

1 p-value obtained from a linear regression model with voting outcome regressed on RMSE. I used an alpha threshold of .05.

2 I used a Shapiro-Wilk test for normality and determined that a p-test is a valid test in this scenario.

3 p-value obtained from a linear regression model with voting outcome regressed on white/black voting percent. I used an alpha threshold of .05.