Showing posts with label polls. Show all posts
Showing posts with label polls. Show all posts

Monday, October 5, 2020

The DeSart & Holbrook 2020 Presidential Election Forecast

 For the past 5 presidential elections, Tom Holbrook and I have been generating forecasts of the national popular vote and Electoral College vote using a model that we developed in 1999.  Applying that same model to the data from 2020, we are projecting that Joe Biden will win a majority of the national two-party popular vote by a fairly wide margin of 54.4 to 45.6 over Donald Trump.   It also suggests that Biden will win 358 Electoral Votes to Trump's 180.

Our model uses state- and national-level polling in the month of September, along with each state's electoral history, to generate a prediction for each state's outcome.  In addition, we can use this data to calculate a probability for each candidate on how likely it is that they will win each state. We can then extrapolate those predictions up to the national-level to project the popular vote and Electoral College vote outcomes a month in advance of the election.  

Using this method, we have correctly predicted the winner of the national popular vote in each election from 2000 to 2016. Our track-record for our Electoral College projections is a little more mixed. We incorrectly projected that Al Gore and Hillary Clinton would win in 2000 and 2016, respectively. On the other hand, we went 51 for 51 in predicting the winner of all 50 states and the District of Columbia in 2012.

In 2016, our model projected that Hillary Clinton would win 52.05% of the national two-party popular vote, an error of just .95.  Our Electoral College projection estimated that Clinton would defeat Donald Trump 326 to 212.  We got five states wrong: Florida, Michigan, Ohio, Pennsylvania, and Wisconsin.  It's notable to point out that our model gave Hillary Clinton less than a 90% chance of winning of each of those states.  It's important to point that out because our model has never incorrectly predicted a state where it predicted a greater than 90% chance that it will be won by a candidate.

This is particularly relevant, because as you can see from the figure of Predicted Win Probabilities below, there are a number of states that the model predicts Joe Biden has a greater than 90% chance of winning.  More importantly, the number of Electoral Votes associated with these states totals 279, nine more than a candidate needs in order to win the election.  This suggests that even if Biden were to lose every other state, he would still win an Electoral College majority.


Of note is the location of the "tipping point" state.  This is the state where, when all states are arranged in order of their win probabilities, either candidate would achieve an Electoral College majority.  That state is Pennsylvania.  The problem for Trump is that it is located well within the Biden column. The model suggests that Biden has a 93.5% chance of winning Pennsylvania.  To win the election, Donald Trump is going to have to win seven states that the model suggests Joe Biden has a better than 50-50 chance of winning, two of which are over 90%.  That's not an impossible task, but it just doesn't seem likely.


Simulated Election Outcomes

As I mentioned above, we are able to take each of the model's predicted state-level outcomes and extrapolate national-level outcomes from them.  We can project the national popular vote by calculating a national popular vote total by taking each state's predicted outcome, weighting it by its contribution to the total national vote in the previous election, and summing it up.  That's how we derived the national popular vote projection of 54.4% for Biden.

We derive the Electoral College vote by simply awarding a state's Electoral Vote on the basis of the model's point estimates. One thing new this year, is that we are also incorporating projections for the Electoral Votes tied to the Maine and Nebraska Congressional Districts.  Doing this yields the Electoral College map pictured below.


You can see that the model suggests a split result in Nebraska.  While Trump is the clear favorite in the statewide vote, as well as in the First and Third Congressional District, the model indicates that Biden has a 74% chance of winning the Second Congressional District just as he and Barack Obama did in 2008.  On the other hand, the model suggests that Biden will win all four Electoral Votes in Maine, denying Trump the Electoral Vote associated with the the Second District that he won in 2016.

When we factor in the the uncertainty of the model, we can create confidence intervals around these projections. We run simulated elections while allowing each state to randomly vary around the standard error of the estimate, and aggregate the outcome in 100,000 simulated elections. Doing this yields the frequency distributions displayed below.



As you can see, Donald Trump does not win a majority of the national two party popular vote in any of the 100,000 simulations. 95% of Biden's outcomes fall within a range of 53.5 and 55.4%, so that serves as our Confidence Interval. When we expand the level of confidence out to 99%, the interval ranges from 53.1 to 55.8.

The distribution of Electoral College outcomes reveals the only ray of hope for Trump, but even so, the ray is very dim.  The most frequently occurring outcome is the one based on the point estimates of the model: Biden over Trump 358-180.  However, given the uneven relationship between the popular vote and the Electoral College Vote, the average outcome is 338-200. 95% of the outcomes fall within a range where Biden wins between 291 and 384 Electoral Votes.  99% fall between 279 and 401.



In other words, there is a greater than 99% chance based on this analysis, that Joe Biden will win and Electoral College majority and win the election. There is a small collection of outcomes where Donald Trump manages to win an Electoral College majority. In only 105 of the 100,000 simulations does Donald Trump win re-election, but it would entail a repeat of the 2016 election where he wins the Electoral College but fails to win the popular vote.


But what about 2016?

It's legitimate to question this prediction given that we, and a lot of other forecasters, missed the mark with our model in 2016. But there are some key substantive differences between 2016 and 2020 that leads us to have a bit more confidence in this forecast despite the 2016 misfire.

First, the lead that Joe Biden has in the national polls is substantively different than the lead that Hillary Clinton had in 2016.  Despite the widespread perception that "polls are broken" after what happened in 2016, national polls were not really as inaccurate as people think they were.  In September 2016, Hillary Clinton had an average two-party share of just 51.9% in national polls.  This is remarkably close to the actual result.  She ultimately won 51.1% of the two-party popular vote.

In contrast, Joe Biden's average share of the national polls in September was 53.8%, considerably higher than that of Clinton in 2016.  The polls have been remarkably stable over the past 12 months.  Biden has consistently led Trump since October of last year.  Over that time, Biden's monthly average two-party share in the polls has never dropped below 52.3% Given that, it seems highly unlikely that this lead will simply evaporate in the campaign's final weeks.

Of course, as we learned in 2016, it's what happens in the states that really matters when it comes to the deciding the Electoral College outcomes. It was the state-level polling that had the biggest issues in 2016, missing the mark in key states that ultimately tipped the balance in favor of Donald Trump.

Here, again, the situation is different than it was in 2016. The figure below shows the comparison of average September poll results for each of the 50 states.  In general, there has been an average shift towards Biden of a little over 2% across the states.  




Over the past five elections, without exception, when a candidate has a statistically significant lead in a poll in a state (ie, the lead is beyond the margin of error) for the month of September, they end up winning that state.

The table below dives a little further into this  comparison of the polling in 2016 with that of 2020. It shows how the September state polls compared to the eventual outcome.  Generally speaking, you can see evidence of what I mentioned above: September polls in 2016 actually did a reasonably good job of telling us what was going to happen in November, even at the state-level.  To be sure, there were polls in key states that ended up over-estimating Clinton's support, but her lead in those states was not statistically significant.  


Simply put, we could not be confident that she was actually leading in those states given the margin of error, so it should not have really been a surprise that she did not win those states. Most important is the fact that the disposition of many of these states this year is different than they were in 2016.  I have marked those states with an arrow showing how they've shifted.

All 12 states that have shifted since 2016 have moved away from Trump and towards Biden.  Trump does not have statistically significant leads in two states that were statistical locks for him in 2016: Alaska and Texas.  That doesn't mean he will lose those states, but it suggests that his position there, at least according to the polls, isn't as firm as it was four years ago.

Four states where Trump held statistically insignificant leads in 2016 have shifted towards Biden as well: Arizona, Georgia, Nevada, and Ohio. Biden holds slight leads in all four. Again, we can't say with any confidence that Biden will necessarily win those states simply based on these polls, but it is indicative of the general shift away from Trump compared to 2016.

Most relevant are the six states that have moved from being states that Clinton held insignificant leads in 2016 to states where, in 2020, Biden has a lead that is beyond the margin of error: Colorado, Maine, Michigan, Minnesota, New Hampshire, and Virginia. As I've stated above, in every election we've looked at going back to 2000, a candidate goes on to win a state where they hold a statistically significant lead in September. 

The states where Biden holds statistically significant leads account for a total of 240 Electoral Votes, meaning he only needs to find 30 more in order to win a majority.  A combination of just two or three of the eight states where he holds slight leads is all he needs to get him across the finish line. At this point it seems improbable, but not impossible, that Trump could win re-election. The map, and the context, looks considerably more difficult for him than it was in 2016. 

If 2016 taught us anything, however, it's that you shouldn't take anything for granted.  When the results come in next month, I'll be there looking at the numbers, because that's what a nerd does.

Saturday, October 29, 2016

Think This Month Has Been A Rollercoaster? Think Again.

If you've been keeping up with the news of the campaign this month, you probably feel like you've riding a roller coaster with all its twists and turns and ups and downs.

Each day there's been some new revelation about one or the other's campaign: Trump made another troubling remark; another Wikileaks dump of embarrassing Clinton campaign emails; another accuser comes forward alleging Trump groped/kissed/said a rude thing to her; another new poll showing the race is running neck-and-neck.

The undulations in this thrill ride we call a presidential election campaign have seemed to be relentless.  And now we have the latest stomach-flipping corkscrew in the track: The FBI has new evidence related to Hillary Clinton's email investigation that it's looking into.

There's a dramatic Breaking News alert every time one of these "bombshells" drop, and it's treated by the news media as if it is the game-changer that will potentially change the race, and fundamentally alter the course of U.S. history.   

Except when it doesn't. 

...which is most, if not all, of the time.

You see, voters don't really make sudden, roller coaster moves during a campaign, even when you would think otherwise, especially this late in a campaign.  The reason for this is pretty simple, and has some pretty deep roots in a psychological phenomenon known as cognitive dissonance.

You've most likely heard of cognitive dissonance before, even if not by name. If your parents, or a significant other, or even you, have ever said "you only hear what you want to hear," you're referencing the idea of cognitive dissonance.

About 60 years ago, a psychologist named Leon Festinger discussed what happens when humans encounter information that conflicts with something they already believe. It creates an uncomfortable state he referred to as cognitive dissonance. Since it is uncomfortable, people naturally try to minimize the importance of that information.  They'll ignore it, or try to rationalize it away.

On the other hand, when information reinforces what people already believe, there's no dissonance. Quite the contrary: they are all too willing to accept it because it makes them feel more justified in holding their views. It's not uncomfortable; it's like curling up with a cozy blanket.

Commentators on TV news programs like to use the term "baked in" a lot to talk about voters' attitudes about certain considerations, like their impressions of a candidate's qualifications, temperament, or honesty.  It's an overused term, but it does get to the heart of what I'm discussing here.

In Political Science there has been a long-standing proposition that campaigns actually have a minimal effect on the outcome of elections.  It's undergone some slight revision in the past couple of decades, but the simple point remains: By this point of the campaign the vast majority of voters have already made up their mind, and when events happen they usually hit the wall of people's pre-existing perceptions.

People will view campaign events through their own psychological filters.  If the events reinforce their predispositions, they will hold them up as highly relevant. If they create dissonance, they will be discounted as irrelevant.

So, what does that tell us about how the events of the past several weeks will affect the election on November 8th?  What it suggests is that "October surprises" often do very little to shift voter intentions.  As a case in point, take a look at the figure below. It presents the movement of the polls so far this month along with the "Breaking News" events we've seen so far.


You'll notice that the individual polls (the red dots) show a great deal of variability.  That's normal. Each individual poll is likely to have some error to it. That's to be expected, because each one represents a small subset of the population, and it is reasonable to expect it to have some error.  That's why we talk about a "margin of error." It's an acknowledgement that it's just a sample and could be wrong.  Even if a poll is conducted perfectly, it's likely to have some unintended bias in its sample.

That's why any savvy consumer of poll data will tell you that instead of focusing on any single poll, you should take a look at the average of polls.  Some polls will have error overstating one candidate's level of support, while others will have error overstating his/her opponents level of support. Sampling bias happens. It's not necessarily a nefarious attempt at generating a phony result.  It is just a simple by-product of not talking to everyone, because no one has time for that. But, generally speaking, that bias is going to be random.

That's why I've added the moving average line to the chart. It shows the ebb and flow of public opinion while smoothing out the random sampling errors of each of the individual polls.  The most important thing to notice is how little movement it shows.  Yes, it has moved and it has moved in somewhat predictable fashion.  It is possible to discern that there may have been an uptick of support for Clinton following the Access Hollywood tape coming out, but what's more telling is that it was pretty small. Only 2 points separate the peak (the week following the Access Hollywood tape) from the valley (the week right before the Access Hollywood tape).  For the most part, however, the polls have hovered within that range.

This is pretty consistent with research that has been done on the subject.  My forecasting partner, Tom Holbrook at the University of Wisconsin-Milwaukee, published a book a number of years ago entitled Do Campaigns Matter? and his conclusion was remarkably similar to what I've shown you here: That events in the course of a Fall campaign can move the polls somewhat in predictable fashion, but generally only do so a couple points, and often they move in opposite directions and cancel each other out.

It's this basic fact that makes it possible for an election forecast model like ours to generate fairly accurate predictions of the outcome a month in advance of the election.  By the end of September, most people already know how they're going to vote so they've already got their cognitive defenses built up.  So whatever comes in October often is blunted by their perceptual screens.

So, yes, the polls do move in response to "Breaking News" events, but not to degree that you would think, especially given the attention they are given on the cable news programs.  That's because, at this point in a campaign, the die is pretty much cast for the vast majority of voters.