Latest national poll median date: October 20
Projections reflect recent polling graciously made publicly available by pollsters and media organizations. I am not a pollster, and derive no income from this blog.

Wednesday, August 7, 2019

Projection Update: Mainstreet 7/30-31

Update Aug. 9: Due to small modifications to the model, parts of this post have been revised; the revisions are underlined.

The following poll has been added to the model:
Mainstreet's 7/30-31 national poll (current weight among national polls: 38%)
The high weight is due to this poll's relatively large sample size and the dearth of other recent polls. Because Mainstreet's earlier poll now has a weight of -12%, it does not disproportionately dominate the polling average. View the poll weighting methodology here.
For a full list of included polls, see previous Projection posts.

Seats projected ahead as of the latest national poll (midpoint: July 30.5)
LIB - 155 152 (32.8%)
CON - 150 151 (36.1%)
NDP - 17 18 (13.0%)
BQ - 11 (4.1%)
GRN - 4 5 (10.7%)
IND - 1 (Wilson-Raybould)

If you're new to my 2019 projections, view key interpretive information here.

Relative to the last projection (which has also been revised):
- Tories pick up one seat in NL from Liberals: Avalon.
- Tories pick up one seat in NS from Liberals: South Shore--St. Margaret's.
- Tories pick up two seats in PE from Liberals: Egmont, Charlottetown.
- The NDP picks up one seat in QC from Liberals: Rimouski-Neigette--Témiscouata--Les Basques.
- Liberals pick up seven seats in ON from Tories: Mississauga--Streetsville, Oakville, Mississauga--Lakeshore, Markham--Stouffville, Kitchener South--Hespeler, St. Catharines, Vaughan--Woodbridge.
- Liberals pick up one two seats in ON from the NDP: Nickel Belt, Ottawa Centre.
- Tories pick up one seat in MB from Liberals: Charleswood--St. James--Assiniboia--Headingley.
- Tories pick up both the last NDP seats in SK: Desnethé--Missinippi--Churchill River, Saskatoon West.

Monday, August 5, 2019

First Projection for the 2019 Election

Update Aug. 9: Due to small modifications to the model, parts of this post have been revised; the revisions are underlined.

And we're off! Here is the first Canadian Election Watch projection for the 2019 election.

As of the latest national poll (midpoint: July 27)
CON - 153 152 (36.4%)
LIB - 149 (32.4%)
NDP - 20 (13.0%)
BQ - 11 (4.1%)
GRN - 4 5 (10.9%)
IND - 1 (Wilson-Raybould)
With the polls being this close, it comes as no surprise that it's almost 50/50 regarding who gets the most seats. But wait a minute... I actually have the Tories up 4%, which seems at odds with the polling averages on other websites. I explain the discrepancy toward the end of this post: multiple factors may be pushing in the same direction.

Update Aug. 11: I have revised up the pollster common variance used to derive confidence intervals. The ones below are a bit too narrow - add roughly 5-10 on each side.

CRUDE 80% and 95% confidence intervals for election as of midpoint of latest poll
CON: 130-175, 120-185
LIB: 125-170, 115-180
NDP: 12-28, 8-35
BQ: 2-20, 1-25

CRUDE 80% and 95% confidence intervals for actual election (at this stage, this is really just for fun)
CON: 70-265, 40-310
LIB: 55-235, 30-275
NDP: 2-60, 0-75
BQ: 0-50, 0-60
These numbers may seem crazy. But remember: 95% confidence interval means you'd do better only once every 40 times, and worse once every 40 times. (EDIT: To be precise, it actually means that if the truth is at the upper/lower end of the interval, then the point estimate would be lower/higher once every 40 times.) And this is the 43rd General Election in our history. Just in the past eight general elections, we've had two parties losing over 90% of their seats (PCs from 169 to 2 in 1993, BQ from 49 to 4 in 2011), and a party gain 150 seats (Liberals from 34 to 184 in 2015).

This post contains some important information for interpreting these numbers and confidence intervals, as well as links to a fairly detailed description of the methodology. As the post indicates, the seat numbers are the number of ridings where each party is projected ahead, which is different from each party's expected number of seats.

(In case you're wondering about the background colour: I change it periodically according to which party leads the projection, and how close it is to a majority. Right now, it's only very slightly blue because the projection is virtually tied.)

This projection is intended to use all national polls satisfying the inclusion criteria with field dates ending on July 1 or later:

Mainstreet's 6/27-7/2 poll
Nanos' poll ending 7/12 (presumed 6/17-7/12 Update Sept. 9: Actually 6/15-7/12 as Nanos does poll on weekends)
Abacus' 6/28-7/2 poll
EKOS' 7/7-9 6-8 poll
Angus Reid's 7/5-12 poll
Campaign Research's 7/9-12 poll
Ipsos' 7/12-15 poll (numbers from Global News article only)
Abacus' 7/12-17 poll
Research Co.'s 7/15-17 poll
EKOS' 7/17-19 16-18 poll (national, Ontario and Québec numbers only*)
Léger's 7/19-23 poll
EKOS' 7/24-26 23-25 poll (national and Québec numbers only*)
Forum's 7/26-28 poll
*Updates Sept. 9 and 10: The full regional breakdowns for these polls were included when made available in early September.

Regional and riding-level polls also factor into the projection; they are listed in this post. If there is a poll that I have missed (and satisfies the inclusion criteria), please let me know!

Here is a map of this projection. Most projections won't be mapped, but since this is the first one, I figure it could be fun to look back in October and marvel at how things have (or haven't) changed. (The model modification caused offsetting NDP-Conservative changes in St. John's East and Desnethé--Missinippi--Churchill River, as well as the Greens adding Cowichan--Malahat--Langford. The map has been updated.)


If the situation on election night is anywhere near this projection, it'll be a long and exciting (or nerve-wracking) night! Let's play along for a bit:
- The Conservatives re-establish themselves in the Atlantic, aided by vote splitting between the Liberals and the Greens; Jack Harris' comeback bid falls just short succeeds after a tight three-way race. 10 11 of the 32 Atlantic races are decided by 5 points or less.
- The Liberals pick up 9 seats in QC without increasing their vote share at all, thanks to the NDP's collapse. The Tories pick up 5, while the Bloc falls just short of official party status again. Alexandre Boulerice is the only NDP MP left in La Belle Province.
- ON proves very competitive, with a slight edge to the Liberals. The NDP gains a few seats thanks to Liberal weakness relative to 2015 - its only gains of the night.
- The Tories consolidate their position in MB/SK, and re-take all 4 Liberal seats in AB.
- BC decides who gets the plurality, and thanks to vote splitting on the left (this is the strongest region for both the Greens and the NDP), the Tories win a majority of BC seats on just 35% of the vote, similar to the Liberals in QC. The Greens capture two three more seats, establishing dominance over southern Vancouver Island. Jody Wilson-Raybould wins as an independent, but Jagmeet Singh bites the dust as Burnaby South goes Conservative, thanks to the People's Party petering out; he resigns as NDP leader.
- Government formation is complicated. The Liberals and NDP have exactly half of the seats. So the Tories must get support (or abstention) from either of those parties - Bloc, Greens and independent aren't enough. Similarly, NDP support isn't enough for the Liberals.

Now, I don't expect most of this to happen - see the "actual election" confidence intervals above!

This projection is more generous to the Conservatives than many other websites'. Keep in mind that I make a turnout adjustment of +1% in favour of the Tories, which explains part of the difference. Still, even without this adjustment, I have the Tories up by 3% over the Liberals, more than some other sites. This could be due to my more restrictive poll inclusion criteria (specifically, not using data kept behind a paywall) and/or to my more aggressively discounting older polls (the three most recent polls all show Conservative leads of 3% or more, and have over 60% of the weight).

Also, the Research Co. poll has a regional breakdown that is inconsistent with the top line national numbers: a weighted average of the regional numbers gives a 3-point Tory lead rather than a 3-point Liberal lead. This can happen as response weights may differ when calculating different sets of numbers, but it's rather unusual for the gap to be this big. I went with the regional breakdown, which may also contribute to the polling average being tilted to the Tories.

I do not believe that my actual seat model is more friendly to the Tories than, say, 338Canada's. In fact, with 338's current regional polling averages, I get the Liberals leading in 16 (this figure may have changed; I did not check after the model modification) more seats than the Conservatives rather than 6 more seats.

Sunday, August 4, 2019

Important Information for Interpreting 2019 Projections

If you're new to this blog, welcome! This post provides information that's useful to keep in mind for interpreting the forthcoming projections.
Update Aug. 11: Change in the common pollster variance used to derive confidence intervals. New value underlined.
Update Aug. 13: Major change in how I am presenting projections. Changes underlined. I have also added the first paragraph containing information about inputs into the model.
Update Aug. 21: Change to turnout adjustment. Changes underlined.
Update Aug. 26: Separate polling weights are now derived for each region. Link to relevant post added.
Update Sept. 13: Belatedly added a missing link to Round 2 of polling-related adjustments.
Update Oct. 11: Belatedly added missing links to Rounds 3 and 4 of polling-related adjustments.

My model is relatively parsimonious. It uses the following information:
- Results of the last General Election and by-elections since then
- Results of earlier elections, but ONLY for "special cases" (see part II of this post) and to calibrate model parameters (e.g. turnout adjustment (as discussed below), uncertainty)
- Polling data
In particular, the model does NOT use the following information:
- Election results before the last General Election, except for special cases and calibration
- Incumbency effect, except when a star MP leaves, which is treated as a "special case" (incumbency effect tends to be insignificant for low-profile politicians, unlike in the U.S.)
- Demographic data
- Data about online activity (e.g. search data, social media followers, etc.)
Despite using a limited set of data, I have a good track record relative to others making seat projections: see the 2011 and 2015 comparisons. I have also enhanced the model in several ways for the 2019 election (see the links below), so I hope to do well again!

The projections are made on vote shares after the following turnout adjustments:
CON: +1.5 pp
NDP: -1 pp
LIB: +0.5 pp
GRN: -1 pp
This may explain why my projection is more favourable to the Conservatives than other projections. The rule of thumb in this election appears to be, very roughly, 1 point nationwide = 10 seats (CLARIFICATION: this is for the gap between the two main parties).

The turnout adjustments are based on how election results compared with the final polls in recent federal general elections. The Conservatives have consistently outperformed the polls: by about 1 point in 2015 (they did worse seat-wise because the Liberals also outperformed, and did so very efficiently), and by several points in 2008 and 2011. The Liberals roughly matched the final polls in 2008 and 2011, and outperformed them by 1-2 points in 2015. The Greens have consistently underperformed by 1 point or more. The NDP has usually somewhat underperformed as well, but since their support is currently much lower than usual, the appropriate adjustment may also be lower (and thus almost zero). although the pattern is not as clear as for the Greens. Therefore, I took at look at recent provincial elections, and the only cases where the NDP did not underperform are when pollsters were in high agreement (BC, SK) or, in the case of MB, when the Greens weren't running a full slate of candidates (presumably causing some Green voters to vote NDP once they realized their preferred party isn't on the ballot). As neither of these is currently the case, I have decided to add a small adjustment for the NDP. As you can see, I am erring on the cautious side when making these adjustments, since the past doesn't always repeat itself. I reserve the right to tinker with these as the campaign progresses based on a subjective appreciation of various indicators of voter enthusiasm, and will let you know if I do.

In the interest of transparency (and geekiness), here are the posts describing the methodology for these seat projections. Many of these features are new for 2019!
Poll inclusion criteria
How national polling weights are derived (see this post for an illustration)
How regional polling weights are derived
General modification to uniform swing and riding-level adjustments to the 2015 baseline
Polling-related adjustments: Round 1, Round 2, Round 3, Round 4

The seat numbers on the left are the number of ridings where each party is projected ahead, which is different from while those on the right are rough estimates of each party's expected number of seats. For example, if party A has a 60% chance of winning two ridings, while party B has a 99% chance of winning another riding, the projection table on the left will show 2 seats for A and 1 seat for B, even though the expected number of seats while the numbers on the right would be 1.21 for A and 1.79 for B. Nationally, the gap between these numbers should usually (but NOT always) be small relative to the totals, as each party typically has its share of narrow wins. However, the provincial breakdown in the left column of the blog should be treated with caution.

The shortcut I take to estimate the expected number of seats will lead to a slight underestimation of parties that are disproportionately often competitive despite being third or fourth. My simple method is explained here:
Estimates for expected number of seats

Confidence intervals, provided from time to time, are crudely estimated rather than based on simulations. This is merely to save myself some work - ideally, they should be based on simulations. I differentiate between intervals for a hypothetical election happening at the time of the latest poll (specifically, the midpoint) and those for the actual election. Ranges for the latter are wider because public opinion may shift. All intervals are speculative; for large parties, I round them to the nearest 5 to draw attention to their very rough nature.

The assumptions behind the confidence intervals are the same as those behind poll weighting. In addition, I assume a variance of 0.0002 0.0006 (i.e. a standard deviation of about 1.4% 2.45%) on the common component of pollster error. I won't go into the specifics of the (definitely not rigorous) method in which I estimate the confidence intervals, but on the eve of the election, the "actual election" confidence intervals should be consistent with the magnitudes given in this 2015 post (somewhat wider for the Liberals and Conservatives, and narrower for the NDP, due to the changed political landscape); they'll likely be pretty similar to the ones given by 338Canada, whose calibration seems solid.

Some other projection websites also provide confidence intervals, often using simulations, which is the way it should ideally be done. However, simulations can still be junk if they're miscalibrated. Other sites aren't always clear about the timing of the hypothetical election for which confidence intervals are given. If they're based on how accurate the last polls before an election have historically been, the confidence intervals would be for a hypothetical election taking place just after the latest poll. They should therefore be wider than the confidence intervals I give for a hypothetical election taking place during the latest poll, and narrower than the ones I give for the actual election (except just before the election, when they should be similar to the latter). If another projection site gives confidence intervals for an election today/tomorrow narrower than those that I give for an election as of the last poll, then the other site is probably not taking the uncertainty seriously enough.

Polling-Related Adjustments 2019, Round 1

Update Aug. 9: Beauce poll added (updated adjustments, including knock-on effects and minor corrections on Quebec regional adjustments, underlined; minor additional adjustments may be made as contemporaneous provincial numbers become available)

First, a quick word about my polling averages. In 2015, 0.6% of the national vote went to minor (below 5% of the riding vote) independent and "other" candidates. This year, I will assume the same. The People's Party is not treated as a regular party since there is no baseline for it. I will give it a basic vote share of 2.4%, to be adjusted from time to time on an ad hoc basis. In total, this leaves 97% of the vote to the major parties. This is reduced to 96.8% in ON (Jane Philpott), 96.6% in QC (Maxime Bernier adding to the PPC base), and 96.2% in BC (Jody Wilson-Raybould). I then adjust the raw polling averages from polls (weighting described here) proportionally so that the totals match these figures. A turnout adjustment is then applied - it will be described in a future post.

The rest of this post reviews polls of smaller areas than the six standard polling regions (BC, AB, SK/MB, ON, QC, Atlantic). Resulting adjustments to the model are indicated in bold. You can view the current non-poll-based adjustments to the model here.

I. Regional Polls (from Jan. 1, 2019)

Atlantic Canada
MQO Winter 2019 polls (1/16-28 NL, 1/21-27 PEI, 1/30-2/10 NS, 1/30-2/10 NB)
MQO Spring 2019 polls (4/11-16 PEI, 4/25-5/4 NL, 4/23-5/6 NB, 5/3-13 NS)
Narrative's 5/6-24 poll
These polls consistently suggest that the Conservatives are up more than expected in NL (the Harper effect going away?), and that the Greens are not catching on there. Adjustments:
NL: CON +9, NDP -2, LIB -1, GRN -6
NS: CON -3, NDP +3.5, LIB -1, GRN +0.5
NB: CON -2.5, NDP -2, LIB +2, GRN +2.5
PE: CON +1, LIB -3.5, GRN +2.5

Québec
Léger's 1/25-28 poll
Léger's 3/8-11 poll
Forum's 6/11-12 poll
Forum's 7/22-24 poll
All these polls suggest that, even though the Conservatives are gaining in Québec as a whole, they are gaining less (or not at all) in the Quebec City area. One explanation could be that the Quebec City area already moved to the Tories prior to the 2015 election. Instead, the Tories appear to be gaining a little more in rural/small-town Quebec, where the Liberals and NDP do a little worse than expected. Adjustments:
Quebec City CMA: CON -3, NDP +1.5, LIB +1.5
Montreal CMA: CON -1, LIB +1
Rest of Quebec North: CON +2, NDP -1 -0.5, LIB -1.5
Rest of Quebec South: CON +2.5, NDP -0.5, LIB -2
(Salaberry--Suroît is mostly outside of Montreal CMA but shares some characteristics with it. Since the Montreal CMA and Rest of QC S adjustments go in opposite directions, I have left this riding unadjusted.)
This are only partial adjustments, which will be increased in size if future polls point in the same direction. Even so, they are flipping many seats right now: Liberals get to keep Québec and Louis-Hébert, while Saint-Hyacinthe--Bagot, Jonquière and Trois-Rivières go to the Tories instead of the Liberals.

Ontario
Forum's 7/3-6 poll (Toronto only)
Corbett Communications' 7/9-10 poll
Both these polls suggest that the NDP is doing worse in Toronto - specifically in Old Toronto, if you believe the Forum poll - than uniform swing implies. For now, I'll be cautious and go with a relative small adjustment: NDP -5, CON +2, LIB +3 in the seven central ridings in Toronto. This flips three seats (Davenport, Parkdale--High Park and Toronto--Danforth) that the model would have had the NDP just barely taking back from the Liberals.

Manitoba/Saskatchewan
Probe's 3/12-24 poll (Manitoba only)
Probe's 6/4-17 poll (Manitoba only)
These polls have the Liberals down more in MB than national polls have them down in MB/SK, and have the Conservatives and NDP up in MB while they're flat or down in MB/SK in national polls. They also show the NDP up in Winnipeg (but down in the part of Winnipeg roughly corresponding to Elmwood--Transcona) and down in the rest of MB. However, these polls also asked about provincial politics, and there is an upcoming provincial election from MB. So I will be cautious for now and not adjust the model; if these patterns continue after the MB election in mid-September, I will revisit the issue.

British Columbia
Justason's 4/4-5 poll (City of Vancouver only)
This polls shows a huge drop off for the Liberals in the City of Vancouver, and a huge rise for "Other" (much more than Jody Wilson-Raybould can account for). Given the tiny sample size, I won't make any adjustment here yet.


II. Riding Polls (by Mainstreet unless otherwise noted)

Note: A riding poll will result in a model adjustment ONLY IF it shows a substantial deviation from regional swing and/or it is in a riding with unusual factors. Most riding polls will therefore be ignored: they generally did not help much in 2015.

Beauce, 11/10-11
Beauce, 8/5
CON -30 -25, LIB -5, PPC +30 (Maxime Bernier)
This is a two-way Conservative/People's Party race.

Oakville, 5/27
This poll has a different party ahead depending on whether party leader or actual candidate names are given (the Liberals did not nominate the most popular candidate). The projection has Oakville as a very tight race, and is not adjusted.

Vancouver Granville, Justason 4/4-5 and Mainstreet 5/29-30
CON -10, NDP -8, LIB -10, GRN -2, Wilson-Raybould +30 
It'll be interesting to see whether Wilson-Raybould successfully maintains her lead during the campaign: it's always a struggle for independents not to be squeezed out by party politics. This could turn out to be a three-way Wilson-Raybould/Liberal/Conservative race.

Markham--Stouffville, 5/29-30
LIB -5, NDP -2, CON -5, GRN -3, Philpott +15
This poll somewhat surprisingly suggests Philpott pulls more from the Conservatives than the Liberals. I am not prepared to make such an adjustment without corroborating evidence, as this riding is a very tight two-way race. So for now, the adjustment assumes that Philpott pulls equally from the two main parties, and the two-way race is unaffected.

Niagara Centre, 7/15-16
This poll is bad news for the NDP and good news for the Liberals. It departs significantly from model expectations. I will adjust the model about halfway toward the poll: CON +1, NDP -6, LIB +6, GRN -1. This shifts the projection from Tories to Liberals.

Whitby, 7/18-19
This poll is very close to the projection. No adjustment.

Québec, 7/23-24
This poll is close to the projection. The Liberal lead is a bit bigger than projected, but that's because I'm being cautious for now on the regional adjustment. No further adjustment at this point.

*Leaked internal polls will not be included in the projection. That said, you may be interested in these.

The Not-Quite Uniform Swing Model for 2019

Update Aug. 9: I've made some changes (indicated in italics) to Section I. The model is slightly simplified and hopefully more robust.

Update Aug. 10: Added Edmonton Strathcona to Section II.e. This change does not affect any past projections.

Update Aug. 13: Added Kenora to Section II.d. This change flips Kenora to the Liberals as of Aug. 13. Also added a clarification at the end of Section I.

Update Aug. 15: Changed the adjustment for Victoria in Section II.a, and added Thunder Bay--Superior North to Section II.d. (see Sept. 8 update)

Update Aug. 18: Reviewed the adjustment for Skeena--Bulkley Valley in Section II.e.

Update Aug. 20: Added Longueuil--Saint-Hubert to Section II.c to reflect news.

Update Aug. 21: Reviewed the adjustment for Burnaby South (Section II.b) to offset new NDP turnout adjustment, and refined where the extra NDP vote comes from. Also added adjustments for Halifax, Sackville--Preston--Chezzetcook and Ottawa Centre (Section II.d) following the same approach as used for Skeena--Bulkley Valley in an earlier update.

Update Aug. 22: Added Windsor West to Section II.f. Adjustments further updated as of Sept. 8.

Update Sept. 8: Removed Thunder Bay--Superior North from Section II.d due to Bruce Hyer candidacy.

Update Sept. 16: Reviewed Lac-Saint-Jean adjustments.

Update Sept. 27: Added Section II.f. adjustment to Laurier--Saint-Marie.

Update Oct. 5: Expanded Section II.a. to include candidates ejected by their party or withdrawing after the 2019 nomination deadline, and added an adjustment for Burnaby North--Seymour.

As those of you familiar with this blog know, my seat projections are primarily based on assuming uniform swing within each of Canada's six "polling regions" (BC, AB, MB/SK, ON, QC, Atlantic) from the previous election to the current polling average. This post describes two categories of non-polling-based* changes that I am making to uniform swing.
*Other than comparing by-election results to contemporaneous polls for adjustments in II(b)

Part I is geeky (light math involved), while Part II looks at many specific races. They're independent from each other, so if you don't want to eat your veggies before dessert, skip to Part II!

[EDIT Aug. 5: Added Ottawa Centre to the list, though no adjustment there for now.]

I. Smaller Swings in Uncompetitive Ridings

As my post analyzing model performance for the 2015 election noted, uniform swing tends to shift party support too much when it is either very low or very high. This is not a new discovery, but due to the size of the swings in 2015, it actually mattered. Going to proportional swing doesn't solve the problem: while it would reduce shifts when party support is very low, it would increase them when party support is very high.

The starting point of my modification is the assumption that shifts between two parties are roughly proportional to the sizes of both parties - a sort of gravity model of voting. This isn't crazy: a shift from party A to party B has greater potential if A has more supporters and B has more salience. If there were only two parties, this means that shifts are biggest when the parties are at 50-50, still similarly big when they're at 70-30, but substantially smaller when the support levels are 80-20 or more extreme.

Of course, Canada has more than two major parties, so things are more complicated. First, for every pair of "significant" parties (big 5 in Québec, big 4 elsewhere)*, I multiply the two parties' pairwise vote shares. For example, if a riding is LIB 40%, CON 30%, NDP 20%, GRN 10%, then the three pairwise vote shares involving the Liberals for that riding are:
LIB-CON: (4/7)*(3/7) = 12/49
LIB-NDP: (4/6)*(2/6) = 2/9
LIB-GRN: (4/5)*(1/5) = 4/25

Then, I calculate a "coefficient of variation" for each party within each riding, equal to the average of the pairwise vote shares involving that party, weighted by the vote share of the other party in each pair. So in the example above, the Liberal coefficient of variation would be:
(30*(12/49) + 20*(2/9) + 10*(4/25))/60
This is approximately equal to 0.223. Note that the maximum possible coefficient is 0.25, which happens when a party is tied for first place, and any party not tied for the lead has no support. This corresponds to the intuition that competitive ridings may be especially fluid because many voters see the parties in contention as roughly equally good (or bad).

The swing I apply to a party's vote share in riding X is then its regional swing multiplied by its coefficient of variation in riding X, divided by its average coefficient of variation in the region. So if the Liberals have a coefficient of 0.22 in Abitibi--Baie James--Nunavik--Eeyou (it's my favourite riding name), an average coefficient of 0.2 in Québec, and increase by 5% in Québec, then I increase their vote share by 5%*0.22/0.2 = 5.5% in Abitibi--Baie James--Nunavik--Eeyou.

Then, for each riding, I proportionally increase or decrease the support levels so that they sum to the same number as they would if uniform swing were applied (i.e. 95-100% for most ridings).

Finally, I compute the average swing for each party within each region. This should be close to the target uniform swing. Any remaining gap is closed via uniform swing.

Now, a question arises: should the coefficients of variation be calculated based on support levels in the previous election, or hypothetical support levels after applying uniform swing (with a minimum threshold at, say, 2%)? Testing this on the shifts from 2011 to 2015 using actual regional support levels, I find that the variance of errors is minimized when calculating the coefficient both ways, and using 10% of the former and 90% of the latter. However, testing on a hypothetical reverse shift (if we were going back in time from 2015 to 2011) suggests that putting more weight on the former is better. Therefore, I have decided to proceed in a symmetric way, and simply take the arithmetic mean of the two. I also find that the results are best when they're pulled back 15% of the way toward what uniform swing would have given. (This crossed-out sentence no longer applies after the formula change.)

Next question: does this actually work? Again, testing this on the shifts from 2011 to 2015 using actual regional support levels, I find that, outside Québec, the variance of vote share errors is reduced by 10-20% around 10% for most parties relative to using uniform swing (with a minimum threshold). (Within Québec, it doesn't seem to make much of a difference.) The regional projected seat counts become slightly closer to the actual result on average. Québec is the main exception to the latter as this method would have overshot on the Liberal seat count by even more than uniform swing (57 instead of 47 vs. 40 actual), but half of the additional Liberal overshoot actually improves the riding-level accuracy (either the Liberal actually won the riding, or the Liberal was 2nd while uniform swing projected the 3rd place candidate as the winner). This problem no longer arises in Québec.

Now, since this change was inspired by observations made in 2015, one should worry that the method wouldn't work on another election. While I didn't do extensive testing, I did look at what would happen if one were to project the 2011 results on the 2015 to 2011 shifts (i.e. pretend that we're going back in time). Again, the variance of vote share errors is reduced, and the seat projection is improved both regionally and nationally. (though not by as much). Although the seat projection ends up slightly farther from the actuals, the number of ridings projected correctly slightly improves.

I will adopt this model for my 2019 projections, and assess its performance relative to uniform swing after the election.

*Clarification: For 2019, I am assigning a flat vote share of 2.4% for the People's Party of Canada (outside of Beauce) and of 0.6% for minor candidates (all candidates outside the six largest parties, except for star independents - currently Wilson-Raybould and Philpott). Thus the five large parties' support will always add up to 96.7-96.8% (97% minus adjustments for Bernier and star independents). This is to correct for the tendency of polls to overstate the support of minor parties and "others." Of course, if the People's Party takes off, I will review this assumption.

II. Adjustments to 2015 Riding Baselines

These adjustments do NOT include poll-based adjustments to swing, which will be discussed in a future post. At issue here are adjustments based on information about candidates and past election results. These are somewhat subjective, so let me know if you see something that you think is way off!

a) Party not running or withdrawing in a riding in 2015 or 2019

Labrador: Greens did not run in 2015. Baseline adjustments: CON -0.2, NDP -0.3, LIB -1, GRN +1.5. Obviously, this adjustment makes little difference.

Mississauga--Malton: Conservative candidate dropped by party in 2015 (after deadline, so still got many votes). Baseline adjustments: CON +2, NDP -0.3, LIB -1.4, GRN -0.3. Again, this adjustment makes little difference.

Kelowna--Lake Country: Greens did not run in 2015. Baseline adjustments: CON -1.3, NDP -2.3, LIB -5.7, GRN +9.3. This will make it even harder for the Liberals to keep the riding, which is already an uphill climb.

Victoria: Liberal candidate withdrew in 2015 (after deadline, so still got some votes). Baseline adjustments: CON -1.3, NDP -3.7, LIB +16.3, GRN -11.3. This is still a Green/NDP race with the Greens favoured, but the race is tighter. NDP -6.5, LIB +22.8, GRN -16.3. (This adjustment used to be a blend of different approaches, and was modified after new polling suggests that one of the approaches is more appropriate.) This is a tight three-way race.

Burnaby North--Seymour: The Conservative candidate has been ejected from her party, but will remain on the ballot, much like the Liberal candidate in Victoria in 2015. The Victoria adjustments imply that the Liberal candidate that dropped out still got about 1/3 of the vote that she would otherwise have received. For this year's situation, since Conservative voters are less likely to have a second choice, I assume that this candidate will retain 1/2 of the vote that she would have otherwise received - around 15% per the current projection. The remaining 15% is distributed among the other parties through a combination of two effects: (i) Tories actually voting for their second choice, and (ii) Tories staying home. Factor (i), according to second choice data in recent surveys, might slightly favour the PPC and the NDP over the Liberals and Greens - but it comes pretty close to an even four-way split. Factor (ii) will increase the vote share of parties proportionally, thereby favouring the Liberals and NDP. To obtain a result consistent with these assumptions, the baseline adjustment is: CON -14, NDP +6, LIB +5.5, GRN +2.5, and will be coupled with a 3-point increase for the PPC projection taken proportionally from other parties. These changes will be applied on projections where the midpoint date of the most recent poll is October 4 or later.

b) By-elections

CON holds (no adjustment)
Medicine Hat--Cardston--Warner
Calgary Heritage
Calgary Midnapore
Sturgeon River--Parkland
Battlefords--Lloydminster
Leeds--Grenville--Thousand Islands and Rideau Lakes
York--Simcoe

LIB holds (no adjustment)
Markham--Thornhill
Ottawa--Vanier
Saint-Laurent
Bonavista--Burin--Trinity
Scarborough--Agincourt

NDP hold
Burnaby South (Jagmeet Singh's riding): -1.5 CON, +5 +6 NDP, -5 -3.5 LIB, -1 GRN. This puts the Liberals further behind, making this riding an NDP/Conservative race (the People's Party candidate, who got over 10% in the by-election, is running in Alberta in the General Election).

CON to LIB
Lac-Saint-Jean: -10 CON, -5 -8 NDP, +10 +13 LIB, +5 BQ. This is shaping up to be a Liberal/Conservative contest.
South Surrey--White Rock: No adjustment

LIB to CON
Chicoutimi--Le Fjord: +25 CON, -10 NDP, -5 LIB, -10 BQ. This riding now looks solid for the Tories.

NDP to LIB
Outremont: No adjustment

NDP to GRN
Nanaimo--Ladysmith: -5 LIB, +5 GRN. This helps the Greens hold the NDP at bay.

c) MP changed affiliation and seeking re-election
- Beauce: rolled into poll-based adjustment
- Longueuil--Saint-Hubert: -3 NDP, +3 GRN. This is a relatively cautious adjustment, and will be increased if polls warrant. The riding remains a two-way LIB/BQ race.
- Aurora--Oak Ridges--Richmond Hill: No adjustment
- Markham--Stouffville: rolled into poll-based adjustment
- Vancouver Granville: rolled into poll-based adjustment
Raj Grewal (Brampton East) and Darshan Kang (Calgary Skyview), two Liberals turned Independent, have not yet announced (as far as I know) whether they will run again. If they do, there may be adjustments.

d) Star candidates or strong independent candidates that lost in 2015 and may not run in 2019
Note: If a candidate ends up running again, the corresponding adjustment will be cancelled.

Avalon (Scott Andrews): CON +6, NDP +6, LIB +5.5. This won't make much of a difference.

Egmont (Gail Shea): CON -10, LIB +10. Since the 1980s, this seat had consistently been around the PEI average or more Liberal until Gail Shea ran in 2008. The adjustment pushes the seat most (but not all) of the way toward the PEI average, and implies a competitive race.

Halifax (Megan Leslie) and Sackville--Preston--Chezzetcook (Peter Stoffer): No adjustment for now since there is no reasonable base (Leslie was preceded by Alexa McDonaugh, while the Sackville riding has been represented by Stoffer since its creation), but I will be less hesitant to adjust these ridings away from the NDP based on riding polls. Halifax: NDP -1, LIB +0.5, GRN +0.5. No new information, but I took another look at the situation by examining how things have evolved since Leslie inherited the riding from McDonaugh in 2008. There's plenty of uncertainly as to the size of the Leslie effect (it could be a lot larger than the adjustment implies), but I went with the most cautious method. Sackville--Preston--Chezzetcook: NDP -2, LIB +2. No new information either, but similarly, I took a look at how things have evolved since the riding was created in 1997. This adjustment is very cautious: while it looks like Peter Stoffer had a large positive effect on the NDP vote, most of it came between 1997 and 2004. So it's possible that most of those voters became habitual NDP supporters, or that the effect gradually dissipated, and Stoffer's vote held up due to other factors. In both cases, we have tight races (three-way in Sackville--Preston--Chezzetcook, LIB/NDP in Halifax), though I wouldn't be surprised if constituency polls end up dramatically increasing the adjustment and favouring the Liberals.

Avignon--La Mitis--Matane--Matapédia (Jean-François Fortin): NDP +3.7, LIB +3.7, GRN +0.5, BQ +3.7. This won't make much of a difference.

Laurier--Sainte-Marie (Gilles Duceppe): NDP +2, LIB +1, BQ -3. This will reduce the odds of the Bloc retaking this riding.

Pierre-Boucher--Les Patriotes--Verchères (JiCi Lauzon): NDP +1, LIB +1, GRN -6, BQ +4. This will increase the odds of the Bloc keeping this riding.

Ottawa Centre (Paul Dewar): Somewhat surprisingly, historical patterns do not suggest Dewar did better electorally than an "average" New Democrat would have done in this riding. Therefore, there is no adjustment for now, but I will be less hesitant to adjust these ridings away from the NDP based on riding polls. NDP -2.5, LIB +2, GRN +0.5. No new information, but I took another look at the situation by examining how things have evolved since Dewar inherited the riding from Broadbent in 2006, and, if you squint, Dewar 2015 seems slightly more popular than Dewar 2006.

Eglinton--Lawrence (Joe Oliver): CON -1, LIB +1. This will increase the odds of the Liberals keeping this riding.

Renfrew--Nipissing--Pembroke (Hec Clouthier): CON +5, NDP +1, LIB +5. High uncertainty here (vote pattern suggests Clouthier's vote came from Tories, but he's an ex-Liberal), but it doesn't matter much because this is a very conservative riding (one of two ON Canadian Alliance seats in 2000).

Thunder Bay--Superior North (Bruce Hyer): CON +3, NDP +5, GRN -8. Hyer was elected as an NDP MP in 2008 and 2011, but ran under the Green banner in 2015. This adjustment makes the NDP stay the main contender to the (still favoured) Liberals. (Update Sept. 8: Hyer is running again per this article.)

Kenora (Howard Hampton): CON +2, NDP -4, LIB +2. This will increase the odds of the Liberals keeping this riding.

Dauphin--Swan River--Neepawa (Inky Mark): CON +4.5, NDP +0.5, LIB +2.5, GRN +0.5. Still safe Conservative.

St. Albert--Edmonton (Brent Rathberger): CON +5, NDP +3, LIB +10, GRN +1.5. Still safe Conservative.

e) Star MPs retiring in 2019

Cumberland--Colchester (Bill Casey): CON +10, LIB -10. This traditionally Conservative seat becomes a very likely Tory pickup (likeliest in Atlantic Canada outside NB).

Kings--Hants (Scott Brison): CON +5, LIB -5. This is maybe half the voters Scott Brison "took with him" from the PCs to the Liberals back in the day - the assumption being that the other half may stay in the Liberal camp now that it has been so long. The adjustment makes this riding more like West Nova and potentially competitive.

Edmonton Strathcona (Linda Duncan): NDP -3, LIB +3. Linda Duncan has been around for a long time, so this adjustment is pretty cautious.

Skeena--Bulkley Valley (Nathan Cullen): No adjustment for now since there is no reasonable base (this riding has been represented by Cullen since its creation), but I will be less hesitant to adjust this riding away from the NDP based on riding polls. CON +1.6, NDP -5, LIB +2.9, GRN +0.5. No new information, but I took another look at the situation by examining how things have evolved since the riding was created in 2004, and Cullen 2015 seemed quite a bit more popular than Cullen 2004. Thus, an adjustment is in order even without comparing to pre-Cullen days. This adjustment is cautious: much like in Kings--Hants, I'm going with just half of the apparent effect due to how long the retiring MP has been around.

f) New star candidates running in 2019

Laurier--Sainte-Marie (Steven Guilbault): No (further - see above) adjustment for now, but this is a riding where I will be less hesitant to adjust in the Liberal direction based on riding polls. And we finally have a riding poll in this riding, which suggests a sizeable Guilbeault effect: NDP -7, LIB +9, GRN -1, BQ -1. (Combined with the Duceppe adjustment above, it's NDP -5, LIB +10, GRN -1, BQ -4 in this riding.)

Windsor West (Sandra Pupatello): NDP -12.5 -11.3, LIB +14 +12.8, GRN -1.5 to 2015 baseline, and additionally NDP +1.25, LIB +1.25, GRN -3 -2.5 to projection. The reason for a two-step adjustment is that I couldn't fully adjust the Greens in the baseline without going negative. This adjustment is guided by this poll (but is classified as a model rather than poll-based adjustment since it's clearly due to candidate's star status). This riding goes from safe NDP to a close NDP/LIB race. Update Sept. 8: Adjustments modified to avoid double-counting regional adjustment to Southwestern ON.

Thursday, August 1, 2019

Why Was the 2015 Projection Off?

Unlike in past elections, where a uniform swing model (applied regionally and with suitable adjustments based on additional information) would have performed very well using actual vote shares, both polling misses (mainly in Québec) and model shortcomings played a significant role in the 2015 projection being quite off. Unfortunately, those sources of error happened to push in the same direction.

On actual numbers for each of the 6 standard polling regions of the country, the model would have projected 168 Liberals, 114 Conservatives, 51 NDP, 4 Bloc and 1 Green. That is quite a bit better than the actual projection, but still farther away than had been the case in earlier elections (where misses with actual numbers tended to be around 5 seats for the main parties, rather than around 15).

Here's an analysis of what was off for each polling region. As you can see, there isn't much of a common thread across the country. The only issue that is somewhat recurring is that vote shares tend to move less in seats where a party has either very low or very high support (which makes sense: you can't drop by much if you're already really low, and you have fewer potential new supporters if you're already really high), and thus tend to move a bit more in other seats. Thus, the 2019 model will attempt to take this into account. A sneak preview is that, without any adjustments (e.g. for specific ridings or areas such as the 905), it would have projected 181 Liberals, 114 Conservatives, 41 NDP, 1 Bloc and 1 Green on 2015 numbers. (Update Aug. 9: The figures have changed as the 2019 model has been slightly modified.)

Projected Actual [Projection on actual popular vote, without ad hoc last-minute adjustments]

Atlantic Canada
LIB - 26 (53.5%) 32 (58.7%) [27]
CON - 3 (21.1%) 0 (19.0%) [3]
NDP - 3 (20.2%) 0 (18.0%) [2]
GRN - 0 (4.2%) 0 (3.5%) [0]

In Atlantic Canada, the Liberals outperformed the polls. Moreover, the model included an adjustment based on historical results that had New Brunswick partly swinging with the rest of Atlantic Canada, and partly with Ontario. That backfired, as the Liberal swing in NB was actually similar to the swing in the rest of Atlantic Canada (and much higher than in ON). Finally, the Liberals gained more in tight ridings than in blowouts. Those gains were from both the NDP and Tories, so I don't see this as evidence of strategic voting.

Lessons: Simply model NB with the rest of Atlantic Canada. If a party is gaining in a region and has extremely high support in some ridings, further gains in those ridings may be smaller - and thus, gains in other ridings may be larger and bring more seats.

Québec
NDP - 31 (26.3%) 16 (25.3%) [17]
LIB - 26 (30.1%) 40 (35.6%) [47]
CON - 11 (20.5%) 12 (16.8%) [10]
BQ - 10 (19.4%) 10 (19.4%) [4]
GRN - 0 (2.6%) 0 (2.3%) [0]

In Québec, the Liberals again outperformed the polls. Note that here, with actual results, the projection would actually have given them too many seats! The Bloc did exactly as I projected, but the model would have given it too few seats had the actual percentages been used for other parties.

The reasons why the model would have been pretty off even with actual vote percentages are a bit complicated. In the 450 area, the Bloc lost less support than the model expected while the NDP lost more. In the 514 area, the NDP held up better than expected, while the Liberals gained less. Put these together, and you have the Bloc winning more seats than expected, and the Liberals winning fewer.

Also, in the Quebec City area, the Conservatives did much better than expected (making big gains relative to 2011 while actually sliding in the rest of the province) at the expense of everyone else.

Lessons: Perhaps the smaller changes in the 514 could be picked up by a model with regional elasticities as the 514 vote often moves less than the rest of Québec. But I'm not sure how it would have been possible to project what happened in the 450 and the Quebec City area without access to regional polling breakdowns...

Ontario
LIB - 71 (44.3%) 80 (44.8%) [73]
CON - 36 (34.3%) 33 (35.0%) [34]
NDP - 14 (16.7%) 8 (16.6%) [14]
GRN - 0 (3.7%) 0 (2.9%) [0]

In Ontario, the polls were pretty much bang on, at least after my pro-CPC turnout adjustment! The seat projection, though, was too charitable to the NDP, who didn't resist as well as expected. The problem was almost entirely in Toronto and Northern Ontario. Was it strategic voting in these strong NDP/Liberal areas? The Liberals targeting smartly? Or was it simply that the NDP had more votes to lose?

The model did very well in the rest of the province. An adjustment based on regional breakdowns from polls amplifying the swing in the 905 area (from Tories to Liberals) helped, as raw uniform swing would have left the Tories with too many seats (around 40).

Lesson: Assuming somewhat proportional swing, rather than uniform swing, would have helped predict bigger NDP drops in Toronto and Northern Ontario.

Manitoba/Saskatchewan
CON - 17 (41.7%) 15 (43.0%) [17]
LIB - 5 (32.6%) 8 (34.4%) [5]
NDP - 6 (20.7%) 5 (19.1%) [6]
GRN - 0 (4.1%) 0 (2.7%) [0]

In the Prairies, the polls did well too. The Liberals picked up 3 more seats than the model expected, all in the Winnipeg area. In those 3 seats (and nowhere else in MB/SK), the Liberals outperformed the projection by more than 10 points! And this didn't really look like strategic voting as the gains were from both the Tories and the NDP.

In fact, the Liberals overperformed across the board in MB, and underperformed across the board in SK, while it was the reverse for the Conservatives. Not exactly sure why - the provincial Tories were quite popular in MB at the time. The Liberals usually get more votes in MB than SK, but the gap has exploded: ~4% in 2006 and 2008, ~8% in 2011, ~21% in 2015. The Tories went from just 3% more in SK than MB in 2011 to 11% more in 2015. Perhaps SK went from "more like MB" to "more like AB" due to its resource boom?

Lesson: Not sure... This looks like a case of two provinces no longer moving together politically, but unfortunately lumped together in most polls.

Alberta
CON - 32 (56.3%) 29 (59.6%) [32]
LIB - 1 (24.4%) 4 (24.6%) [1]
NDP - 1 (15.0%) 1 (11.6%) [1]
GRN - 0 (3.3%) 0 (2.5%) [0]

The story is pretty straightforward and intuitive in AB: Liberals support picked up more in cities, which is where they had a chance of winning seats.

Lesson: Here, the proportional model would have worked better if you looked at the Liberal number. But it would have worked worse if you looked at the Conservative number: Tory support dropped more in ridings where they were already weaker. So really, the lesson is that a party's support moves most when it's neither too high nor too low.

British Columbia
CON - 20 (32.1%) 10 (30.0%) [18]
NDP - 11 (25.1%) 14 (25.9%) [11]
LIB - 10 (32.8%) 17 (35.2%) [12]
GRN - 1 (8.9%) 1 (8.2%) [1]

In BC, the Liberals slightly outperformed the polls, at the expense of the Conservatives. The bigger issue, though, was that the model was way too generous to the Conservatives seat-wise. Here, we have perhaps the patterns most suggestive of strategic voting and/or excellent targeting by Liberals/NDP. The Liberals way underperformed uniform swing in Vancouver Island and the Kootenays, where it is weaker than the NDP, and somewhat exceeded it elsewhere. Where Liberals underperformed, the Greens were the main beneficiaries EXCEPT in tight NDP-Conservative races, where the NDP benefited. This helped the NDP win/solidify seats on Vancouver Island and the Kootenays.

The above paragraph also means that the Liberals overperformed elsewhere, which helped the Liberals win seats. That alone does not, however, explain why the Liberals gained so many extra seats. It appears that the far suburbs of Vancouver (seats like Mission--Matsqui--Fraser Canyon, Pitt Meadows--Maple Ridge and Cloverdale--Langley City) swung especially strongly toward the Liberals.

Also, on the basis of regional breakdowns from polls, I had the NDP being weaker than uniform swing on Vancouver Island, to the benefit of the Greens. As it turns out, while this was correct for the Greens, it was wrong for the NDP - it was the Liberals that undershot there. This was significant as it made the model incorrectly give two NDP seats to the Tories. I also made an adjustment based on a riding poll that incorrectly shifted a seat from the Liberals to the Tories.

Finally, the model did not account for James Moore leaving politics and the Greens not running in Kelowna--Lake Country, both of which likely added seats to the Liberals.

So, in short, uniform swing would have been somewhat off in BC, perhaps due to strategic voting. Moreover, counterproductive inclusion of regional breakdown/riding polls, as well as idiosyncratic riding-level factors, happened to push in the same direction. As a result, the model was frankly embarrassingly inaccurate in BC.

Lesson: Not sure how to guard against inaccurate regional breakdowns, other than ignoring them. But then again regional breakdowns elsewhere (e.g. in Québec) would have been so useful! Other than that, it looks like strategic voting might have been a substantial phenomenon in BC.

Wednesday, July 31, 2019

Criteria for Poll Inclusion

As some of you may have noticed, some pollsters are pushing back against their polling results being used in seat projections. To the extent that some media outlets may be free-riding on their work for financial gain, I understand their frustration. On the other hand, if these firms release their political polling results publicly, they gain notoriety, which may help, for example, their market research business. So one could argue that polling firms can't have it both ways, especially given the fair dealing provision of Canadian copyright law.

This argument poses an obvious ethical dilemma for this blog. First, I would like to point out that I am not making a penny from making these projections: this blog is not monetized (as you can see, there are no ads), and because I am doing this anonymously, there is no prospect of a media outlet hiring me.

Nevertheless, as a producer of knowledge in my real job, I sympathize with pollsters' desire to protect their intellectual property. Therefore, going forward, my projections will only include polls satisfying both of these conditions:
1. the poll results have been publicly disclosed by the polling firm, a member of the polling firm, or a person/organization that directly obtained permission to disclose the results (such as a news outlet that commissioned the poll), and
2. it is not the case that every disclosure covered under #1 states that the use of the results for seat projections is prohibited.

For example, I will not be using:
- polling results that are behind a paywall and leaked (including regional results, even when the same poll's national results are public);
- results from a pollster that does not want their numbers fed into a projection model, if this is clearly indicated in ALL disclosures of poll results by the pollster or an authorized entity.

However, I will feel free to use:
- polling results that were put behind a paywall after being publicly disclosed;
- results that are disclosed (by the pollster or an authorized entity) without a mention that they should not be used in a projection model, even if other disclosures of that same poll carry such a mention (e.g. pollster says "don't use these," but their client news organization doesn't).

To show appreciation to pollsters that remain open to having their numbers used (which is still, fortunately, most of them), I will, for any poll used in the projection:
- refrain from giving the poll's top line numbers on this blog (except in rare cases where it is necessary for my commentary);
- link, as much as possible, either to the pollster's website or to the website of a news outlet having commissioned the poll (rather than directly to the PDF or to a third party source);
- avoid, to the best of my ability, providing commentary on that poll that overlaps a lot with the pollster's commentary on that poll. (I may still provide commentary on aspects the pollster is unlikely to comment on, such as how the poll compares with my polling average.)

I hope that these measures will encourage you, the reader, to visit the pollster or the commissioning news agency's website. If pollsters or their clients benefit from projection websites such as this one, they will be less inclined to restrict the use of their numbers or hide them behind a paywall.

Finally, from time to time, I will be removing from the left right-hand column of this page pollsters that restrict, on their own website, the use of both their latest poll and all federal voting intention polls they conducted within the past month. Some polls from such firms may still be included in the projection if, for example, a news outlet commissioning them made the results available.

New Poll Weighting Methodology

Update Aug. 26: This post discusses national weights. Regional weights are discussed here.

One of the main methodological changes that I have made to the model concerns the weighting of polls. In the old formula, poll weights depended on sample size, recency and whether the same pollster has a more recent poll. Going forward, they will also depend on the presence of ALL other polls.

The old formula consisted of a number of ad hoc discount factors based on what seems sensible. It seemed to work OK most of the time, but when shifts in poll numbers occurred, one always wondered if the formula was too slow/quick to react.

The new weighting formula tries to approximate the optimal variance-minimizing (linear) formula under certain straightforward assumptions. That is, it is grounded in statistical theory rather than just being what seems reasonable to me. These assumptions will also allow me to propose approximate confidence intervals (to be explained in another post) not just for "what would happen if an election took place at the same time as the most recent poll," but also for Election Day.

Unfortunately, the new weights are derived through solving a system of (linear) equations (one equation per poll), so I can't tell you exactly when a poll would be discounted by 30% vs. 50%. Instead, I'll point out some notable effects of the changes:

- An old poll will be now discounted more aggressively if there are many other old polls. This makes sense as old polls' errors relative to current voting intentions are correlated through the change in voting intention since they were conducted. This can be very significant: the 6/27-7/2 Mainstreet poll of 2,651 would have a weight of over 50% of the most recent poll (Forum 7/26-28, 1,733 respondents) if they were the only two polls. But due to all the intervening polls, that Mainstreet poll currently counts less than 1/8 as much as the Forum poll.

As a result of this change, I will not need to change the formula to discount more aggressively when polls become more frequent: the model will automatically take care of this, and do it right.

- In certain cases, a poll can have negative weight! This appears surprising at first (my reflex was that I made a programming mistake), but it's actually not that hard to understand. Suppose there are two pollsters, A and B. Pollster A has conducted one recent poll and one old poll. Pollster B has conducted only one old poll. In order to guard against pollster A's potential in-house bias, the model wants to put significant weight on pollster B's poll. But that might put too much weight on an old data point, so it may be optimal to assign a slightly negative weight to pollster A's old poll.

Here are the details. I assume that polls have four potential independent sources of error for estimating current support:
1. Sampling variance: this is the pure statistical error, which is what a poll's "margin of error" refers to.
2. Changes in public opinion since the field dates: I mostly assume that voting intentions follow a random walk (a slight adjustment is made for potential short-term momentum).
3. The specific pollster's in-house bias: pollsters vary in their methodology.
4. Bias common to all pollsters: pollsters' methodologies may have common flaws.

Of course, there's no way to reduce #4 through averaging polls, so the weighting formula ignores it. (But it is very important in getting the right confidence intervals.)

#2 implies that any two polls' errors have a common component: the evolution of public opinion since the more recent poll. The more a poll's error is correlated with other polls', the less informative it is, and the less it should be weighted. Therefore, all polls' weights should depend on each other (over and above other polls "changing the denominator"). How much this is the case depends on how fast one thinks public opinion evolves.

Voting intentions tend to be much more volatile during a campaign - especially late - than before a campaign. Therefore, I will compute each poll's normalized age a by valuing the number of calendar days since the poll's median field date ("calendar age") as follows (updated with the tentative date of the first debate: October 7):
- Before September 1: 0.1
- September 1-20, before writs are issued: 0.2
- September 1-20, after writs are issued: 0.5
- September 21-30: 1
- October 1-7: 1.5
- October 8-21: 2
For example, a July 15 poll will be considered 3 days old on August 14. A September 15 poll will be considered 7.5 days old on September 25 (5*0.5 + 5*1), as writs must be issued by September 15. An October 10 poll will be considered 14 days old on October 17.

I assume that the variance of the evolution of public opinion over a normalized time period of length a is V(a) = 0.0001a. Assuming a normal distribution, this implies that, in a given week, for a given main party (i.e. 30%+ in the polls), there is approximately a 30% chance of a change in support exceeding:
- 0.8% before September 1
- 2.6% toward the end of September
- 3.7% in October, after the debate
I didn't do any formal analysis for this calibration, but to me, eyeballing how things evolved in the past two campaigns, this passes the smell test.

This formula implies that changes in voting intentions are uncorrelated over time. On a weekly scale (which is the scale of my "eyeball calibration"), the assumption is not too crazy. But on a daily scale, it's a poor assumption: "momentum" is clearly a thing when key movements happen during a campaign. To partly correct for this:
- Before the writs are issued, I adjust a poll's calendar age down by 1 day if it's older than 1.5 days, and down by 2/3 if it's less than 1.5 days old.
- After the writs are issued, I adjust a poll's calendar age down by 2 days if it's older than 4.5 days, down by 1-2 days if it's 1.5-4.5 days old, and down by 2/3 if it's less than 1.5 days old.

Finally, factor #3 is taken into account by assigning a individual-pollster variance of 0.0001 0.0002 (updated per this Aug. 10 post) (that is, 1 ~1.4% standard deviation). I don't have any great confidence in this parameter value, so let me know if you think it's unreasonable. But remember that this does not take into account any sampling variance, so it's normal that polls (even those in the field at the same time) vary by much more than this implies. The effect of this is that a poll will have more weight if there are fewer other polls - especially if there are no more recent polls - by the same firm. (The old formula crudely approximated this by applying a 50% discount to any poll that is not the firm's most recent one.)

What I'm NOT doing (and would ideally be doing if I had unlimited time)
- Except for the small adjustment mentioned above, I do not take into account serial correlation in changes in voting intention (in particular, I do not project forward from past trends), and the parameters I use are from "eyeballing" past patterns rather than a rigorous analysis.
- I am not adjusting pollsters' results for bias. That is, if pollster X consistently has better results for Party A than other pollsters, I am not adjusting X's numbers for A downward. Getting such adjustments right would require an analysis of recent years' polls for which I don't have time. Moreover, they should roughly cancel out across pollsters at crucial times in the campaign when most pollsters publish a poll. At quieter times, though, this will cause the average to be a bit more topsy turvy than it should be.
- I am not grading the quality of pollsters (and accordingly modifying the weights). As above, lack of time.
- I will maintain an ad hoc approach when incorporating provincial/regional polls. They contain useful information, so I won't ignore them and will "eyeball" the weight they should receive. However, dealing with them systematically (in particular, separately computing the weights on the provincial numbers of national polls) doesn't seem worth the trouble. I now deal with regional polls systematically (though not quite optimally). Details in the regional weights post.
- Riding polls will not factor into polling averages (though they can inform riding-level adjustments).

Sunday, July 28, 2019

Comparison of 2015 Projections

This is a post I started writing after the 2015 election, but never finished. I provide it in its unfinished state, "for the record."

2015 Final Projections (italics = not model-based)
178-115-44-  1-0-0 (37%, 31%, 22%, 4%, 4%) CVM Election Model
177-  95-53-11-1-1 (38.0%, 30.9%, 21.3%, 5.7%, 3.4%) Teddy on Politics
160-120-50-  7-1-0 (36.7%, 32.0%, 20.4%, 5.2%, 4.0%) The Signal
149-105-81-  2-1-0 David Akin's Predictionator
147-115-67-  7-2-0 (39.3%, 32.4%, 19.4%, 4.1%, 4.1%) Sauder Prediction Market
146-118-66-  7-1-0 (37.2%, 30.9%, 21.7%, 4.9%, 4.4%) ThreeHundredEight
142-119-66-10-1-0 (37.3%, 32.4%, 20.1%, 4.9%, 4.3%) Canadian Election Watch
142-116-68-11-1-0 Election Atlas
140-115-79-  3-1-0 LISPOP
138-117-76-  6-1-0 (36.8%, 31.8%, 22.7%, 4.2%, 4.5%) Le calcul électoral
138-120-75-  1-1-0 Election Almanac
137-120-75-  8-1-0 (36.8%, 32.5%, 21.4%, 4.6%, 4.1%) Too Close to Call
128-120-83-  5-2-0 Election Prediction Project

Sum of absolute seat deviations (divided by 2)
11 Teddy on Politics
16 CVM Election Model
27 The Signal
40 Sauder Prediction Market
41 ThreeHundredEight
42 Canadian Election Watch
42 Election Atlas
43 David Akin's Predictionator
49 Too Close to Call
50 Le calcul électoral
51 LISPOP
55 Election Almanac
61 Election Prediction Project

This time, I was in the middle of the pack. (Here is how I did in 2011.) Note that my projection without the turnout adjustment was off by 34, which is significantly better than average.

Three models did particularly well by this measure. If I understand correctly, Teddy on Politics injects a fair bit of personal judgment into his projections - they're not as heavily based on polls as the other projections. This time, he got things very right (which happens fairly often, if you've been following him, although there are also big misses), foreseeing which way the wind was blowing. The CVM election model appeared to have been lucky: the seat model was too extreme, but the Liberal support was underestimated, and the two errors canceled out. For example, it correctly projected the Liberal Atlantic sweep, but on percentages with which the sweep would not have happened (as the Liberals won some seats narrowly). And it incorrectly predicted an NDP wipeout in Ontario, which didn't even come close to happening.

This leaves us with The Signal, which did very well. A big part of this is that it correctly foresaw that the Liberal Québec vote would be quite efficient. Unfortunately, its description of methodology is not very detailed, so it is hard to see why exactly it did better than everyone else - especially given that it had the lowest Liberal popular vote projection. Still, congrats to The Signal!

Average absolute popular vote deviation for 5 parties
0.46 Sauder Prediction Market
0.84 Canadian Election Watch
0.94 The Signal
1.02 Teddy on Politics
1.16 Too Close to Call
1.30 ThreeHundredEight
1.40 CVM Election Model
1.48 Le calcul électoral

By this measure, I provided the best poll-based projection. The unadjusted numbers were off by 0.90 on average. The polls did quite well this time, and adjusting polls to reflect the latest trends helped. In fact, had my adjustments not been so cautious, the vote (and seat) projections would have been even more accurate - perhaps close to the almost-on-the-dot Sauder market prediction.

Speaking of the Sauder prediction market, feeding its popular vote prediction into most seat models would probably have produced more accurate seat counts than it predicted. This suggests that while market participants were good at guessing the overall vote, the seat count conversion proved challenging. This is why quantitative seat models help!

Number of ridings correctly projected
275 Teddy on Politics
269 ThreeHundredEight
268 Canadian Election Watch
261 Election Prediction Project

Friday, July 26, 2019

2015 Result Map

Here is a map of the results of the 2015 General Election. I am getting the blog ready for the 2019 General Election, so stay tuned!