sociological methods & research size matters: quantifying...

33
Article Size Matters: Quantifying Protest by Counting Participants Michael Biggs 1 Abstract Since the 1970s, catalogs of protest events have been at the heart of research on social movements. To measure how protest changes over time or varies across space, sociologists usually count the frequency of events, as either the dependent variable or a key independent variable. An alter- native is to count the number of participants in protest. This article investigates demonstrations, strikes, and riots. Their size distributions manifest enormous variation. Most events are small, but a few large events contribute the majority of protesters. When events are aggre- gated by year or by city, the correlation between total participation and event frequency is low or modest. The choice of how to quantify protest is therefore vital; findings from one measure are unlikely to apply to another. The fact that the bulk of participation comes from large events has positive implications for the compilation of event catalogs. Rather than worrying about the underreporting of small events, concentrate on recording large ones accurately. Keywords protest, social movements, strikes, riots, demonstrations 1 Department of Sociology, University of Oxford, Oxford, UK Corresponding Author: Michael Biggs, Department of Sociology, University of Oxford, Oxford OX1 3UQ, UK. Email: [email protected] Sociological Methods & Research 2018, Vol. 47(3) 351-383 ª The Author(s) 2016 Reprints and permission: sagepub.com/journalsPermissions.nav DOI: 10.1177/0049124116629166 journals.sagepub.com/home/smr

Upload: others

Post on 17-Jul-2020

2 views

Category:

Documents


0 download

TRANSCRIPT

Page 1: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Article

Size Matters QuantifyingProtest by CountingParticipants

Michael Biggs1

Abstract

Since the 1970s catalogs of protest events have been at the heart of researchon social movements To measure how protest changes over time orvaries across space sociologists usually count the frequency of events aseither the dependent variable or a key independent variable An alter-native is to count the number of participants in protest This articleinvestigates demonstrations strikes and riots Their size distributionsmanifest enormous variation Most events are small but a few largeevents contribute the majority of protesters When events are aggre-gated by year or by city the correlation between total participation andevent frequency is low or modest The choice of how to quantify protestis therefore vital findings from one measure are unlikely to apply toanother The fact that the bulk of participation comes from large eventshas positive implications for the compilation of event catalogs Ratherthan worrying about the underreporting of small events concentrate onrecording large ones accurately

Keywords

protest social movements strikes riots demonstrations

1 Department of Sociology University of Oxford Oxford UK

Corresponding AuthorMichael Biggs Department of Sociology University of Oxford Oxford OX1 3UQ UKEmail michaelbiggssociologyoxacuk

Sociological Methods amp Research2018 Vol 47(3) 351-383

ordf The Author(s) 2016Reprints and permission

sagepubcomjournalsPermissionsnavDOI 1011770049124116629166

journalssagepubcomhomesmr

For the study of social movements the crucial empirical innovation was thecatalog of protest events Such catalogs really originated in the late nine-teenth century when governmentsmdashin response to the emerging labor move-mentmdashbegan publishing statistics on strikes (Franzosi 1989) Other forms ofprotest however were not subject to official investigation American socialscientists began collecting their own data on events from the 1960s Inpolitical science cross-national time series of protest and violence werecompiled from the New York Times Index (eg Hibbs 1973 Rummel1963) In sociology Tilly used newspapers to catalog lsquolsquocontentious gather-ingsrsquorsquo in Britain in the late eighteenth century and early nineteenth century(1978appendix 3 1995) Tillyrsquos work proved enormously influential Eventcatalogs have been at the core of landmark studies of social movements (egKriesi et al 1995 McAdam 1982 Olzak 1992 Tarrow 1989) Large-scaleresearch projects have now compiled national data sets extending over sev-eral decades The United States is covered from 1960 to 1995 using the NewYork Times As data accumulate analyses proliferate Seven leading sociol-ogy journals since 2000 have published over 40 papers that quantify protestas either the dependent variable or a key independent variable To capturehow protest changes from year to year or how it varies across cities or statesthe standard procedure is to count the frequency of events1

Methodological scrutiny has focused on one problem The vast majorityof protest events are never reported by the news media and so the frequencyof protest is severely underestimated (Barranco and Wisler 1999 Earl et al2004 Franzosi 1987 Maney and Oliver 2001 McCarthy McPhail andSmith 1996 McCarthy et al 2008 Myers and Caniglia 2004 Oliver andManey 2000 Oliver and Myers 1999 Ortiz et al 2005) This lacuna leadsMyers and collaborators to argue that catalogs compiled from newspapersmdashespecially from a single newspapermdashhave lsquolsquovery serious flawsrsquorsquo (Ortiz et al2005398 also Myers and Caniglia 2004) Such skepticism is rare Mostsocial scientists analyzing protest events concur with the conclusions of Earlet al (200477) lsquolsquonewspaper data does not deviate markedly from acceptedstandards of qualityrsquorsquo Debate over sources of datamdashabout reliability ratherthan validitymdashhas displaced a more fundamental question How should pro-test be quantified

This article emphasizes a statistical property shared by almost all kinds ofprotest Events vary enormously in size An event may comprise a singlepersonrsquos action or it may combine the actions of a million participants Suchenormous variation would hardly matter if it tended to even out when manyevents are aggregated into time intervals or geographical units Then thefrequency of protest events would correlate highly with the total number

352 Sociological Methods amp Research 47(3)

of protesters and the choice between the two would not be significant Infact this article demonstrates a low or modest correlation in several data setsranging from 10 to 64 Therefore counting events and counting participantswill yield very different conclusions

Once the extreme variation in event size is appreciated it is clear that thelargest eventsmdashrelatively few in numbermdashcontribute the majority of totalparticipants That most events go unreported by the media is far less of aproblem if we measure total participation because these small events con-tribute so little to the total National series of strikes and demonstrations forexample are dominated by events of at least 10000 participants Effortshould therefore concentrate on accurately recording these large events byeliminating erroneous duplication and estimating their size more consis-tently this effort is feasible because there are so few of them

The theoretical rationale for quantifying protest by the number of parti-cipants is worth sketching briefly Conceptualizations of the object of socio-logical interest refer to actions For Tarrow (20117) lsquolsquo[t]he irreducible actthat lies at the base of all social movements protests rebellions riots strikewaves and revolutions is contentious collective actionrsquorsquo Opp (200938)defines protest as lsquolsquojoint (ie collective) action of individuals aimed atachieving their goal or goals by influencing the decisions of a targetrsquorsquo Socialmovements are conceived by Oliver and Myers (20033) as lsquolsquopopulations ofcollective actionsrsquorsquo By implication actions should be quantified The mostbasic measure is the number of protest actions in other words the number ofparticipants in protest

The number of participants appears as a crucial variable in theories of howsocial movements bring about change DeNardo (198536-37) takes lsquolsquothedisruptiveness of protests demonstrations and uprisings to be first and fore-most a question of numbersrsquorsquo thus one key parameter is lsquolsquothe percentage ofthe population mobilizedrsquorsquo (see also Lohmann 1994) According toOberschall (199480) lsquolsquothe crucial resource for obtaining the collective goodis the number of participants or contributorsrsquorsquo Practitioners echo this pointlsquolsquoRemember in a nonviolent struggle the only weapon that yoursquore going tohave is numbersrsquorsquo (Popovic and Miller 201552) The importance of numbersis reinforced by considering the orchestration of contentious gatherings Tilly(1995370) postulates that leaders are lsquolsquomaximizing the multiple of fourfactors numbers [of participants] commitment worthiness and unityrsquorsquoWhen demonstrators converge on one location (such as the capital city) ona single day they are clearly maximizing the number of participants If theywanted to maximize the frequency of events then they would disperse todifferent places to perform separate demonstrations at different times

Biggs 353

Indeed the size of the demonstrations would not mattermdashten demonstrationsof a dozen people would be preferable to a single demonstration of onehundred thousand

Several data sets are used to develop the argument First is Dynamics ofCollective Action in the United States 1960ndash1995 (DCAUS) compiled fromthe New York Times (McAdam et al nd) This has been used in a score ofarticles By contrast the European Protest and Coercion Dataset (EPCD)spanning the years 1980 to 1995 is unduly neglected (Francisco nd) UnlikeDCAUS it is derived from multiple newspapers and newswires The UnitedKingdom as the most similar country to the United States is selected forcomparison Both data sets encompass heterogeneous events from prisonriots to press conferences from kneecapping to general strikes Lumpingthese together makes little sense These two data sets also cover a differentrange of events routine strikes are included in EPCD but excluded fromDCAUS Therefore I will focus on demonstrations including marches ral-lies and vigils In DCAUS demonstrations contribute half the total numberof participants in EPCD for the United Kingdom a third of the total Thesedata sets are usefully compared to a comprehensive catalog of demonstra-tions in Washington DC in 1982 and 1991 compiled primarily from therecords of three police forces (McCarthy et al 1996) Although the authorshave not made the data available published tabulations are valuable forrevealing the mass of small events that never make the news Strikes areparticularly valuable for my purpose because they are less prone to under-reporting and because their size is measured more accurately The UnitedKingdom consistently tabulated the size distribution from 1950 to 19842 TheUnited States recorded every strike from 1881 to 1893 which permits anal-ysis of a subset of events in small cities and towns The final data set isCarterrsquos catalog of black riots in the United States from 1964 to 1971(1986220-21) This allows investigation of variation across cities to com-plement the time series for strikes and demonstrations

This article begins by reviewing the quantification of protest events inrecent literature The second section conceptualizes various ways of measur-ing the size of protest events and outlines the challenge of measuring par-ticipation Size distributions of demonstrations strikes and riots arepresented in the third section These distributions are heavy-tailed meaningthat most participants are concentrated in a few huge events How does thismatter The fourth section compares event frequency and total participationand shows that they are not strongly associated conclusions drawn from onecannot be assumed to apply to the other This negative finding is offset bypositive implications in the fifth section The problem of unreported small

354 Sociological Methods amp Research 47(3)

events diminishes large events provide a remarkably accurate measure oftotal participation The conclusion draws implications for future research

Protest Events in the Literature

Fifty years ago Tilly and Rule (1965) wrote at length on Measuring PoliticalUpheaval Their major precedent was strikes which were quantified in threeways (elegantly explicated by Spielmans 1944) the frequency of events thetotal number of participants workers involved in strikes and the total num-ber of participant-days working-days lost in strikes Tilly and Rule(196574) concluded that lsquolsquo[t]he most useful general conception of the mag-nitude of a political disturbance seems to be the sum of human energyexpended in itrsquorsquo This they argued was best approximated by participant-days In the subsequent half century data on events have accumulated whileconsideration of what to quantify has disappeared A review article from1989 devotes a single page to this issue (Olzak 1989127) and it is notmentioned in a subsequent review (Earl et al 2004)

How then is protest quantified in recent literature Consider articlespublished from 2000 to 2014 in seven leading Anglophone journals Amer-ican Journal of Sociology American Sociological Review British Journal ofSociology European Sociological Review Mobilization Social Forces andSocial Problems (following Amenta et al 2010)3 Articles are selected if theymeasure protest as either the dependent variable or an independent variableor both4 This excludes variables defined by the occurrence of protest ratherthan its magnitudemdashas for example the dates at which cities experiencerioting used in event-history analysis5 Also excluded are studies that takethe event as the unit of observation Predominantly qualitative articles thatdefine the explanans by graphing a time series are included (eg Biggs2013) The literature search yields 41 articles (Online Table S1) Most arti-cles construct a time series at the national level which is usually annual butin a few cases quarterly or monthly Some articles measure how protestvaries across cities or other geographical units A few combine time andspace so the unit of observation is the city-year for example There are singleinstances of other combinations

The quantification of protest requires two choices The first is what sort ofprotest events to cover Some articles focus on distinct types of protest suchas petitions or demonstrations or strikes or riots Other articles aggregatediverse types of action from sit-ins to litigation into a single variable (egJacobs and Kent 2007) Whether it is conceptually meaningful to aggregate

Biggs 355

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 2: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

For the study of social movements the crucial empirical innovation was thecatalog of protest events Such catalogs really originated in the late nine-teenth century when governmentsmdashin response to the emerging labor move-mentmdashbegan publishing statistics on strikes (Franzosi 1989) Other forms ofprotest however were not subject to official investigation American socialscientists began collecting their own data on events from the 1960s Inpolitical science cross-national time series of protest and violence werecompiled from the New York Times Index (eg Hibbs 1973 Rummel1963) In sociology Tilly used newspapers to catalog lsquolsquocontentious gather-ingsrsquorsquo in Britain in the late eighteenth century and early nineteenth century(1978appendix 3 1995) Tillyrsquos work proved enormously influential Eventcatalogs have been at the core of landmark studies of social movements (egKriesi et al 1995 McAdam 1982 Olzak 1992 Tarrow 1989) Large-scaleresearch projects have now compiled national data sets extending over sev-eral decades The United States is covered from 1960 to 1995 using the NewYork Times As data accumulate analyses proliferate Seven leading sociol-ogy journals since 2000 have published over 40 papers that quantify protestas either the dependent variable or a key independent variable To capturehow protest changes from year to year or how it varies across cities or statesthe standard procedure is to count the frequency of events1

Methodological scrutiny has focused on one problem The vast majorityof protest events are never reported by the news media and so the frequencyof protest is severely underestimated (Barranco and Wisler 1999 Earl et al2004 Franzosi 1987 Maney and Oliver 2001 McCarthy McPhail andSmith 1996 McCarthy et al 2008 Myers and Caniglia 2004 Oliver andManey 2000 Oliver and Myers 1999 Ortiz et al 2005) This lacuna leadsMyers and collaborators to argue that catalogs compiled from newspapersmdashespecially from a single newspapermdashhave lsquolsquovery serious flawsrsquorsquo (Ortiz et al2005398 also Myers and Caniglia 2004) Such skepticism is rare Mostsocial scientists analyzing protest events concur with the conclusions of Earlet al (200477) lsquolsquonewspaper data does not deviate markedly from acceptedstandards of qualityrsquorsquo Debate over sources of datamdashabout reliability ratherthan validitymdashhas displaced a more fundamental question How should pro-test be quantified

This article emphasizes a statistical property shared by almost all kinds ofprotest Events vary enormously in size An event may comprise a singlepersonrsquos action or it may combine the actions of a million participants Suchenormous variation would hardly matter if it tended to even out when manyevents are aggregated into time intervals or geographical units Then thefrequency of protest events would correlate highly with the total number

352 Sociological Methods amp Research 47(3)

of protesters and the choice between the two would not be significant Infact this article demonstrates a low or modest correlation in several data setsranging from 10 to 64 Therefore counting events and counting participantswill yield very different conclusions

Once the extreme variation in event size is appreciated it is clear that thelargest eventsmdashrelatively few in numbermdashcontribute the majority of totalparticipants That most events go unreported by the media is far less of aproblem if we measure total participation because these small events con-tribute so little to the total National series of strikes and demonstrations forexample are dominated by events of at least 10000 participants Effortshould therefore concentrate on accurately recording these large events byeliminating erroneous duplication and estimating their size more consis-tently this effort is feasible because there are so few of them

The theoretical rationale for quantifying protest by the number of parti-cipants is worth sketching briefly Conceptualizations of the object of socio-logical interest refer to actions For Tarrow (20117) lsquolsquo[t]he irreducible actthat lies at the base of all social movements protests rebellions riots strikewaves and revolutions is contentious collective actionrsquorsquo Opp (200938)defines protest as lsquolsquojoint (ie collective) action of individuals aimed atachieving their goal or goals by influencing the decisions of a targetrsquorsquo Socialmovements are conceived by Oliver and Myers (20033) as lsquolsquopopulations ofcollective actionsrsquorsquo By implication actions should be quantified The mostbasic measure is the number of protest actions in other words the number ofparticipants in protest

The number of participants appears as a crucial variable in theories of howsocial movements bring about change DeNardo (198536-37) takes lsquolsquothedisruptiveness of protests demonstrations and uprisings to be first and fore-most a question of numbersrsquorsquo thus one key parameter is lsquolsquothe percentage ofthe population mobilizedrsquorsquo (see also Lohmann 1994) According toOberschall (199480) lsquolsquothe crucial resource for obtaining the collective goodis the number of participants or contributorsrsquorsquo Practitioners echo this pointlsquolsquoRemember in a nonviolent struggle the only weapon that yoursquore going tohave is numbersrsquorsquo (Popovic and Miller 201552) The importance of numbersis reinforced by considering the orchestration of contentious gatherings Tilly(1995370) postulates that leaders are lsquolsquomaximizing the multiple of fourfactors numbers [of participants] commitment worthiness and unityrsquorsquoWhen demonstrators converge on one location (such as the capital city) ona single day they are clearly maximizing the number of participants If theywanted to maximize the frequency of events then they would disperse todifferent places to perform separate demonstrations at different times

Biggs 353

Indeed the size of the demonstrations would not mattermdashten demonstrationsof a dozen people would be preferable to a single demonstration of onehundred thousand

Several data sets are used to develop the argument First is Dynamics ofCollective Action in the United States 1960ndash1995 (DCAUS) compiled fromthe New York Times (McAdam et al nd) This has been used in a score ofarticles By contrast the European Protest and Coercion Dataset (EPCD)spanning the years 1980 to 1995 is unduly neglected (Francisco nd) UnlikeDCAUS it is derived from multiple newspapers and newswires The UnitedKingdom as the most similar country to the United States is selected forcomparison Both data sets encompass heterogeneous events from prisonriots to press conferences from kneecapping to general strikes Lumpingthese together makes little sense These two data sets also cover a differentrange of events routine strikes are included in EPCD but excluded fromDCAUS Therefore I will focus on demonstrations including marches ral-lies and vigils In DCAUS demonstrations contribute half the total numberof participants in EPCD for the United Kingdom a third of the total Thesedata sets are usefully compared to a comprehensive catalog of demonstra-tions in Washington DC in 1982 and 1991 compiled primarily from therecords of three police forces (McCarthy et al 1996) Although the authorshave not made the data available published tabulations are valuable forrevealing the mass of small events that never make the news Strikes areparticularly valuable for my purpose because they are less prone to under-reporting and because their size is measured more accurately The UnitedKingdom consistently tabulated the size distribution from 1950 to 19842 TheUnited States recorded every strike from 1881 to 1893 which permits anal-ysis of a subset of events in small cities and towns The final data set isCarterrsquos catalog of black riots in the United States from 1964 to 1971(1986220-21) This allows investigation of variation across cities to com-plement the time series for strikes and demonstrations

This article begins by reviewing the quantification of protest events inrecent literature The second section conceptualizes various ways of measur-ing the size of protest events and outlines the challenge of measuring par-ticipation Size distributions of demonstrations strikes and riots arepresented in the third section These distributions are heavy-tailed meaningthat most participants are concentrated in a few huge events How does thismatter The fourth section compares event frequency and total participationand shows that they are not strongly associated conclusions drawn from onecannot be assumed to apply to the other This negative finding is offset bypositive implications in the fifth section The problem of unreported small

354 Sociological Methods amp Research 47(3)

events diminishes large events provide a remarkably accurate measure oftotal participation The conclusion draws implications for future research

Protest Events in the Literature

Fifty years ago Tilly and Rule (1965) wrote at length on Measuring PoliticalUpheaval Their major precedent was strikes which were quantified in threeways (elegantly explicated by Spielmans 1944) the frequency of events thetotal number of participants workers involved in strikes and the total num-ber of participant-days working-days lost in strikes Tilly and Rule(196574) concluded that lsquolsquo[t]he most useful general conception of the mag-nitude of a political disturbance seems to be the sum of human energyexpended in itrsquorsquo This they argued was best approximated by participant-days In the subsequent half century data on events have accumulated whileconsideration of what to quantify has disappeared A review article from1989 devotes a single page to this issue (Olzak 1989127) and it is notmentioned in a subsequent review (Earl et al 2004)

How then is protest quantified in recent literature Consider articlespublished from 2000 to 2014 in seven leading Anglophone journals Amer-ican Journal of Sociology American Sociological Review British Journal ofSociology European Sociological Review Mobilization Social Forces andSocial Problems (following Amenta et al 2010)3 Articles are selected if theymeasure protest as either the dependent variable or an independent variableor both4 This excludes variables defined by the occurrence of protest ratherthan its magnitudemdashas for example the dates at which cities experiencerioting used in event-history analysis5 Also excluded are studies that takethe event as the unit of observation Predominantly qualitative articles thatdefine the explanans by graphing a time series are included (eg Biggs2013) The literature search yields 41 articles (Online Table S1) Most arti-cles construct a time series at the national level which is usually annual butin a few cases quarterly or monthly Some articles measure how protestvaries across cities or other geographical units A few combine time andspace so the unit of observation is the city-year for example There are singleinstances of other combinations

The quantification of protest requires two choices The first is what sort ofprotest events to cover Some articles focus on distinct types of protest suchas petitions or demonstrations or strikes or riots Other articles aggregatediverse types of action from sit-ins to litigation into a single variable (egJacobs and Kent 2007) Whether it is conceptually meaningful to aggregate

Biggs 355

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 3: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

of protesters and the choice between the two would not be significant Infact this article demonstrates a low or modest correlation in several data setsranging from 10 to 64 Therefore counting events and counting participantswill yield very different conclusions

Once the extreme variation in event size is appreciated it is clear that thelargest eventsmdashrelatively few in numbermdashcontribute the majority of totalparticipants That most events go unreported by the media is far less of aproblem if we measure total participation because these small events con-tribute so little to the total National series of strikes and demonstrations forexample are dominated by events of at least 10000 participants Effortshould therefore concentrate on accurately recording these large events byeliminating erroneous duplication and estimating their size more consis-tently this effort is feasible because there are so few of them

The theoretical rationale for quantifying protest by the number of parti-cipants is worth sketching briefly Conceptualizations of the object of socio-logical interest refer to actions For Tarrow (20117) lsquolsquo[t]he irreducible actthat lies at the base of all social movements protests rebellions riots strikewaves and revolutions is contentious collective actionrsquorsquo Opp (200938)defines protest as lsquolsquojoint (ie collective) action of individuals aimed atachieving their goal or goals by influencing the decisions of a targetrsquorsquo Socialmovements are conceived by Oliver and Myers (20033) as lsquolsquopopulations ofcollective actionsrsquorsquo By implication actions should be quantified The mostbasic measure is the number of protest actions in other words the number ofparticipants in protest

The number of participants appears as a crucial variable in theories of howsocial movements bring about change DeNardo (198536-37) takes lsquolsquothedisruptiveness of protests demonstrations and uprisings to be first and fore-most a question of numbersrsquorsquo thus one key parameter is lsquolsquothe percentage ofthe population mobilizedrsquorsquo (see also Lohmann 1994) According toOberschall (199480) lsquolsquothe crucial resource for obtaining the collective goodis the number of participants or contributorsrsquorsquo Practitioners echo this pointlsquolsquoRemember in a nonviolent struggle the only weapon that yoursquore going tohave is numbersrsquorsquo (Popovic and Miller 201552) The importance of numbersis reinforced by considering the orchestration of contentious gatherings Tilly(1995370) postulates that leaders are lsquolsquomaximizing the multiple of fourfactors numbers [of participants] commitment worthiness and unityrsquorsquoWhen demonstrators converge on one location (such as the capital city) ona single day they are clearly maximizing the number of participants If theywanted to maximize the frequency of events then they would disperse todifferent places to perform separate demonstrations at different times

Biggs 353

Indeed the size of the demonstrations would not mattermdashten demonstrationsof a dozen people would be preferable to a single demonstration of onehundred thousand

Several data sets are used to develop the argument First is Dynamics ofCollective Action in the United States 1960ndash1995 (DCAUS) compiled fromthe New York Times (McAdam et al nd) This has been used in a score ofarticles By contrast the European Protest and Coercion Dataset (EPCD)spanning the years 1980 to 1995 is unduly neglected (Francisco nd) UnlikeDCAUS it is derived from multiple newspapers and newswires The UnitedKingdom as the most similar country to the United States is selected forcomparison Both data sets encompass heterogeneous events from prisonriots to press conferences from kneecapping to general strikes Lumpingthese together makes little sense These two data sets also cover a differentrange of events routine strikes are included in EPCD but excluded fromDCAUS Therefore I will focus on demonstrations including marches ral-lies and vigils In DCAUS demonstrations contribute half the total numberof participants in EPCD for the United Kingdom a third of the total Thesedata sets are usefully compared to a comprehensive catalog of demonstra-tions in Washington DC in 1982 and 1991 compiled primarily from therecords of three police forces (McCarthy et al 1996) Although the authorshave not made the data available published tabulations are valuable forrevealing the mass of small events that never make the news Strikes areparticularly valuable for my purpose because they are less prone to under-reporting and because their size is measured more accurately The UnitedKingdom consistently tabulated the size distribution from 1950 to 19842 TheUnited States recorded every strike from 1881 to 1893 which permits anal-ysis of a subset of events in small cities and towns The final data set isCarterrsquos catalog of black riots in the United States from 1964 to 1971(1986220-21) This allows investigation of variation across cities to com-plement the time series for strikes and demonstrations

This article begins by reviewing the quantification of protest events inrecent literature The second section conceptualizes various ways of measur-ing the size of protest events and outlines the challenge of measuring par-ticipation Size distributions of demonstrations strikes and riots arepresented in the third section These distributions are heavy-tailed meaningthat most participants are concentrated in a few huge events How does thismatter The fourth section compares event frequency and total participationand shows that they are not strongly associated conclusions drawn from onecannot be assumed to apply to the other This negative finding is offset bypositive implications in the fifth section The problem of unreported small

354 Sociological Methods amp Research 47(3)

events diminishes large events provide a remarkably accurate measure oftotal participation The conclusion draws implications for future research

Protest Events in the Literature

Fifty years ago Tilly and Rule (1965) wrote at length on Measuring PoliticalUpheaval Their major precedent was strikes which were quantified in threeways (elegantly explicated by Spielmans 1944) the frequency of events thetotal number of participants workers involved in strikes and the total num-ber of participant-days working-days lost in strikes Tilly and Rule(196574) concluded that lsquolsquo[t]he most useful general conception of the mag-nitude of a political disturbance seems to be the sum of human energyexpended in itrsquorsquo This they argued was best approximated by participant-days In the subsequent half century data on events have accumulated whileconsideration of what to quantify has disappeared A review article from1989 devotes a single page to this issue (Olzak 1989127) and it is notmentioned in a subsequent review (Earl et al 2004)

How then is protest quantified in recent literature Consider articlespublished from 2000 to 2014 in seven leading Anglophone journals Amer-ican Journal of Sociology American Sociological Review British Journal ofSociology European Sociological Review Mobilization Social Forces andSocial Problems (following Amenta et al 2010)3 Articles are selected if theymeasure protest as either the dependent variable or an independent variableor both4 This excludes variables defined by the occurrence of protest ratherthan its magnitudemdashas for example the dates at which cities experiencerioting used in event-history analysis5 Also excluded are studies that takethe event as the unit of observation Predominantly qualitative articles thatdefine the explanans by graphing a time series are included (eg Biggs2013) The literature search yields 41 articles (Online Table S1) Most arti-cles construct a time series at the national level which is usually annual butin a few cases quarterly or monthly Some articles measure how protestvaries across cities or other geographical units A few combine time andspace so the unit of observation is the city-year for example There are singleinstances of other combinations

The quantification of protest requires two choices The first is what sort ofprotest events to cover Some articles focus on distinct types of protest suchas petitions or demonstrations or strikes or riots Other articles aggregatediverse types of action from sit-ins to litigation into a single variable (egJacobs and Kent 2007) Whether it is conceptually meaningful to aggregate

Biggs 355

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 4: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Indeed the size of the demonstrations would not mattermdashten demonstrationsof a dozen people would be preferable to a single demonstration of onehundred thousand

Several data sets are used to develop the argument First is Dynamics ofCollective Action in the United States 1960ndash1995 (DCAUS) compiled fromthe New York Times (McAdam et al nd) This has been used in a score ofarticles By contrast the European Protest and Coercion Dataset (EPCD)spanning the years 1980 to 1995 is unduly neglected (Francisco nd) UnlikeDCAUS it is derived from multiple newspapers and newswires The UnitedKingdom as the most similar country to the United States is selected forcomparison Both data sets encompass heterogeneous events from prisonriots to press conferences from kneecapping to general strikes Lumpingthese together makes little sense These two data sets also cover a differentrange of events routine strikes are included in EPCD but excluded fromDCAUS Therefore I will focus on demonstrations including marches ral-lies and vigils In DCAUS demonstrations contribute half the total numberof participants in EPCD for the United Kingdom a third of the total Thesedata sets are usefully compared to a comprehensive catalog of demonstra-tions in Washington DC in 1982 and 1991 compiled primarily from therecords of three police forces (McCarthy et al 1996) Although the authorshave not made the data available published tabulations are valuable forrevealing the mass of small events that never make the news Strikes areparticularly valuable for my purpose because they are less prone to under-reporting and because their size is measured more accurately The UnitedKingdom consistently tabulated the size distribution from 1950 to 19842 TheUnited States recorded every strike from 1881 to 1893 which permits anal-ysis of a subset of events in small cities and towns The final data set isCarterrsquos catalog of black riots in the United States from 1964 to 1971(1986220-21) This allows investigation of variation across cities to com-plement the time series for strikes and demonstrations

This article begins by reviewing the quantification of protest events inrecent literature The second section conceptualizes various ways of measur-ing the size of protest events and outlines the challenge of measuring par-ticipation Size distributions of demonstrations strikes and riots arepresented in the third section These distributions are heavy-tailed meaningthat most participants are concentrated in a few huge events How does thismatter The fourth section compares event frequency and total participationand shows that they are not strongly associated conclusions drawn from onecannot be assumed to apply to the other This negative finding is offset bypositive implications in the fifth section The problem of unreported small

354 Sociological Methods amp Research 47(3)

events diminishes large events provide a remarkably accurate measure oftotal participation The conclusion draws implications for future research

Protest Events in the Literature

Fifty years ago Tilly and Rule (1965) wrote at length on Measuring PoliticalUpheaval Their major precedent was strikes which were quantified in threeways (elegantly explicated by Spielmans 1944) the frequency of events thetotal number of participants workers involved in strikes and the total num-ber of participant-days working-days lost in strikes Tilly and Rule(196574) concluded that lsquolsquo[t]he most useful general conception of the mag-nitude of a political disturbance seems to be the sum of human energyexpended in itrsquorsquo This they argued was best approximated by participant-days In the subsequent half century data on events have accumulated whileconsideration of what to quantify has disappeared A review article from1989 devotes a single page to this issue (Olzak 1989127) and it is notmentioned in a subsequent review (Earl et al 2004)

How then is protest quantified in recent literature Consider articlespublished from 2000 to 2014 in seven leading Anglophone journals Amer-ican Journal of Sociology American Sociological Review British Journal ofSociology European Sociological Review Mobilization Social Forces andSocial Problems (following Amenta et al 2010)3 Articles are selected if theymeasure protest as either the dependent variable or an independent variableor both4 This excludes variables defined by the occurrence of protest ratherthan its magnitudemdashas for example the dates at which cities experiencerioting used in event-history analysis5 Also excluded are studies that takethe event as the unit of observation Predominantly qualitative articles thatdefine the explanans by graphing a time series are included (eg Biggs2013) The literature search yields 41 articles (Online Table S1) Most arti-cles construct a time series at the national level which is usually annual butin a few cases quarterly or monthly Some articles measure how protestvaries across cities or other geographical units A few combine time andspace so the unit of observation is the city-year for example There are singleinstances of other combinations

The quantification of protest requires two choices The first is what sort ofprotest events to cover Some articles focus on distinct types of protest suchas petitions or demonstrations or strikes or riots Other articles aggregatediverse types of action from sit-ins to litigation into a single variable (egJacobs and Kent 2007) Whether it is conceptually meaningful to aggregate

Biggs 355

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 5: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

events diminishes large events provide a remarkably accurate measure oftotal participation The conclusion draws implications for future research

Protest Events in the Literature

Fifty years ago Tilly and Rule (1965) wrote at length on Measuring PoliticalUpheaval Their major precedent was strikes which were quantified in threeways (elegantly explicated by Spielmans 1944) the frequency of events thetotal number of participants workers involved in strikes and the total num-ber of participant-days working-days lost in strikes Tilly and Rule(196574) concluded that lsquolsquo[t]he most useful general conception of the mag-nitude of a political disturbance seems to be the sum of human energyexpended in itrsquorsquo This they argued was best approximated by participant-days In the subsequent half century data on events have accumulated whileconsideration of what to quantify has disappeared A review article from1989 devotes a single page to this issue (Olzak 1989127) and it is notmentioned in a subsequent review (Earl et al 2004)

How then is protest quantified in recent literature Consider articlespublished from 2000 to 2014 in seven leading Anglophone journals Amer-ican Journal of Sociology American Sociological Review British Journal ofSociology European Sociological Review Mobilization Social Forces andSocial Problems (following Amenta et al 2010)3 Articles are selected if theymeasure protest as either the dependent variable or an independent variableor both4 This excludes variables defined by the occurrence of protest ratherthan its magnitudemdashas for example the dates at which cities experiencerioting used in event-history analysis5 Also excluded are studies that takethe event as the unit of observation Predominantly qualitative articles thatdefine the explanans by graphing a time series are included (eg Biggs2013) The literature search yields 41 articles (Online Table S1) Most arti-cles construct a time series at the national level which is usually annual butin a few cases quarterly or monthly Some articles measure how protestvaries across cities or other geographical units A few combine time andspace so the unit of observation is the city-year for example There are singleinstances of other combinations

The quantification of protest requires two choices The first is what sort ofprotest events to cover Some articles focus on distinct types of protest suchas petitions or demonstrations or strikes or riots Other articles aggregatediverse types of action from sit-ins to litigation into a single variable (egJacobs and Kent 2007) Whether it is conceptually meaningful to aggregate

Biggs 355

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 6: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

such heterogeneous forms of action is outside the scope of this article Itdoes however have implications for the second choice

The second choice is how to quantify protest It is usually measured bycounting the frequency of events occurring in each observation This variableis used in 83 percent of articles and 66 percent use this measure alone Fourarticles use the average number of participants per event (eg Soule and Earl2005) Four articles use the total number of participants in protest events forexample workers involved in strikes (eg Checchi and Visser 2005) Twoarticles on riots combine participation duration and severity by using factoranalysis to condense several characteristicsmdashthe number of people arrestedinjured and killed the number of buildings set alight and the duration indaysmdashto a single score (eg Myers 2010)6 Various other variables areconfined to a single article

These two methodological choices are fundamental but most articles donot justify them There is methodological discussion but it concerns relia-bility rather than validity Thus an article will claim that the frequency ofevents reported in the New York Times reliably tracks the actual frequency ofevents while eliding the more fundamental questionsmdashwhat kinds of protestshould be included and how should they be measured The majority ofarticles that measure only frequency do not explain this choice An exceptionis Budrosrsquo (2011444) analysis of petitions from the late eighteenth centurywhich notes the difficulty of obtaining the number of signatories In somecases data on size have presumably not been collected as when the source isthe New York Times Index rather than the original news articles (eg JenkinsJacobs and Agnone 2003) The use of event frequency may follow from aheterogeneous definition of events press conferences and boycotts forexample seem to lack a common size metric Conversely all four articlesthat measure total participation focus on one type of event (strikes or demon-strations or suicide protest)

Conceptualizing and Measuring Size

Size can be conceived in various ways The first dimension of size is thenumber of participants Some forms of protest exhibit little variation Theultimate example is suicide protest where the majority of events comprise asingle participant In a study of 500 events the largest involved 12 individ-uals A group of monks and nuns in Vietnam in 1975 who killed themselvestogether (Biggs 2005a188) By contrast participation varies appreciably formost types of protest such as hunger strikes demonstrations and riots Someevents extend significantly over time strikes can continue for months and

356 Sociological Methods amp Research 47(3)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 7: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

occupations for years Duration is the second dimension of size It creates thenumber of participant-days as a two-dimensional measure In strike statisticsthis is the number of working-days lost The total number of days that everystriker was out on strike A rather different way of conceptualizing size isseverity or disruptiveness The severity of a riot for example can be mea-sured by the number of fatalities or the number of properties destroyedSeverity is specific to the type of protest

Whether measured by participants participant-days or severity eventsare naturally aggregated over time and space to construct quantitative vari-ables such as the total number of demonstrators per month or the totalnumber of riot fatalities per year Note that the total number of participantsis not identical to the total number of individual people who protested in theperiod because some protested more than once If we are interested in con-tentious collective action it is appropriate to count actions A worker whogoes on strike twice in the course of a year properly contributes two actionsto the annual total One point to emphasize is that aggregation entails takingthe sum and not computing the average The literature sometimes takes theaverage number of protesters per event as a measure of protest but this doesnot serve to quantify participation (The same objection applies to averageduration) A simple numerical example clarifies this point One city has twodemonstrations with 100 and 100000 participants Another city has threedemonstrations two of 100 participants and one of 100000 Measuringaverage participation implies that there is 50 percent more protest in theformer city than in the latter In fact of course there is marginally moreprotest in the latter city If protesters wanted to maximize average rather thantotal participation then they would never hold small events

Aggregate measures are often used to trace change over a long periodwhen population grew appreciably or to compare across spatial units varyingin population Total participants and total participant-days are naturallydivided by the population of potential protesters to create proportional mea-sures of propensity and intensity respectively Propensity neatly correspondsto the individualrsquos continuous-time hazard rate of protesting It is commonusage to also use population as the denominator for event frequency (egRosenfeld 2006) or equivalently to specify this as the exposure term inPoisson or negative binomial regression (eg Inclan 2009) Events per per-son however is less easily interpreted More seriously the ratio implies thatevent frequency increases in linear proportion to population which is math-ematically implausible For a given propensity to protest double the popu-lation means double the total number of participants (and hence double thetotal number of participant-days) But we would not expect twice as many

Biggs 357

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 8: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

events because that would imply no change in the average number of parti-cipants per event Presumably event frequency and average size would eachincrease by a factor of between 1 and 2 with their product equal to 2 In otherwords event frequency should be denominated by population raised to thepower of 05 (the square root of population) or similar fraction

This article focuses on the first and most basic dimension of size thenumber of participants The two-dimensional measure of size participant-days is pertinent for strikes but is not relevant for demonstrations and is notempirically measurable for riots Therefore it is not considered further hereAlso omitted are measures of severity Ultimately it will be necessary tocompare different measures of size but there is no space to undertake thishere This article also ignores the challenge noted above of aggregatingdisparate kinds of protest Conceptually this would seem to require measur-ing the cost (or impact) of different types of protest actions arraying them ona spectrum from signing an e-mail petition to setting oneself on fire Thisproblem will be avoided here by analyzing different types of protestseparately

Measuring the number of participants is often challenging At the reliableend of the spectrum are strikes Data are compiled by specialized officialswho are not involved in the dispute They can use the records of firms and oftrade unions if they provide strike pay Because the point of a strike is toinflict costs on the employer neither side has reason to exaggerate orminimize the number of strikers at least not in statistics published longafter the event is over At the opposite extreme of reliability are riots Asidefrom the question of how exactly to define what counts as participation thenature of a riotmdashfluid dispersed and furtivemdashhinders numerical estima-tion The number of arrests provides a rough proxy for participation thoughthis depends not just on the number of rioters but also the tactics andcapabilities of the police

Estimating the number of demonstrators falls somewhere between riotsand strikes A proper estimate can be calculated from the dimensions of thegathering place and the physical density of demonstrators (McPhail andMcCarthy 2004)7 Such calculations began to be produced by the US ParkPolice in Washington DC in the mid-1970s More usually however sizehas to be taken from the guesstimates of police or reportersmdashwho are usuallyunsympatheticmdashor of organizersmdashwho naturally exaggerate Media sourcescan select which of these estimates to report in accordance with their sym-pathy or antipathy to the protestersrsquo cause (Mann 1974)8 A notorious exam-ple of conflicting estimates was the lsquolsquoMillion Man Marchrsquorsquo in 1995 Theorganizers naturally claimed that it lived up to its name while the Park Police

358 Sociological Methods amp Research 47(3)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 9: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

calculated 400000 The ensuing controversy led Congress to forbid Federalpolice forces from estimating crowd size

Conflicting estimates of size may seem to pose an insuperable problem Itis attenuated however if we conceive size multiplicatively rather than addi-tively on a logarithmic rather than linear scale This point was originallymade by Richardson (1948523) for the size of wars measured by fatalitiesthese too are estimated with considerable error (and also have a heavy-taileddistribution to anticipate the next section) What really matters is the order ofmagnitude the difference between one power of ten and the next To returnto the example of the Million Man March the difference in estimates isconsiderable The Park Policersquos figure is 60 percent less Taking the loga-rithm to the base 10 the competing estimates are 56 and 60 on a scalebeginning at 0 (for a single demonstrator) Thus transformed the differenceshrinks To put this another way even the lower estimate makes the MillionMan March one of the very largest demonstrations in the United StatesThinking of size in terms of orders of magnitude is intuitive for commenta-tors often describe size in this waymdashreferring to lsquolsquohundredsrsquorsquo or lsquolsquothousandsrsquorsquoof participants for example

Size Distributions

We can now investigate the number of participants in demonstrationsstrikes and riots Table 1 presents summary statistics for five data sets Acrude way of gauging variation is to compare the maximum to the medianAnother is to divide the standard deviation by the mean which yields thecoefficient of variation Following mean and variance the third and fourthmoments of the distribution are skewness and kurtosis The latter is signif-icant because it indicates the heaviness of the tail of the distribution Table 1presents L-kurtosis one of the L-moments that are derived from order sta-tistics (Hosking 1990) Unlike the conventional measure of kurtosis thisdoes not increase with sample size and it is less sensitive to extreme valuesL-kurtosis can range from25 to 1 for a normal distribution it is 12 and foran exponential distribution it is 17 (The Gini index will be discussed later)A heavy tail may be defined as one that is heavier than the exponentialdistribution meaning that there is a nontrivial probability of extremely largevalues Heavy tails are less familiar to our statistical intuition nurtured on thenormal distributionmdashthe name reveals its hegemonymdashand developed byexperience with thin-tailed distributions such as age and years of education9

For demonstrations a single estimate of size is reported in the majority ofevents in the United States and the United Kingdom (63 percent in DCAUS

Biggs 359

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 10: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Tab

le1

Size

Dis

trib

utio

nsofPr

ote

stEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

Dem

ons

trat

ions

inD

C

Stri

kes

inth

eU

nite

dK

ingd

om

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

es

Arr

ests

Rio

tsin

the

Uni

ted

Stat

esA

rres

ts

Nonw

hite

s

Num

ber

of

even

ts7

878

104

73

065

783

7859

355

5

Min

imum

11

2lt

101

000

01pe

rcen

tM

edia

n20

130

023

7017

010

33pe

rcen

tM

ean

282

46

693

929

628

116

032

20pe

rcen

tM

axim

um60

000

030

000

050

000

01

750

000

777

24

0307

perc

ent

Stan

dard

devi

atio

n19

061

256

9510

931

169

9656

50

5266

perc

ent

Coef

ficie

ntofva

riat

ion

67

38

118

270

49

16

L-sk

ew8

88

28

78

25

4L-

kurt

osi

s7

86

77

96

92

8G

inii

ndex

91

88

94

87

84

69

360

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 11: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

and 71 percent in EPCD) The remainder are described approximately as byorders of magnitude EPCD employs the sensible convention of coding hun-dreds as 300 thousands as 3000 and so on10 I apply this to DCAUS11 In bothcountries the typical or median size was in the hundreds while the maximumwas in the hundreds of thousands These distributions underestimate the degreeof variation because the smallest events were much less likely to be reportedIt is valuable to compare the exceptionally comprehensive record of demon-strations in Washington DC (McCarthy et al 1996484 488) Here the num-ber of participants is usually the number anticipated by organizers whenapplying for a permit to demonstrate sometimes it is the number as subse-quently amended by the police The median size was two dozen The largestdemonstration exceeded the median by over four orders of magnitude

Data on strikes in the United Kingdom cover those lasting for at least aday and involving at least 10 workers along with small or brief strikes if atleast 100 working days were lost Size is measured by the number of workersinvolved12 Published tabulations use broad size intervals but fortunately itis possible to identify each individual strike reaching 10000 workers (seeOnline Appendix S1) The median was under 100 which is comparable todemonstrations when they are measured comprehensively The maximumwas larger because a strike does not require the physical assembly of parti-cipants in a single place

For riots size is proxied by the number of arrests (Events where no onewas arrested are omitted) The median number of arrests was 17 The max-imummdasha riot in Washington DC in April 1968mdashwas greater by over twoorders of magnitude From retrospective surveys in three cities it is esti-mated that between one in six and one in three rioters were arrested (Fogel-son 197136-37) By implication participation in the largest riots was anorder of magnitude smaller than participation in the largest demonstrationsThis makes sense given that a major demonstration in a capital city alwaysattracts people from elsewhere whereas the vast majority of rioters are localBecause a riot is geographically circumscribed we can also consider thenumber of arrests in relation to the population of potential rioters in thiscase the number of nonwhites in the city13 The median was 01 percent Themaximum was greater by an order of magnitude at 4 percent Dividing bypopulation significantly reduces variance (as manifested in the coefficient ofvariation) and kurtosis Note that the data on riots comprise fewer events thanthe other data sets In principle of course the maximum will increase withthe number of observations and markedly so for a heavy-tailed distributionIn practice though it is hard to conceive of a much greater maximumnumber of arrests in one city

Biggs 361

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 12: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

All these distributions exhibit enormous variation in size The standarddeviation always exceeded the mean The least variation occurred in thedistribution of riot arrests relative to population even then the maximumwas 40 times the median The ratio of maximum to median exceeded 20000for participants in demonstrationsmdashwhen not filtered through news mediamdashand in strikes The typical size is profoundly misleading

A heavy tail implies lsquolsquomass-count disparityrsquorsquo (Crovella 2001) Most of thetotal size comes from a small number of huge events The literature onprotest occasionally notes this feature (Franzosi 1989352 Koopmans1995251 Rucht and Neidhardt 199876-77) but its importance has not beenfully appreciated Figure 1 illustrates the disparity by comparing two cumu-lative distributions (after Feitelson 2006) One is the distribution of events bysize The other shifted to the right is the distribution of participants by sizeof event The horizontal scale must be logarithmic of course to encompassthe extreme range Discontinuities in the graphs for demonstrations in theUnited States and the United Kingdom are due to size ranges (like thousands)being translated into a number

Comparison of the three graphs for demonstrations reveals the gap left byunderreporting In DC demonstrations with no more than 25 participantsaccounted for over half the events When events were filtered through thenews media the total contained far fewer of these tiny events 5 percent inthe United Kingdom and 7 percent in the United States Paradoxically how-ever the distribution of participation shows that when small events werecomprehensively recordedmdashdemonstrations in DC and strikesmdashthey stillcontributed almost nothing to the total Total participation was dominatedby the largest events which constituted a tiny fraction of all events In DConly 1 percent of demonstrations had more than 10000 participants butthey accounted for 69 percent of the total demonstrators For strikes only04 percent involved as many as 10000 workers but they accounted for56 percent of the total workers involved In only 18 percent of riots didarrests reach 1000 but they accounted for just over half of the total arrestsFor riots as before we can adjust for population One in ten riots led to thearrest of at least 1 percent of the nonwhite population and these riotsaccounted for half of the total percentage of nonwhites arrested

Mass-count disparity is visualized as the gap between the two cumulativedistributions It can be captured in a familiar statistic the Gini index ofinequality shown in Table 114 For demonstrations the index ranged from88 to 94 it was highest for DC because more small events were included15

For strikes the index was 87 These figures indicate extreme inequality Forcomparison the distribution of wealth in the contemporary United States has

362 Sociological Methods amp Research 47(3)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 13: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Even

ts

Part

icip

ants

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

USA

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

UK

01

110

010

000

Dem

onst

rato

rs

Dem

onst

ratio

ns in

DC

01

100

100

001

000

000

Wor

kers

invo

lved

Stri

kes

in U

K

01

110

0 Arr

ests

Riot

s in

USA

01

110

010

000

Arr

ests

m

illio

n no

nwhi

tes

Riot

s in

USA

Proportion size

Fig

ure

1

Cum

ulat

ive

size

dist

ribu

tions

mas

s-co

unt

disp

arity

363

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 14: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

a Gini index of about 8 The index for arrests was 84 When arrests aredenominated by population the index was 69 The latter is still higher thanthe index for the distribution of income (before taxes and transfers) in thecontemporary United States about 5

Thus far I have referred generically to heavy-tailed distributions withoutconsidering the shape of the tail The archetypal heavy tail is the power lawwhere the probability of an event of size x is proportional to xa In apioneering study Richardson (1948) argued that the size of wars measuredby fatalities followed a power law The same distribution has been identifiedin terrorist and insurgent attacks (Bohorquez et al 2009 Clauset Young andGleditsch 2007) According to Biggs (2005b) strikes in two cities in the latenineteenth century followed a power law with a estimated as 19 to 20 (forx 100ndash150) Discriminating a power law from other heavy-tailed distribu-tionsmdashsuch as the lognormal and the power law with exponential cutoffmdashisempirically demanding Only a small fraction of events comprises the tailand so the total number of observations must be very large (Clauset Shaliziand Newman 2009) Therefore it makes sense to concentrate on strikesFigure 2 plots the complementary cumulative distribution with both axeson a logarithmic scale (Online Figure S1 depicts the other data sets)

Exponential

Power

10

10

001

Prop

ortio

n s

ize

100 10000 1000000Workers involved

Figure 2 Strikes in the United Kingdom complementary cumulative sizedistribution

364 Sociological Methods amp Research 47(3)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 15: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Following the method of Virkar and Clauset (2014) the power law with thebest fit has a of 22 starting at 1000 to 2499 workers It appears on the graphas a straight diagonal line Clearly this power law does not describe the veryupper tail of the distribution it predicts too few huge strikes But alternativeheavy-tailed distributions (fit to the same tail starting at 1000 to 2499) areinferior16 The graph also serves to illustrate the significance of a heavy tailRecall that heavy means heavier than the exponential distribution The graphshows the best-fitting exponential distribution which resembles an invertedJ Such a distribution would predict many more medium-sized strikes andfewer large ones a strike involving more than a hundred thousand workerswould be vanishingly rare A heavy tail by contrast reflects the fact thathuge eventsmdashmore than four orders of magnitude greater than the median inthis casemdashcan occur

Such huge events occur in data sets that cover a population of manymillions over decades these data sets are the staple of sociological analysisDoes the size distribution differ for small populations Answering this ques-tion requires comprehensive data on small events in minor places whichrules out the media as a source The best candidate is strikes in the UnitedStates between 1881 and 1893 (see Online Appendix S2) At that timealmost all employers were confined to a single location with the majorexception of railroads Table 2 shows the size distribution for strikes inIllinois outside the industrial metropolis of Chicago (Cook County to beprecise) A few large railroad strikes are omitted because the Commissioner

Table 2 Size Distributions of Protest Events

Strikes in Illinois Outside ChicagoWorkers Involved

Strikes in PeoriaWorkers Involved

Number ofevents

551 29

Minimum 2 4Median 80 40Mean 236 95Maximum 8929 600Standard

deviation647 145

Coefficient ofvariation

27 15

L-skew 66 57L-kurtosis 49 33Gini index 71 66

Biggs 365

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 16: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

did not specify all their locations the period ends before the massive Pullmanrailroad strike in 1894 The size distribution exhibits much less variation thanthe large national data sets The coefficient of variation is 27 compared to27 for strikes in the United Kingdom Mass-count disparity is still significantOnly 5 percent of strikes involved at least 1000 workers but these accountedfor almost half (44 percent) the total number of participants Although thelargest strike fell short of 10000 workers note that the number of events isrelatively small Sampling the same number of events from strikes in theUnited Kingdom we would expect only two (04 percent) to reach 10000Focusing on a particular location further reduces variation in the size ofevents Table 2 shows strikes in Peoria the statersquos second city which hada working-class population of about 8000 The largest strike involved 600workers With so few events the upper tail of the distribution cannot beestimated a much larger strike could have been possible Even so the Giniindex shows greater inequality than is found in the distribution of income

Event Frequency and Total Participation

Sociologists usually choose to count the frequency of events in each timeinterval or geographical unit An alternative is to sum the number of parti-cipants to yield total participation per interval or unit This takes into accountthe enormous variation in the size of events like demonstrations and strikesIt also avoids a problem that has escaped attention in the literature17 How isone event to be demarcated from another In principle this should bestraightforward when dealing with a contentious gathering that is character-ized by continuity and contiguity of action like a march In practice how-ever event catalogs often treat multiple gatherings in different locations asone event Thus vigils and processions in 21 cities on lsquolsquoNational Free SharonKowalski Dayrsquorsquo (to support a disabled woman whose lesbian partner wasdenied access by her family) become a single event in DCAUS Why notcount 21 events EPCD likewise classifies the annual Loyalist Orange par-ades throughout Northern Ireland on July 12 as a single event (sometimesentering a particularly large or contentious parade as a second event) Whynot consider the parades in each town or city as separate events Thesecoding decisions apparently reflect the level of detail provided by newsreports The problem of demarcation vanishes if we measure total participa-tion in a time interval When a newspaper reported lsquolsquoscores of Orangedemonstrations in which an estimated 100000 people took partrsquorsquo (The Timesof London July 13 1990) whether we treat this as one event or scores makesno differencemdasheither way total participation increases by 100000

366 Sociological Methods amp Research 47(3)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 17: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Does the choice between total participation and event frequency makea difference in practice One might expect a high correlation especiallyin an annual national time series After all there are many protest eventsin each year on average for example over 200 demonstrations in theUnited States and over 2000 strikes in the United Kingdom When somany events are aggregated we might suspect that size differenceswould tend to average out

Figure 3 traces both time series for demonstrations in the United StatesTotal participation spikes in 1969 and peaks in 1982 Event frequency peaksin the mid-1960s declines to the mid-1970s and then continues at a lowlevel The two variables follow a completely different trajectory The scat-terplots in Figure 4 portray the association between event frequency and totalparticipation the linear regression line is dashed For demonstrations in theUnited States the correlation coefficient is only 10 The correlation issomewhat higher for demonstrations in the United Kingdom albeit for only16 years An estimated correlation coefficient automatically increases withfewer observations in the United States the correlation for each 16-yearsubperiod (1960ndash1975 to 1980ndash1995) averages 19 For strikes in the UnitedKingdom the correlation is low For strikes in Illinois outside Chicago

Tota

l dem

onst

rato

rsD

emon

stra

tion

freq

uenc

y

1960 1970 1980 1990

r = 10

Figure 3 Demonstrations in the United States total participation and eventfrequency

Biggs 367

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 18: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 1

0

Dem

onst

ratio

ns in

USA

by

year

Total demonstrators

Dem

onst

ratio

n fr

eque

ncy

r = 2

7

Dem

onst

ratio

ns in

UK

by

year

Total workers involved

Stri

ke fr

eque

ncy

r = 2

3

Stri

kes

in U

K b

y ye

ar

Total arrests

Riot

freq

uenc

y

r = 5

1

Riot

s in

USA

by

city

Total arrests nonwhites

Riot

freq

uenc

y

nonw

hite

s

r = 6

4

Riot

s in

USA

by

city

Fig

ure

4

Even

tfr

eque

ncy

and

tota

lpar

tici

pation

368

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 19: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

(not graphed) the correlation is 43 again for a much shorter period Forthese time series the population of potential protesters did not increasesufficiently to require adjustment18

Riots are mapped on to the 460 urban places with a population of at least25000 and a nonwhite population of at least 1000 in 1960 Few events aredistributed across many observations whereas the time series have manyevents distributed across few observations The maximum number of riotswas 14 in Washington DC Half of the cities experienced no riots and so theregression line (in Figure 4) is anchored at the bottom left-hand cornerThe correlation between total arrests and riot frequency is only modestBut it is really necessary to adjust for the tremendous variation in thesize of cities Total participation obviously scales with the population atrisk As argued above event frequency must scale to population raisedto a fractional power the square root is used here19 Thus denominatedriot frequency and total arrests reach the highest correlation 64 Thismeans that frequency predicts only 41 percent of the variation in percapita arrests20

In sum when events like riots and demonstrations are aggregatedover time or space the frequency of events is only minimally or mod-estly associated with total participation in those events The divergencebetween event frequency and total participation partly reflects theheavy-tailed size distribution of events the occurrence of a huge eventsignificantly increases total participation while only incrementingevent frequency Furthermore the four annual time series reveal neg-ative correlations between event frequency and average event size inyears with many events events tended to be smaller Whether this iscoincidental or reflects a more general pattern must await research onother data sets

These findings do not prove that total participation always diverges fromevent frequency For events with minimal variation in size (exemplified bysuicide protest) total participation will be practically the same as eventfrequency For the staple tactics of social movements however there is nojustification to assume a high correlation over time or across spatial unitsSimilar disassociation is evident for example for demonstrations inBelarus in the 1990s aggregated by quarter (Titarenko et al 2001137)It follows that the findings from multivariate analysis using event fre-quencymdashas either the dependent variable or an independent variablemdashcannot be assumed to apply to total participation Likewise findings usingtotal participation cannot be assumed to apply to event frequency Thechoice of measurement is crucial

Biggs 369

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 20: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

The Dominance of Large Events

The fact that protest events vary enormously in size can mitigate a majorproblem identified in the literature Most events are not reported by the newsmedia and so are omitted from event catalogs In Washington DC newspa-pers reported only one in ten demonstrations Worse the extent of this under-reporting will fluctuate over timemdashdepending for example on thesignificance of other newsmdashin ways that cannot be ascertained Hence theargument that newspapers are a seriously flawed source of data (Myers andCaniglia 2004 Ortiz et al 2005) They are indeed flawed if one wishes tocount the frequency of events But the problem evaporates if one wants totrace total participation over time After all the probability of an event beingreported increases with the number of participants Newspapers reportedonly 3 percent of the demonstrations involving 2 to 25 participants in DCbut they reported 52 percent of the demonstrations with 10001 to 100000and both the demonstrations exceeding 100000 (McCarthy et al 1996488)

Moreover mass-count disparity has a surprising implication Whenaggregated over time or space total participation can be predicted by lookingonly at the large events What counts as lsquolsquolargersquorsquo depends on the actual sizedistribution of course We can use the threshold of 10000 for strikers anddemonstrators in the national data sets and 1000 for strikers in Illinois andfor arrests Table 3 reports the correlation coefficients Aggregated by yearthe correlations are almost perfect The result for strikes is especially com-pelling because these data do not omit small events Figure 5 details the totalnumber of workers involved in UK strikes reaching various size thresholdsRemarkably to trace how the total number of strikers fluctuated from year toyear it is possible to ignore 996 percent of events just track the total numberof workers involved in strikes of at least 10000 (This understates the abso-lute level of participation of course but it accurately captures relativechange) Increasing the threshold to the next order of magnitude hardly altersthe graph Even confining attention to strikes involving at least a millionworkersmdashonly nine eventsmdashcaptures the salient peaks

Given the dominance of large events it may be tempting to count theirfrequency rather than to sum their participants McAdam and Su (2002) forexample count the frequency of events exceeding 10000 participants Aheavy tail however implies that large events also manifest mass-count dis-parity most are moderately large while a few are massive Table 3 comparesthe correlation between total participation and the frequency of large eventsThe correlation is consistently lower and it is greatly inferior for UKdemonstrations

370 Sociological Methods amp Research 47(3)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 21: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Tab

le3

Tota

lPar

tici

pation

and

Larg

eEv

ents

Dem

ons

trat

ions

inth

eU

nite

dSt

ates

byY

ear

Dem

ons

trat

ions

inth

eU

nite

dK

ingd

om

byY

ear

Stri

kes

inth

eU

nite

dK

ingd

om

byY

ear

Work

ers

Invo

lved

Stri

kes

inIll

inois

Out

side

Chi

cago

byY

ear

Work

ers

Invo

lved

Rio

tsin

the

Uni

ted

Stat

esby

City

Arr

ests

Thr

esho

ldde

finin

gla

rge

even

ts10

000

100

0010

000

100

01

000

Larg

eev

ents

al

leve

nts

41

perc

ent

113

perc

ent

039

perc

ent

47

perc

ent

18

perc

ent

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

part

icip

atio

nin

larg

eev

ents

98

99

98

92

97

Corr

elat

ion

betw

een

tota

lpa

rtic

ipat

ion

and

freq

uenc

yofla

rge

even

ts

76

45

68

84

87

371

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 22: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

The problem of underreporting which has preoccupied the literature onprotest events shrinks when mass-count disparity is understood For analysisof protest over time at the national level the problem is readily overcome ifwe measure total participation rather than event frequency Total participa-tion is dominated by large events and large events are most likely to bereported This reassurance does not hold for the cross-sectional analysis ofprotesters denominated by population because a national newspaper may notbother reporting a relatively large event in a small city

Mass-count disparity does however highlight another issue Becauselarge events are so important errors affecting them have severe repercus-sions Both DCAUS and EPCD erroneously duplicate some demonstrationswith hundreds of thousands of participants One duplicate for exampleincreases the yearrsquos total participants by 55 percent Coding errors are inev-itable in assembling a large data set from news reports Happily mass-countdisparity means that validation can focus on large events which constituteonly a tiny fraction of all events In DCAUS for example only 4 percent ofdemonstrations reached 10000 participants Aside from errors uncertaintyin estimating the size of large events is also a significant concern Whenestimates are wildly discrepantmdashif for example the higher is more thandouble the lowermdashthe decision on which size to use will make a difference

all

N =

78

378

10

000

N =

304 r = 98

100

000

N =

43 r = 96

100

000

0N

= 9

1950 1960 1970 1980

r = 87

Figure 5 Strikes in the United Kingdom Workers involved in events of varying size

372 Sociological Methods amp Research 47(3)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 23: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

When collecting data therefore it will be crucial to record varying estimatesand their provenance Estimates derived from spatial calculation (area occu-pied and density of demonstrators) are obviously preferable Even withoutthe benefit of such calculation one could consistently select estimates withthe same direction of bias for example using the optimistic figures providedby the organizers

Conclusion

The study of social movements requires the quantification of protest Whatexplains protest What does protest explain Whether we treat protest as aneffect and seek its causes or treat protest as a cause and seek its effects weneed to differentiate less protest from more This basic distinction betweenless and more is required even for explanations that are pursued with quali-tative evidence rather than statistical analysis an historical narrative typi-cally graphs the time series of protest The pioneers of protest event analysisdevoted serious thought to what should be measured and how (eg Tilly andRule 1965) As data accumulated thanks to their efforts these crucial issuesfaded from view21

The choice of how to quantify protest is fundamental and this has notbeen fully appreciated Aggregated over time intervals or across geographi-cal units there is no high correlation between event frequency and totalparticipation Four time series yield correlation coefficients from 10 to43 with city as the unit of observation the coefficient does not exceed64 Perhaps other data sets will reveal higher correlations but this will needto be demonstrated As it stands the frequency of events and the total numberof participants diverge so much that findings for one are unlikely to apply tothe other

Event frequency has become the default variable in studies of protestThere are numerous points in its favor It does not require the measurementof the size of events which is always subject to uncertainty (and may even beimpossible when using fragmentary historical records) It enables the aggre-gation of diverse phenomena from hunger strikes to press conferences into asingle metric Using the most commonly used variable has the advantage offacilitating comparison with previous work Nevertheless event frequencyrequires the assumption that size does not matter To put this tangibly itassumes that two events represent twice as much as protest as one event evenif the two events are each attended by only 10 participants whereas the singleevent attracts a million people The findings of most quantitative work onprotest rest on this assumption Whether this assumption is justified depends

Biggs 373

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 24: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

ultimately on theory Event frequency is obviously appropriate for testingtheories that explain variation in the number of events and theories in whichthe effect of protest depends on the number of events irrespective of sizeWhen events are counted from news media underreporting will continue topose a fundamental methodological challenge Event frequency as mea-sured represents a small proportion of all events and this proportion fluc-tuates over time and across space Two other challenges must also be solvedOne is to explicate consistent criteria to demarcate one event from anotherAnother is to choose an appropriate fractional power when adjusting forpopulation size

Total participation has several countervailing advantages Measuring thenumber of actionsmdashthus when divided by population the individualrsquoshazard of participatingmdashfollows naturally from the conception of protestas collective action Theories often postulate that the effect of protestincreases with the total number of participants In practice movements actas if they are trying to mobilize more people Alongside these theoreticalconsiderations empirical findings highlight the significance of size Mea-sured by participants the size distribution of eventsmdashlike demonstrationsstrikes and riotsmdashdoes not range modestly around a typical value Suchprotest events have a heavy tail most events are small but a few aremassive Even when we confine attention to events in one small city (likestrikes in Peoria) or divide participants by population (as with riots) thereis pronounced inequality in size How far does this generalization hold Itdoes not apply to every form of protest an exception is suicide protest Itmay not apply in some social contexts22 But the generalization does holdfor the protest events that are the staple of sociological analysis23

Extreme variation in size has positive implications for measurement Witha heavy-tailed distribution the weight of the distribution is concentrated atthe top Most protesters participate in large events The underreporting ofsmall events matters less for participation than for event frequencybecause small events contribute only a small fraction of the total numberof participants The national series of demonstrations and strikes ana-lyzed here are dominated by events involving at least 10000 participantsIt is important to emphasize that large is relative rather than absoluteand so it varies with the population of potential protesters In Peoria inthe 1880s for example a large strike was one that involved hundreds ofworkers When aggregated over time or space the total number of parti-cipants is very highly correlated with the total number in the largestevents By implication attention should focus on the accurate recordingof these events One issue is the estimation of size Varying estimates of

374 Sociological Methods amp Research 47(3)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 25: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

the same event establish lower and upper bounds which can be used tocheck the sensitivity of empirical findings For contemporary events thehuge volume of photographic and video evidence can surely be exploitedto refine estimates of crowd size24 Another issue is the validation ofdata The duplication of large events in data setsmdashidentified here in bothDCAUS and EPCDmdashcan severely distort empirical results Fortunatelyof course there are only a small number of large events and so it isfeasible to check them thoroughly

My argument has two further implications for future research I havefocused on one measure of size the number of participants Arguably thisvariable is most appropriate when explaining the origins of protest as itmeasures the number of individual decisions to participate When usingprotest as an explanation for subsequent outcomes however participant-daysmdashincorporating the second dimension durationmdashcould be preferableThere are alternative measures of size such as severity (eg Carter 1986)which also deserve further examination The point of my argument is not toimpose one single variable for all theoretical purposes but to encourage theinvestigation of size

The size distribution of protest events has implications beyond method Asignificant task for theory is to explain why most protest events are smallwhile a few are huge One simple explanation is that this distribution reflectsthe division between populations of potential protesters The population ofcities for example follows a power law with a of 2 (Rozenfeld et al 2011)Thus the distribution of riot arrests is partly a function of the distribution ofpopulation the heavy tail is significantly diminished when arrests areexpressed per capita Evaluating this explanation for other types of protestis more complicated because it is hard to conceive of preexisting populationdivisions for strikesmdashthough the distribution of workers by industry or occu-pation could be consideredmdashor for demonstrations A second explanation isthat the size distribution reveals something about the process generatingprotest events Biggs (2005b) argues for positive feedback for an event inprogress the larger it becomes the more likely people are to join it Thiskind of processmdashsynonymous with cumulative advantage or preferentialattachmentmdashis often used to explain heavy-tailed size distributions (egSeguin 2016) Other models of the generative process are also possible Forviolent attacks a model of the fusion and fission of insurgent groups pre-dicts the distribution of fatalities (Clauset and Weigel 2010 Johnson et al2006) Modeling the size of protest events is a promising avenue for futureresearch

Biggs 375

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 26: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Authorrsquos Note

This article is possible because other scholars have generously shared their dataGregg Carter Ronald Francisco Doug McAdam John McCarthy Susan Olzak andSarah Soule Helen Margetts and Scott Hale It uses software written by Nicholas JCox Zurab Sajaia and Yogesh Virkar Kenneth T Andrews and Charles Seguinsubjected a draft to judicious critique while Tak Wing Chan and Neil Ketchleysupplied refreshment and encouragement Andrew Gelmanrsquos blog convinced me thatmeasurement is worth taking seriously The unexpurgated version is available asOxford Sociology Working Paper No 2015-05 (httpwwwsociologyoxacukworking-paperssize-matters-the-problems-with-counting-protest-eventshtml)

Declaration of Conflicting Interests

The author(s) declared no potential conflicts of interest with respect to the researchauthorship andor publication of this article

Funding

The author(s) received no financial support for the research authorship andor pub-lication of this article

Notes

1 The article consistently refers to the frequency of events and the number of

participants simply to help the reader discriminate between them

2 In this period the size classification pertains to strikes begun each year from

1985 it switches to strikes in progress each year In the latter scheme some

strikes spanned two calendar years and thus were counted twice

3 Space limitations preclude consideration of a parallel literature in political sci-

ence (eg Przeworski 2009) which relies on the frequency of strikes riots and

demonstrations in the Banksrsquo Cross-National Time-Series Data Archive (origi-

nating with Rummel 1963)

4 Articles are identified by searching the title and abstract for one of the following

key words protest(s) demonstrations strikes or riots

5 If event-history analysis also measures the magnitude of protest as an indepen-

dent variable then it qualifies for inclusion

6 Factor analysis has a long pedigree for riots (Carter 1986 Morgan and Clark

1973) but has not spread to other sorts of protest events

7 By the same principle participants in a march can be estimated by fixing one

observation point and sampling the number passing per minute adding a second

point enables the estimation of error bounds (Yip et al 2010)

8 McCarthy et al (1998123) compare the number of participants reported by

newspapers with the number anticipated by the organizers when applying for a

376 Sociological Methods amp Research 47(3)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 27: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

permit and find a correlation of 71 Ideally we would discover the correlation

with the actual number of participants

9 Empirical variables encountered in introductory and intermediate statistical

courses are invariably thin-tailed Personal income would be a potential exception

but survey data given to students (like the US General Social Survey) truncate its

heavy tail Resnick (2010) provides a useful primer on heavy-tailed distributions

10 Koopmans (1995261) codes lsquolsquohundredsrsquorsquo as 500 lsquolsquothousandsrsquorsquo as 5000 and so

on This is too generous

11 In the smaller categories I take the midpoint 5 for 2 to 9 30 for 10 to 49 75 for

50 to 99 I exclude 15 demonstrations in the US data set that lack information on

the number of participants

12 There is a distinction between workers directly involved and indirectly involved

The latter are involuntarily thrown out of employment by a strike in their firm If

the explanandum is protest then we ideally should consider only workers directly

involved for they have chosen to participate The tabulation of UK size dis-

tributions however does not separate workers indirectly involved These con-

tributed only a tiny fraction of the total 015 percent

13 Cities are defined as urban places in the 1960 Census (Haines and ICPSR 2010

DS63 US Bureau of the Census 1960table 21) It is not possible to match 38

riots (64 percent) most commonly because the place was too small for the

Census to tabulate population by race

14 The Gini index is also the L-moment equivalent of the coefficient of variation

mean divided by L-scale

15 The index for DC is calculated from binned data which requires interpolation and

so is less reliable than the others

16 The hypothesis that strikes follow a power law is rejected with p frac14 003 But the

power law is superior (p lt 001) to the power law with exponential cutoff the

lognormal the stretched exponential and the exponential distribution

17 An exception is Shorter and Tilly (1974353)

18 The US population grew by 45 percent the UK employed labor force by

1 percent and the Illinois nonagricultural occupied population (Cook County

cannot be separated) by 75 percent

19 Recall that the number of events and the average number of participants per event

must each scale to a fractional power of population with the two fractions adding

to onemdashin order for the total number of participants to scale with population

There is no reason to suppose that either event frequency or average participation

increases faster than the other The most parsimonious assumption is that each

scales with population raised to the power of 5

20 Some analyses use the logarithm of event frequency (eg Jenkins et al 2003) Its

correlation with the logarithm of total participation is barely higher The

Biggs 377

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 28: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

correlation coefficients are 23 for US demonstrations 32 for UK demonstra-

tions 29 for UK strikes and 47 for Illinois strikes The same calculation

cannot be performed for riots because half the cities experienced no riot and zero

has no logarithm

21 The related enterprise of quantifying protest through population surveys followed

a similar trajectory The pioneers carefully formulated survey questions for par-

ticular theoretical purposes the proliferation of survey data subsequently came to

define the object of investigation (Biggs 2015)

22 One candidate could be demonstrations under authoritarian regimes which only

a few dissidents are willing to risk But such regimes also generate on rare

occasions truly massive gatherings as witnessed in East Germany in 1989 and

Egypt in 2011

23 Another example is provided by petitions hosted by official websites in the

United Kingdom and the United States (Margetts et al 2016chapter 3) Out of

32873 petitions submitted to the UK Parliament from 2011 to 2015 only 077

percent reached 10000 signatures but these accounted for 70 percent of the total

signatures (calculated from data kindly supplied by the authors)

24 This will be facilitated by imagery from aerial unmanned vehicles (Choi-

Fitzpatrick Juskauskas and Sabur 2015)

Supplemental Material

Supplementary material for this article is available online

References

Amenta Edwin Neal Caren Elizabeth Chiarello and Yang Su 2010 lsquolsquoThe Polit-

ical Consequences of Social Movementsrsquorsquo Annual Review of Sociology 36

287-307

Barranco Jose and Dominique Wisler 1999 lsquolsquoValidity and Systematicity of News-

paper Data in Event Analysisrsquorsquo European Sociological Review 15301-22

Biggs Michael 2005a lsquolsquoDying without Killing Self-immolations 1963-2002rsquorsquo

Pp 173-208 in Making Sense of Suicide Missions edited by Diego Gambetta

Oxford UK Oxford University Press

Biggs Michael 2005b lsquolsquoStrikes as Forest Fires Chicago and Paris in the Late 19th

Centuryrsquorsquo American Journal of Sociology 1101684-14

Biggs Michael 2013 lsquolsquoHow Repertoires Evolve The Diffusion of Suicide Protest in

the Twentieth Centuryrsquorsquo Mobilization 18407-28

Biggs Michael 2015 lsquolsquoHas Protest Increased since the 1970s How a Survey Ques-

tion Can Construct a Spurious Trendrsquorsquo British Journal of Sociology 66141-62

378 Sociological Methods amp Research 47(3)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 29: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Bohorquez Juan Camilo Sean Gourley Alexander R Dixon Michael Spagat and

Neil F Johnson 2009 lsquolsquoCommon Ecology Quantifies Human Insurgencyrsquorsquo

Nature 461911-14

Budros Art 2011 lsquolsquoExplaining the First Emancipation Social Movements and Abo-

lition in the US North 1776-1804rsquorsquo Mobilization 16439-54

Carter Gregg Lee 1986 lsquolsquoThe 1960s Black Riots Revisited City Level Explanations

of Their Severityrsquorsquo Sociological Inquiry 56210-28

Checchi Daniele and Jelle Visser 2005 lsquolsquoPattern Persistence in European Trade

Union Density A Longitudinal Analysis 1950-1996rsquorsquo European Sociological

Review 211-21

Choi-Fitzpatrick Austin and Tautvydas Juskauskas 2015 lsquolsquoUp in the Air Applying the

Jacobs Crowd Formula to Drone Imageryrsquorsquo Procedia Engineering 107273-81

Clauset Aaron Cosma Rohilla Shalizi and M E J Newman 2009 lsquolsquoPower-law

Distributions in Empirical Datarsquorsquo SIAM Review 51661-703

Clauset Aaron and Frederik W Wiegel 2010 lsquolsquoA Generalized Aggregation-

Disintegration Model for the Frequency of Severe Terrorist Attacksrsquorsquo Journal

of Conflict Resolution 54179-97

Clauset Aaron Maxwell Young and Kristian Skrede Gleditsch 2007 lsquolsquoOn the

Frequency of Severe Terrorist Eventsrsquorsquo Journal of Conflict Resolution 511-31

Crovella Mark E 2001 lsquolsquoPerformance Evaluation with Heavy Tailed Distributionsrsquorsquo

Pp 1-10 in Job Scheduling Strategies for Parallel Processing 7th International

Workshop edited by Dror G Feitelson and Larry Rudolph Berlin Germany

Springer Verlag

DeNardo James 1985 Power in Numbers The Political Strategy of Protest and

Rebellion Princeton NJ Princeton University Press

Earl Jennifer Andrew Martin John D McCarthy and Sarah A Soule 2004 lsquolsquoThe

Use of Newspaper Data in the Study of Collective Actionrsquorsquo Annual Review of

Sociology 3065-80

Feitelson Dror G 2006 lsquolsquoMetrics for Mass-count Disparityrsquorsquo Pp 61-68 in Proceed-

ings of the 14th IEEE International Symposium on Modeling Analysis and Simu-

lation of Computer and Telecommunication Systems (MASCOTS rsquo06)

Washington DC IEEE Computer Society

Fogelson Robert M 1971 Violence as Protest A Study of Riots and Ghettos New

York Doubleday

Francisco Ronald A nd European Protest and Coercion Data 1980ndash1995 (com-

puter file) Retrieved December 18 2012 (httpwebkueduronfranddata

indexhtml)

Franzosi Roberto 1987 lsquolsquoThe Press as a Source of Socio-historical Data Issues

in the Methodology of Data Collection from Newspapersrsquorsquo Historical Meth-

ods 205-15

Biggs 379

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 30: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Franzosi Roberto 1989 lsquolsquoOne Hundred Years of Strike Statistics Methodological

and Theoretical Issues in Quantitative Strike Researchrsquorsquo Industrial and Labor

Relations Review 42348-62

Haines Michael R and Inter-university Consortium for Political and Social

Research 2010 Historical Demographic Economic and Social Data The

United States 1790ndash2002 ICPSR 02896-v3 (computer file) Ann Arbor MI

Inter-university Consortium for Political and Social Research

Hibbs Douglas A 1973 Mass Political Violence A Cross-National Causal Analysis

New York John Wiley and Sons

Hosking J R M 1990 lsquolsquoL-moments Analysis and Estimation of Distributions Using

Linear Combinations of Order Statisticsrsquorsquo Journal of the Royal Statistical Society

Series B 52105-24

Inclan Marıa 2009 lsquolsquoSliding Doors of Opportunity Zapatistas and Their Cycle of

Protestrsquorsquo Mobilization 1485-106

Jacobs David and Stephanie L Kent 2007 lsquolsquoThe Determinants of Executions since

1951 How Politics Protests Public Opinion and Social Divisions Shape Capital

Punishmentrsquorsquo Social Problems 54297-318

Jenkins J Craig David Jacobs and Jon Agnone 2003 lsquolsquoPolitical Opportunities and

African-American Protest 1948ndash1997rsquorsquo American Journal of Sociology 109

277-303

Johnson Neil F Mike Spagat Jorge A Restrepo Oscar Becerra Juan Camilo

Bohorquez Nicolas Suarez Elvira Maria Restrepo and Roberto Zarama 2006

lsquolsquoUniversal Patterns Underlying Ongoing Wars and Terrorismrsquorsquo Retrieved Octo-

ber 27 2010 (httpxxxlanlgovabsphysics0605035)

Koopmans Ruud 1995 Democracy from Below New Social Movements and the

Political System in West Germany Boulder CO Westview

Kriesi Hanspeter Ruud Koopmans Jan Willem Dyvendak and Marco G Giugni

1995 New Social Movements in Western Europe A Comparative Analysis Min-

neapolis University of Minnesota Press

Lohmann Susanne 1994 lsquolsquoThe Dynamics of Informational Cascades The Monday

Demonstrations in Leipzig East Germany 1989ndash91rsquorsquo World Politics 4742-101

Maney Gregory M and Pamela E Oliver 2001 lsquolsquoFinding Collective Events

Sources Searches Timingrsquorsquo Sociological Methods and Research 30131-69

Mann Leon 1974 lsquolsquoCounting the Crowd Effects of Editorial Policy on Estimatesrsquorsquo

Journalism Quarterly 51278-85

Margetts Helen Peter John Scott Hale and Taha Yasseri 2016 Political Turbu-

lence How Social Media Shape Collective Action Princeton NJ Princeton Uni-

versity Press

McAdam Doug 1982 Political Process and the Development of Black Insurgency

Chicago University of Chicago Press

380 Sociological Methods amp Research 47(3)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 31: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

McAdam Doug John D McCarthy Susan Olzak and Sarah A Soule nd Dynamics

of Collective Action in the United States (computer file) Retrieved July 1 2012

(httpwwwstanfordedugroupcollectiveaction)

McAdam Doug and Yang Su 2002 lsquolsquoThe War at Home Antiwar Protests and

Congressional Voting 1965 to 1973rsquorsquo American Sociological Review 67696-721

McCarthy John D Clark McPhail and Jackie Smith 1996 lsquolsquoImages of Protest

Dimensions of Selection Bias in Media Coverage of Washington Demonstrations

1982 and 1991rsquorsquo American Sociological Review 61478-99

McCarthy John D Clark McPhail Jackie Smith and Louis J Crishock 1998

lsquolsquoElectronic and Print Media Representations of Washington DC Demonstra-

tions 1982 and 1991 A Demography of Description Biasrsquorsquo Pp 113-30 in Acts

of Dissent New Developments in the Study of Protest edited by Dieter Rucht

Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edition Sigma

McCarthy John D Larissa Titarenko Clark McPhail Patrick S Rafail and Bogus-

law Augustyn 2008 lsquolsquoAssessing Stability in the Patterns of Selection Bias in

Newspaper Coverage of Protest during the Transition from Communism in

Belarusrsquorsquo Mobilization 13127-46

McPhail Clark and John McCarthy 2004 lsquolsquoWho Counts and How Estimating the

Size of Protestsrsquorsquo Contexts 312-18

Morgan William R and Terry Nichols Clark 1973 lsquolsquoThe Causes of Racial Disor-

ders A Grievance-level Explanationrsquorsquo American Sociological Review 38611-24

Myers Daniel J 2010 lsquolsquoViolent Protest and Heterogeneous Diffusion Processes The

Spread of US Racial Rioting from 1964 to 1971rsquorsquo Mobilization 15289-321

Myers Daniel J and Beth Schaefer Caniglia 2004 lsquolsquoAll the Rioting Thatrsquos Fit to

Print Selection Effects in National Newspaper Coverage of Civil Disorders

1968ndash1969rsquorsquo American Sociological Review 69519-43

Oberschall Anthony R 1994 lsquolsquoRational Choice in Collective Protestsrsquorsquo Rationality

and Society 679-100

Oliver Pamela E and Gregory Maney 2000 lsquolsquoPolitical Processes and Local News-

paper Coverage of Protest Events From Selection Bias to Triadic Interactionsrsquorsquo

American Journal of Sociology 106463-505

Oliver Pamela E and Daniel J Myers 1999 lsquolsquoHow Events Enter the Public Sphere

Conflict Location and Sponsorship in Local Newspaper Coverage of Public

Eventsrsquorsquo American Journal of Sociology 10538-87

Oliver Pamela E and Daniel J Myers 2003 lsquolsquoThe Coevolution of Social Move-

mentsrsquorsquo Mobilization 81-24

Olzak Susan 1989 lsquolsquoAnalysis of Events in the Study of Collective Actionrsquorsquo Annual

Review of Sociology 15119-41

Olzak Susan 1992 The Dynamics of Ethnic Competition and Conflict Stanford CA

Stanford University Press

Biggs 381

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 32: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Opp Karl-Dieter 2009 Theories of Political Protest and Social Movements A Multi-

disciplinary Introduction Critique and Synthesis Abingdon UK Routledge

Ortiz David G Daniel J Myers Eugene N Walls and Maria-Elena D Diaz 2005

lsquolsquoWhere Do We Stand with Newspaper Datarsquorsquo Mobilization 10397-419

Popovic Srdja and Matthew Miller 2015 Blueprint for Revolution How to Use Rice

Pudding Lego Men and Other Non-violent Techniques to Galvanise Commu-

nities Overthrow Dictators or Simply Change the World Melbourne Australia

Scribe

Przeworski Adam 2009 lsquolsquoConquered or Granted A History of Suffrage Extensionsrsquorsquo

British Journal of Political Science 39291-321

Resnick Sidney I 2010 Heavy-tail Phenomena Probabilistic and Statistical Mod-

eling New York Springer

Richardson Lewis F 1948 lsquolsquoVariation of the Frequency of Fatal Quarrels with

Magnitudersquorsquo Journal of the American Statistical Association 43523-46

Rosenfeld Jake 2006 lsquolsquoDesperate Measures Strikes and Wages in Post-Accord

Americarsquorsquo Social Forces 85235-65

Rozenfeld Hernan D Diego Rybski Xavier Gabaix and Hernan A Makse 2011

lsquolsquoThe Area and Population of Cities New Insights from a Different Perspective on

Citiesrsquorsquo American Economic Review 1012205-25

Rucht Dieter and Friedhelm Neidhardt 1998 lsquolsquoMethodological Issues in Collecting

Protest Event Data Units of Analysis Sources and Sampling Coding Problemsrsquorsquo

Pp 65-89 in Acts of Dissent New Developments in the Study of Protest edited by

Dieter Rucht Ruud Koopmans and Friedhelm Neidhardt Berlin Germany Edi-

tion Sigma

Rummel Rudolph J 1963 lsquolsquoDimensions of Conflict Behavior within and between

Nationsrsquorsquo General Systems Yearbook 81-50

Seguin Charles 2016 lsquolsquoCascades of Coverage Dynamics of Media Attention to

Social Movement Organizationsrsquorsquo Social Forces 94997-1020

Shorter Edward and Charles Tilly 1974 Strikes in France 1830-1968 New York

Cambridge University Press

Soule Sarah A and Jennifer Earl 2005 lsquolsquoA Movement Society Evaluated Collective

Protest in the United Statesrsquorsquo Mobilization 10345-64

Spielmans John V 1944 lsquolsquoStrike Profilesrsquorsquo Journal of Political Economy 52319-39

Tarrow Sidney 1989 Democracy and Disorder Protest and Politics in Italy

1965ndash1975 Oxford UK Clarendon Press

Tarrow Sydney 2011 Power in Movement Social Movements Collective Action

and Politics 3rd ed New York Cambridge University Press

Tilly Charles 1978 From Mobilization to Revolution New York McGraw-Hill

Tilly Charles 1995 Popular Contention in Great Britain 1758ndash1834 Cambridge

MA Harvard University Press

382 Sociological Methods amp Research 47(3)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383

Page 33: Sociological Methods & Research Size Matters: Quantifying ...users.ox.ac.uk/~sfos0060/sizematters.pdf · waves, and revolutions is contentious collective action.’’ Opp (2009:38)

Tilly Charles and James B Rule 1965 Measuring Political Upheaval Princeton

NJ Centre of International Studies Princeton University

Titarenko Larissa John D McCarthy Clark McPhail and Boguslaw Augustyn

2001 lsquolsquoThe Interaction of State Repression Protest Form and Protest Sponsor

Strength during the Transition from Communism in Minsk Belarus 1990ndash1995rsquorsquo

Mobilization 6129-50

US Bureau of the Census 1960 Census of Population 1960 Vol 1 Characteristics

of the Population parts 2-52 Washington DC Government Printing Office

Virkar Yogesh and Aaron Clauset 2014 lsquolsquoPower-law Distributions in Binned

Empirical Datarsquorsquo Annals of Applied Statistics 889-119

Yip Paul S F Ray Watson K S Chan Eric H Y Lau Feng Chen Ying Xu Liqun

Xi Derek Y T Cheung Brian Y T Ip and Danping Liu 2010 lsquolsquoEstimation of

the Number of People in a Demonstrationrsquorsquo Australian and New Zealand Journal

of Statistics 5217-26

Author Biography

Michael Biggs is an associate professor of sociology at the University of Oxford Heis fascinated by the dynamics of protest waves from strikes in Chicago in 1886 toriots in London in 2011 He is also interested in the political uses of sufferingincluding self-immolation and hunger strikes

Biggs 383