This tutorial follows the quantitative part of Methods of Social Research week by week. Each section takes the logic of one week of lectures and turns it into a short, runnable analysis in DAX, the browser-based environment Luiss uses for this course. You type Stata-style commands; DAX translates them into R behind the scenes and returns formatted output.
Every section has the same structure:
The best way to use this page. Keep it open next to DAX. Copy one block, run it, and read the output before moving on. Do not run the whole section at once: the point is to understand each step, not to collect output.
use and the full web address. Nothing needs to be
uploaded:use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/worldcup_democracy.dta", clear
describe
DAX is not desktop Stata. It understands a large subset of Stata syntax, but not all of it, and a few commands behave differently. Everything on this page has been written to work in DAX. If you try other commands, check the DAX survival kit at the end first.
All datasets are in the course data folder and load with
use "@DATA@/<file>". You never need to download
them.
| Week | File | What one row is |
|---|---|---|
| 1 | worldcup_democracy.dta |
a national team in one World Cup edition (537 rows) |
| 2 | chile1988.dta, turnout.dta |
a Chilean voter in 1988 (2,700); a US election (14) |
| 3 | vignettes.dta, battery.dta,
chile_samples.dta |
a respondent; a respondent; one simulated sample (500) |
| 4 | resume.dta, social.dta |
a fictitious CV (4,870); a Michigan voter (305,866) |
| 4 | minwage.dta, gini_polarization.dta |
a fast-food restaurant (358); a US Congress (33) |
| 5 | face.dta |
a US Senate race, 2000–2006 (119) |
| 6 | women.dta |
an Indian village (322) |
Science is a method, not a body of facts: a loop that runs from question to theory, hypotheses, observation and evaluation, and then back again. A hypothesis is useful only if it is falsifiable: you must be able to say in advance what evidence would count against it.
Our running question: does a country’s political regime affect its chances of playing in, and winning, the football World Cup? Two rival theories give opposite predictions. Sportswashing says autocracies invest more in sporting prestige (H1a: less democracy, more success). Democratic capacity says sustained success needs institutions that democracies build better (H1b: more democracy, more success).
Software enters only at step 4, and the reason is not speed: it is reproducibility. An analysis that cannot be re-run by someone else is an assertion, not a result.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/worldcup_democracy.dta", clear
describe
summarize goals polyarchy
tab1 year
tab1 winner finalist host
What you should see. 537 observations: one row is
one national team in one World Cup edition, 23 editions from 1930 to
2026. tab1 year shows 13 teams in 1930 and 48 in 2026.
Read it.
describe is always your first command. It tells you
what you actually opened.summarize, read the min and
max first: that is where coding errors hide.summarize polyarchy winner
* DAX draws one bar per distinct value: group polyarchy into bands of 0.05 first
generate poly_band = round(polyarchy * 20) / 20
histogram poly_band
tab1 winner
Read it. polyarchy (the V-Dem electoral
democracy index, 0 to 1) has two humps: many very
autocratic and many very democratic teams, fewer in between. The mean
hides this; the histogram shows it. For a 0/1 variable like
winner, the mean is the proportion of
ones: 23 champions out of 537 team-editions gives 0.043, so
4.3% of appearances end with the trophy. Rare outcomes are hard to
explain statistically.
A concept (democracy) becomes a variable through explicit decisions. Here we turn the Polity5 score (−10 to +10) into three regime types, using the conventional cut-offs:
recode polity2 (-10/-6 = 0) (-5/5 = 1) (6/10 = 2), generate(regime)
label define reglbl 0 "Autocracy" 1 "Anocracy" 2 "Democracy"
label values regime reglbl
tab1 regime
* every champion is also a finalist, so this gives 0 / 1 / 2
generate outcome = winner + finalist
tab2 winner finalist
Always check a variable you just built.
tab2 winner finalist confirms that outcome
means what you think it means. recode leaves missing values
missing, which is what you want. Had you built regime with
comparisons such as polity2 >= 6, every country with no
data at all would have been classified as a democracy, because Stata
treats a missing value as larger than any number.
scatter polyarchy year
ttest polyarchy, by(winner)
ttest polyarchy, by(finalist)
tab2 regime winner, column
Read it.
scatter polyarchy year shows World Cup participants
becoming more democratic over time, but that is because the
world became more democratic. Two variables that both trend
over time will correlate, whether or not they are related. A
research-design problem appeared before any statistics.ttest prints the mean of each group
and then tests whether they differ by more than chance would produce.
Read the means first and the p-value second.Evaluation (step 5). H1a, the sportswashing prediction, finds no support. H1b survives, but only for reaching the final. A theory that survives a test is corroborated, not proved. And “not significant” does not mean “no effect”: it means the evidence is too thin to tell the groups apart.
polyarchy
and polity2) agree?polity2. Does the
conclusion change?ttest goals, by(host)
correlate polyarchy polity2
scatter polyarchy polity2
ttest polity2, by(finalist)
Hosts do score more, but hosts are not a random draw of countries
(they are richer and often strong football nations), and they play at
home. The two democracy measures correlate strongly (r = 0.93) but not
perfectly: every measure is a decision, and different decisions place
some countries differently. And the conclusion does
change: with polity2 the finalists are still more
democratic on average (5.89 vs 4.26), but the difference is no longer
statistically significant (p ≈ 0.12). A finding that depends on which
indicator you pick is a fragile finding.
A concept is a mental construct; it exists only if it draws a boundary between cases that are A and cases that are not-A. To study it empirically we go from concept → property of a unit of analysis → operational definition (the question, the coding rule, the unit of measurement) → variable.
The level of measurement decides which operations make sense:
| Level | Example | Legitimate operations |
|---|---|---|
| Nominal | vote intention | equal / different |
| Ordinal | education level | also greater / less than |
| Interval / ratio | age, income | also differences, means |
The software does not know the level of measurement of your variables. You do.
The chile1988 data come from a survey of 2,700 Chilean
adults interviewed a few months before the 1988 plebiscite on Pinochet’s
rule.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/chile1988.dta", clear
describe
tab1 vote
summarize vote
What you should see. tab1 vote: 889
respondents will vote No, 35.11% of the 2,532 valid answers.
summarize vote returns a mean of 2.844.
Read it.
tab1 computes
percentages on valid answers only: 168 are missing. Out
of all respondents, the No share is 889 / 2,700 = 32.9%. Always
ask: percent of what?vote is arithmetically correct and
substantively meaningless: vote is nominal (1 abstain,
2 No, 3 undecided, 4 Yes). A Yes voter is not “more” of anything than a
No voter.tab1 educ_raw
recode educ_raw (1=1) (3=2) (2=3), generate(educ)
label define edlbl 1 "Primary" 2 "Secondary" 3 "Post-secondary"
label values educ edlbl
tab2 educ_raw educ
tab1 educ
Read it. The original file coded education
alphabetically (Primary, Post-secondary, Secondary), so
post-secondary sat between the other two and the cumulative percentages
mixed apples and oranges. An ordinal variable is only ordinal if its
codes are in the right order. After the recode, 82.82% of respondents
have at most a secondary education: a cumulative percentage that finally
means something. tab2 educ_raw educ checks that every old
value lands in exactly one new value.
summarize income, detail
histogram income
recode age (18/29=1) (30/49=2) (50/70=3), generate(agegrp)
tab1 agegrp
generate female = (sex == 1)
summarize female
What you should see. Income: mean 33,876, median
15,000. Age groups: 929, 1,103 and 667 respondents. female
has a mean of 0.511.
Read it.
Observed value = true value + systematic error +
random error. turnout.dta compares the
turnout reported in the American National Election Studies (ANES) with
official turnout in 14 US elections.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/turnout.dta", clear
generate vep_rate = total / vep * 100
generate bias = anes - vep_rate
summarize bias
What you should see. bias: mean 16.84,
minimum 8.58, maximum 22.49.
Read it. The minimum is positive: the survey overestimates turnout in all 14 elections. The error is systematic, not random. The instrument is reliable (stable) but not valid (stably wrong in the same direction). Repeating the survey would never reveal this, because the bias is reproduced every time. The usual suspects are social desirability and non-response.
In chile1988.dta, statusquo is an index of
support for the regime. If it is a valid measure, people who say they
will vote Yes (for Pinochet) should score much higher than people who
will vote No. This is known-groups validity.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/chile1988.dta", clear
by vote: summarize statusquo
No voters average about −0.91 and Yes voters about +0.94. The measure separates the known groups very clearly. But both measures come from the same survey: if fear biased both answers in the same way, they would agree with each other and both be wrong.
A survey turns concepts into questions, answers into response categories, several answers into scales, and a few hundred respondents into statements about populations. Each step can go wrong in its own way.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/vignettes.dta", clear
describe
* how much say do YOU have in getting the government to address
* issues that interest you? (1 = no say at all ... 5 = unlimited say)
ttest self, by(china)
Read it. On the same 1–5 scale, Chinese respondents report more political say (2.62) than Mexican respondents (1.83), in a year when Mexico had just had its first change of governing party in 71 years and China had no competitive elections. Either Chinese citizens really have more say, or the categories mean something different in Beijing and in Mexico City. Anchoring vignettes exist to tell these two stories apart.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/battery.dta", clear
summarize a1 a2 a3 a4 a5
correlate a1 a2 a3 a4 a5
What you should see. 2,709 respondents. Items a2–a5 have means between 4.55 and 4.80 and correlate positively with each other (0.31 to 0.51). Item a1 has a mean of 2.41 and correlates negatively with all the others.
Read it. a1 (“Am indifferent to the feelings of others”) measures the same trait in the opposite direction. It has to be reversed before the items are added up: reversed = (min + max) − original = 7 − a1.
generate a1r = 7 - a1
tab2 a1 a1r
correlate a1r a2 a3 a4 a5
generate agree = a1r + a2 + a3 + a4 + a5
summarize agree, detail
histogram agree
What you should see. After reversing, all ten correlations are positive, with a mean of 0.3325. The scale runs from 5 to 30, with mean 23.2 and median 24.
Read it. Cronbach’s alpha by hand: α = n·r̄ / [1 + r̄(n − 1)] = 5 × 0.3325 / (1 + 0.3325 × 4) = 0.714, just above the conventional 0.70 threshold. The mean (23.2) below the median (24) signals a long left tail: most people describe themselves as agreeable, a few do not (social desirability may help). Remember that alpha measures reliability: it says nothing about validity or unidimensionality.
chile_samples.dta contains 500 random
samples drawn from the same 2,531 Chilean voters, at three
sample sizes. In that population the true share of No voters is
35.12%.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/chile_samples.dta", clear
* DAX shows two decimals: express the estimates in percentage points first
generate pp50 = pno50 * 100
generate pp200 = pno200 * 100
generate pp1000 = pno1000 * 100
summarize pp50 pp200 pp1000
* round to whole points to draw the histograms
generate pct50 = round(pp50)
generate pct1000 = round(pp1000)
histogram pct50
histogram pct1000
What you should see (in percentage points).
| Sample size | Mean of the 500 estimates | Std. dev. (= sampling error) | Min | Max |
|---|---|---|---|---|
| 50 | 35.10 | 6.67 | 16.0 | 56.0 |
| 200 | 35.26 | 3.24 | 27.0 | 46.0 |
| 1,000 | 35.01 | 1.21 | 31.8 | 40.5 |
Read it. All three averages land on the true 35.1%: random sampling is unbiased. What changes with n is the spread: 6.7, then 3.2, then 1.2 points. One sample of 50 said 16%, another 56%. Theory predicts the same numbers: √(pq/n) × √(1 − n/N) gives 6.7, 3.2 and 1.2 points.
Size does not fix bias. In 1936 the Literary Digest collected 2.4 million answers and missed Roosevelt’s landslide by about 20 points: its sampling frame (telephone directories, car registrations) and its 24% response rate selected wealthier, better-educated respondents. Twice the ballots would have reproduced the same bias twice as precisely.
Run summarize age50 age200 age1000. Compare the three
standard deviations with s/√n, where s = 14.76 is the standard deviation
of age in the population.
summarize age50 age200 age1000
s/√n gives 14.76/√50 ≈ 2.09, 14.76/√200 ≈ 1.04 and 14.76/√1000 ≈ 0.47. DAX gives 2.03, 1.03 and 0.36. The first two are close to the formula; the third is clearly smaller, because 1,000 out of 2,531 is a large fraction of the population. With the finite population correction, 0.47 × √(1 − 1,000/2,531) ≈ 0.47 × 0.78 ≈ 0.36.
Saying that X causes Y says more than “X and Y go together”. It says that changing X would change Y. Empirically we need three things: covariation, the right causal direction, and control of other causes. The third is the hard one.
Every unit has two potential outcomes: what happens with the treatment, Y(1), and without it, Y(0). We only ever observe one of them; the other is the counterfactual. Random assignment solves the problem on average: if units are assigned to treatment by chance, the two groups are equal on everything, measured or not, and the difference in mean outcomes estimates the average causal effect.
Bertrand and Mullainathan (2004) sent about 5,000 fictitious CVs in response to real job ads in Boston and Chicago. The first name on each CV was randomly assigned: stereotypically African-American or stereotypically white. The outcome: did the employer call back?
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/resume.dta", clear
describe
tab1 black
tab2 black call, row
ttest call, by(black)
What you should see. White names: 235 callbacks out
of 2,435 CVs (9.65%). Black names: 157 out of 2,435
(6.45%). In ttest: means 0.0965 and
0.0645, t ≈ 4.11, p < 0.001.
Read it.
ttest
compares two callback rates.* 0 = white man, 1 = Black man, 2 = white woman, 3 = Black woman
generate group = black + 2*female
label define grouplab 0 "white man" 1 "Black man" 2 "white woman" 3 "Black woman"
label values group grouplab
tab2 group call, row
Read it. Women: 9.89% vs 6.63% (3.3 points). Men: 8.87% vs 5.83% (3.0 points). Discrimination is almost identical for both sexes. But only 1,124 CVs carry a male name against 3,746 female ones: smaller subgroups mean less precise estimates.
Gerber, Green and Larimer (2008) randomly assigned 305,866 Michigan voters to four groups before the August 2006 primary: no letter (control), a Civic Duty letter, a Hawthorne letter (“you are being studied”), and a Neighbors letter listing the household’s and its neighbours’ past turnout.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/social.dta", clear
tab1 message
tab2 message primary2006, row
What you should see. Turnout: Control 29.7%, Civic Duty 31.5%, Hawthorne 32.2%, Neighbors 37.8%.
Read it. Being reminded of a duty: +1.8 points. Being watched: +2.6. Being watched by your neighbours: +8.1 points, from a single letter.
by message: summarize primary2004 yearofbirth hhsize
Read it. In all four groups the share who voted in 2004 is 0.40 or 0.41, the mean year of birth is between 1956.15 and 1956.34, and the mean household size is 2.18 or 2.19. The groups were balanced before the letters were sent. We check only pre-treatment variables: a letter sent in 2006 cannot change the past, while anything measured afterwards could be an effect of the letter.
In resume.dta, which of the four sex × race groups is
called back least? In social.dta, why would it make no
sense to add primary2006 to the balance check?
Black men are called back least (5.83%). primary2006 is
the outcome: differences in it are what we want to
explain, not a check on whether the groups were comparable
beforehand.
Most causes we care about (a law, a religion, a regime) cannot be assigned at random. In an observational study the researcher observes who is exposed to X; nobody assigns it. Realism improves, but internal validity is no longer guaranteed: it must be argued. Every comparison builds a counterfactual on an assumption, so name the assumption.
In April 1992 New Jersey raised its minimum wage from $4.25 to $5.05; neighbouring Pennsylvania did not. Card and Krueger (1994) surveyed fast-food restaurants on both sides of the border before (February) and after (November) the increase.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/minwage.dta", clear
describe
tab1 nj
by nj: summarize wageBefore wageAfter
What you should see. 291 restaurants in New Jersey, 67 in Pennsylvania. Average starting wage in New Jersey: $4.61 before, $5.08 after. In Pennsylvania: $4.65 before, $4.61 after.
The law did change wages in New Jersey. Did it reduce jobs? Economic theory predicts that employers would shift from full-time to part-time workers. Compute the share of full-time employees:
generate fullshare_before = fullBefore / (fullBefore + partBefore)
generate fullshare_after = fullAfter / (fullAfter + partAfter)
by nj: summarize fullshare_before fullshare_after
* comparison 1: New Jersey vs Pennsylvania, after the law
ttest fullshare_after, by(nj)
What you should see. Full-time share after the law:
0.27 in Pennsylvania, 0.32 in New Jersey (ttest: means
0.2723 and 0.3204, t ≈ −1.43, p ≈ 0.16). In New Jersey before the law:
0.30.
Read it. Two different counterfactuals:
Neither comparison suggests that the higher minimum wage destroyed full-time jobs. Neither is as clean as a randomised experiment, and the cross-section difference is not statistically significant.
tab2 chain nj, column
Read it. Burger King is 46% of Pennsylvania restaurants and 41% of New Jersey ones; KFC is 15% and 22%. The groups are similar but not identical on what we can observe, and we cannot check what we do not observe. This is the price of not randomising.
Going further (optional). Combine the two comparisons. The full-time share rose by 2.4 points in New Jersey and fell by 3.8 points in Pennsylvania. The difference of the differences is about 6.2 points. This estimator, called difference-in-differences, assumes that without the law the two states would have followed parallel trends.
For two categorical variables we use a contingency
table (tab2), with percentages computed within the
categories of the independent variable. For two interval variables we
use a scatterplot and the correlation
coefficient r, which runs from −1 to +1 and measures
how closely the points cluster around a straight line.
gini_polarization.dta has one row per US Congress,
1948–2012: income inequality (the Gini coefficient) and polarization
(the distance between the median Republican and the median Democrat in
Congress, on the DW-NOMINATE ideological scale).
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/gini_polarization.dta", clear
scatter polarization gini
correlate gini polarization year
scatter polarization year
scatter gini year
What you should see. r(gini, polarization) = 0.94. But also r(polarization, year) = 0.94 and r(gini, year) = 0.87.
Read it. A correlation of 0.94 is very high. But both variables rise steadily over time, and two variables that trend together will correlate whether or not one causes the other. This is the trap you met in Week 1 with the World Cup. A strong correlation is the start of a causal argument, not its end.
Is the correlation between polarization and
gini the same in the first and in the second half of the
period? What would you conclude if it were much weaker within each
half?
correlate gini polarization if year < 1980
correlate gini polarization if year >= 1980
scatter polarization gini if year < 1980
scatter polarization gini if year >= 1980
Before 1980 (16 Congresses) the correlation is negative, about −0.68; from 1980 onwards (17 Congresses) it is about +0.97. The overall 0.94 is driven by the period in which both series climb together. A relationship that changes sign across periods is a warning that a third factor, here the passage of time and everything that changed with it, may be doing the work.
Correlation tells us how closely two variables move together. Regression tells us by how much Y is expected to change when X changes by one unit, and lets us predict Y from X:
Y = α + βX + ε
Ordinary least squares (OLS) chooses α and β so that the squared prediction errors are as small as possible.
Todorov et al. (2005) showed people the faces of the two main
candidates in US Senate races, for one second, and asked which one
looked more competent. They had no other information.
face.dta covers 119 races from 2000 to 2006.
d_comp is the share of participants who rated the
Democrat as more competent.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/face.dta", clear
describe
generate d_share = d_votes / (d_votes + r_votes)
generate r_share = r_votes / (d_votes + r_votes)
generate diff_share = d_share - r_share
summarize diff_share d_comp
Read it. diff_share is the Democrat’s
two-party margin: positive when the Democrat won, negative when the
Republican won. It ranges from −0.77 to +0.70, with a mean close to zero
(0.01).
scatter diff_share d_comp
correlate diff_share d_comp
regress diff_share d_comp
What you should see. r = 0.433. In
regress:
| Coefficient | (Std. error) | |
|---|---|---|
| d_comp | 0.660*** | (0.127) |
| Constant | −0.312*** | (0.066) |
Observations 119, R² = 0.187.
Read it.
d_comp = 0.The regression line is a prediction machine. With α = −0.312 and β = 0.660:
| d_comp | Predicted margin = −0.312 + 0.660 × d_comp |
|---|---|
| 0.3 | −0.114 (Republican ahead by 11 points) |
| 0.5 | +0.018 (almost a tie) |
| 0.7 | +0.150 (Democrat ahead by 15 points) |
Prediction is not explanation. Competence ratings predict election results, but nobody randomised the faces. Incumbents, for example, may look more confident in their photographs and win more often for other reasons. A regression coefficient describes an association; whether it is a causal effect depends on the design, not on the software.
generate d_win = (d_share > r_share)
summarize d_win
The Democrat won 51.3% of the 119 races.
A regression coefficient has a causal interpretation only when the design justifies it. In a randomised experiment, regressing the outcome on the treatment dummy gives exactly the difference in means, that is, the average causal effect. In observational data, the coefficient mixes the effect of X with the effect of everything correlated with X that is not in the model.
In India, since the mid-1990s, one third of village council (Gram Panchayat) head seats have been reserved for women, and the councils were selected at random. Chattopadhyay and Duflo (2004) asked: do women leaders invest in different public goods? In the villages studied, women complained most about drinking water, men about irrigation.
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/women.dta", clear
describe
tab2 reserved female
Read it. All 108 reserved villages have a female council head; only 16 of the 214 unreserved ones do. The randomised policy really changed who governs.
ttest water, by(reserved)
regress water reserved
regress irrigation reserved
What you should see. Drinking-water facilities:
14.74 in unreserved villages, 23.99 in reserved ones. In
regress water reserved: coefficient on
reserved = 9.252 (3.948), constant =
14.738. In regress irrigation reserved: coefficient =
−0.369 (1.122).
Read it.
Why do the p-values differ? ttest in
DAX does not assume equal variances in the two groups (Welch test, p ≈
0.07), while regress does (p ≈ 0.02). Same means, different
assumptions about the noise. With a difference this close to the
conventional threshold, the honest conclusion is “suggestive evidence”,
not a proof.
With a categorical treatment, i. tells DAX to create one
dummy per category, leaving out the first one as the
reference group:
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/resume.dta", clear
* the four sex x race groups from Week 4
generate group = black + 2*female
label define grouplab 0 "white man" 1 "Black man" 2 "white woman" 3 "Black woman"
label values group grouplab
regress call i.group
What you should see. Constant = 0.089 (the callback rate of white men). Coefficients: Black man −0.030, white woman 0.010, Black woman −0.022.
Read it. Each coefficient is the difference from the
reference group (white men, 8.87%): Black men 5.83% (−3.0 points), white
women 9.89% (+1.0), Black women 6.63% (−2.2). These are exactly the
percentages you computed with tab2 group call, row in Week
4. Notice that none of the three coefficients has a star: each subgroup
is small, so each single comparison is imprecise.
regress call black
regress call black female
Read it. The coefficient on black is
−0.032 with or without female in the model. Because names
were randomised, race of the name is unrelated to the sex of the name,
and adding a control does not change the estimated
effect. In observational data, adding a relevant control usually does
change the coefficient, and that change is exactly what multiple
regression is for: estimating the effect of X holding the other
variables constant.
Which controls? Control for variables that come before X and affect Y (confounders). Do not control for variables that are consequences of X: you would remove part of the effect you want to measure.
In face.dta, could you interpret the coefficient on
d_comp as the causal effect of looking competent? Name one
confounder and say in which direction it would bias the estimate.
No. Competence ratings were not assigned at random. Incumbency is a plausible confounder: incumbents may look more experienced and confident, and they win more often for many other reasons (visibility, money, constituency service). If incumbents are both rated more competent and more likely to win, the coefficient overstates the effect of appearance.
If you want one script that touches the core of each week, start from this. Run it in the Script Editor, then go back to each section to interpret the output.
* Week 1 - a first test
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/worldcup_democracy.dta", clear
describe
summarize polyarchy winner
generate poly_band = round(polyarchy * 20) / 20
histogram poly_band
ttest polyarchy, by(finalist)
* Week 2 - levels of measurement
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/chile1988.dta", clear
tab1 vote
recode educ_raw (1=1) (3=2) (2=3), generate(educ)
tab2 educ_raw educ
summarize income, detail
* Week 3 - scales and samples
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/battery.dta", clear
generate a1r = 7 - a1
correlate a1r a2 a3 a4 a5
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/chile_samples.dta", clear
summarize pno50 pno200 pno1000
* Week 4 - experiments
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/resume.dta", clear
tab2 black call, row
ttest call, by(black)
* Week 4 - observational data and correlation
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/gini_polarization.dta", clear
correlate gini polarization year
* Week 5 - regression and prediction
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/face.dta", clear
generate diff_share = (d_votes - r_votes) / (d_votes + r_votes)
scatter diff_share d_comp
regress diff_share d_comp
* Week 6 - regression and causation
use "https://raw.githubusercontent.com/francescovisconti/MSR-data/main/women.dta", clear
regress water reserved
DAX translates Stata-style commands into R. Most of what you need in this course works, but a few things do not, and some outputs look different from desktop Stata.
| Works | Notes |
|---|---|
use "https://…/file.dta", clear |
load data from a web address |
describe, browse |
|
summarize x y z, summarize x, detail |
always list the variables; shows the median, not the full percentiles |
by g: summarize x |
summary statistics within groups |
tab1 x, tab2 x y, row column |
frequency and contingency tables |
generate, recode …, generate() |
arithmetic, log(), comparisons such as
(x > 10) |
label define, label values |
|
ttest y, by(g) |
Welch test: does not assume equal variances |
correlate x y z, scatter y x |
correlate … if and scatter … if work |
histogram x |
draws one bar per distinct value; for continuous variables, group them first (see below) |
regress y x1 x2, regress y i.g |
output with standard errors in parentheses and stars; fails on very large files (see below) |
| Does not work (yet) | Do this instead |
|---|---|
generate … if … |
build the variable for everyone,
e.g. generate d = (x == 3) |
recode x (… = .) (recoding to missing) |
compare groups with by g: summarize |
ttest … if …, regress … if … |
by g: summarize; or a dummy plus
ttest y, by(dummy) |
comparisons with text, e.g. (state == "PA") |
use the numeric version of the variable |
tabstat, codebook, collapse,
preserve, bysort, egen |
summarize, by g: summarize,
tab1 |
histogram x, bin(20) |
bin() is ignored: build bands with
generate xb = round(x * 10) / 10 |
round(x, 0.1) (rounding to a unit) |
round(x * 10) / 10 |
regress on very large datasets
(e.g. social.dta, 305,866 rows) |
fails with “offset is out of range”: use tab2 and
by g: summarize on large files |
log, cd, pwd,
clear all |
not needed: load data from the web |
Reading DAX output. summarize rounds to
two decimals: if you need more precision, rescale the variable first
(e.g. multiply a proportion by 100). In ttest, a p-value
printed as 0 means p < 0.0001.
When a command fails, DAX shows an error from its R
engine and does not add the command to your history. Read the message,
check the table above, and simplify. If DAX starts returning odd errors
after a failed command, reload the data with
use …, clear.
The point of this part of the course is not to learn a software package. It is to learn how to move from a research question to a defensible empirical analysis. The commands are the easy part. The habits are what matter:
Textbooks
Studies used in the examples
Data sources
resume, social,
minwage, face, women,
vignettes and turnout come from Imai &
Bougher (2021); gini_polarization is derived from their
congress and USGini files.