Causal analysis of food insecurity and type 2 diabetes
Food insecurity and type 2 diabetes are strongly associated, and the association is easy to measure. Whether one causes the other is a different question, and answering it requires data that can order the two in time. This project asked what could be established from the National Health and Nutrition Examination Survey, one of the richest sources of population health data in the United States. The estimated effect was close to zero and not significant under any specification, and the more durable contribution is a precise account of why the design, rather than the sample, was the binding constraint.
The question, and the shape it had to take
The substantive question was the average treatment effect of experiencing food insecurity on the subsequent development of type 2 diabetes. Randomising anyone into food insecurity is not available as a research design, so the effect has to be recovered from observational data under assumptions, and the assumptions have to be defensible.
NHANES is cross-sectional. Each respondent is interviewed and examined once, which ordinarily forecloses any causal claim, because exposure and outcome are observed at the same instant and nothing in the data says which came first. The route around this was to build temporality out of the instrument itself. NHANES asks whether a respondent has had their glycohemoglobin measured by a clinician in the previous twelve months, and what that value was, and it separately measures glycohemoglobin directly at examination. A respondent who reports a non-diabetic value from the past year and presents a diabetic value at examination has, within the resolution of the survey, developed diabetes during a known interval. The food insecurity item asks about the same twelve months. That construction gives a genuine before and after inside a single-visit survey, and it is the only reason the question could be asked at all.
It also costs almost the entire sample. Restricting to respondents with both a recalled clinical measurement and an examination measurement reduced three pooled cycles to a few hundred people, which is the first thing to say about the precision of anything that follows.
The assumptions, stated before the estimate
Five conditions have to hold for the estimand to be a causal effect rather than a weighted association. Writing them down first, and reporting honestly on each, is most of the methodological work in a study of this kind.
| Assumption | Status |
|---|---|
| Positivity | Satisfied. Every covariate stratum contains both exposed and unexposed respondents. |
| Conditional exchangeability | Partially satisfied. Age, sex, race and ethnicity, educational attainment, and survey cycle were adjusted for. Dietary intake, physical activity, and family history were not measured in a usable form and remain unmeasured confounders. |
| No reverse causation | Assumed, not established. The construction places the exposure window before the outcome window, but food insecurity is treated as persistent over the interval, and diabetes can itself produce food insecurity through medical cost. The direction is argued rather than observed. |
| SUTVA | Reasonable. One respondent's food security is unlikely to alter another's diabetes risk within this sampling frame. |
| Accurate measurement | Weakest link. The prior glycohemoglobin value is recalled by the respondent from a clinical encounter up to a year earlier, and food insecurity is self-reported. Both are exposed to recall bias, and misclassification of either attenuates the estimate. |
Two of the five are compromised in ways no estimator can repair. That is the finding, and it is worth more than the point estimate that follows from it.
Why a doubly robust estimator
With a set of measured confounders there are two conventional routes to the average treatment effect. Model the outcome given treatment and covariates and average the contrast over the covariate distribution, which is g-computation. Or model the probability of treatment given covariates, the propensity score, and reweight the observed outcomes so that the two arms are comparable, which is inverse probability weighting. Each requires its own model to be right, and there is rarely a good reason to be confident in either.
Augmented inverse probability weighting combines them:
E[ µ1(X) − µ0(X) + A{Y − µ1(X)} / π(X) − (1−A){Y − µ0(X)} / {1 − π(X)} ]
where π(X) is the propensity score and µa(X) the expected outcome under treatment a. The estimator has a property that is genuinely remarkable the first time one meets it: it remains consistent if either the outcome model or the propensity model is correctly specified, without requiring both. Two chances at the truth, needing only one to land. This was the reason for choosing it over either component alone, given how little confidence the covariate set warranted.
Double robustness, demonstrated
Two hundred datasets of 1,000 observations are simulated from a known data-generating process with a true treatment effect of 2. Each is analysed three ways, and the sampling distribution of each estimator is drawn. Turn off either nuisance model and its corresponding single-model estimator moves off target while AIPW stays put. Turn off both and AIPW moves with them, which is the honest boundary of the property: doubly robust is not robust to being wrong twice.
Inverse probability weighting carries a small bias even when its model is correct, visible as a slight rightward shift. This is a finite-sample effect of dividing by estimated propensities near zero, and it is the practical reason the augmented estimator is usually preferred to the weighting estimator alone.
In the analysis itself the estimator was fitted with the AIPW package, with g-computation from RobinCar run alongside as a comparison, over a nested sequence of covariate sets ordered from most to least likely to confound: survey cycle, then age, then race and ethnicity, then sex, then educational attainment. Presenting a sequence rather than a single adjusted model shows how the estimate responds to adjustment, which is more informative than any one specification.
What the analysis returned
The estimated average treatment effect of food insecurity on developing diabetes was close to zero across every covariate set, and no specification reached statistical significance at the five percent level. The two estimators agreed with each other closely, which is reassuring about the fitting and says nothing about the identification.
The correct reading of a null result under these conditions is narrow. With a few hundred respondents, a rare outcome, self-reported exposure, and recalled outcome history, the study had little power to detect a moderate effect, and the measurement error present in both exposure and outcome attenuates estimates toward zero as a matter of course. The result is consistent with no effect and equally consistent with a real effect the design could not see. Published cohort studies with longitudinal follow-up have reported substantially elevated diabetes incidence among food insecure households; nothing here contradicts that, and nothing here supports it either.
The question that could not be asked
A second question, the effect of vitamin D deficiency on diabetes, was planned and abandoned. Serum vitamin D is measured only at examination, and because NHANES draws a fresh sample each cycle there is no way to observe a respondent's earlier level. No construction analogous to the recalled glycohemoglobin measurement exists in the instrument, so exposure and outcome are simultaneous and the effect is not identified.
An instrumental variable was considered, with average regional sun exposure as the leading candidate. It was rejected because every candidate examined violated the exclusion restriction: sunlight and geography influence diabetes risk through physical activity, diet, and the socioeconomic composition of place, not solely through vitamin D. An instrument that fails the exclusion restriction does not weaken an estimate; it invalidates it. Reporting the question as unanswerable was the correct output, and it belongs in the write-up rather than being quietly dropped.
What would make NHANES answer this
The constraint is not the size of NHANES or the quality of its measurement, both of which are excellent. It is that the survey observes each person once. The precedent for the fix already exists in the programme's own history: the NHANES Epidemiologic Follow-up Study re-interviewed participants from an earlier wave and, in doing so, produced a dataset that has supported serious causal work for decades, including the canonical analysis of smoking cessation and weight gain that appears in causal inference textbooks. A follow-up component attached to the modern continuous survey would convert an unmatched collection of cross-sectional measurements into a longitudinal resource, and would make questions of this kind answerable rather than merely askable.
That is the argument the project ended on, and it is the reason a null finding was still worth writing up. A study that establishes precisely why a question cannot be answered with the available data has produced information, provided it is explicit about which assumption failed and what data would repair it.
The manuscript itself is not circulated publicly, at the authors' request. This page describes the design, the estimator, and the methodological conclusions; the simulation above is an independent illustration of double robustness and uses no study data.