What the air is doing to India
A district-level investigation of air pollution and public health across 150 Indian districts — what the raw data shows, what 10 statistical tests confirmed, and what the findings demand from policy.
The scale of the problem
India is home to 21 of the world's 30 most polluted cities. But aggregated rankings obscure what is actually happening at ground level: a sustained, multi-year health emergency that is invisible in any single metric and impossible to miss when you look at 150 districts simultaneously over six years. The first number you encounter when you average the dataset stops the analysis cold.
That average doesn't mean most districts hover around it. The distribution is right-skewed and harsh: the worst days in the worst districts exceed 200 µg/m³, and no district in the panel consistently stays below NAAQS. When sorted by mean PM2.5, the gradient between states is 3–5×:
The Indo-Gangetic Plain states — Punjab, Haryana, UP, Bihar, Delhi — run structurally above 60 µg/m³. Southern states cluster near 20–25. This isn't just urban versus rural. It's topography (the Plain is flanked by the Himalayas, limiting vertical dispersion), industry mix, agricultural burning, and decades of under-investment in enforcement. Geography has done for the South what regulation hasn't managed in the North.
On the health side: — total respiratory presentations and — cardiovascular across the panel. Before any statistical test, the geographic overlap is already damning — the five most polluted states also top the respiratory case rankings.
What we built and how we cleaned it
Three independent data sources were joined: 250,008 daily air-quality readings from 150 CPCB/NDAP monitoring stations (2018–2023), health-facility reports from HMIS covering respiratory, cardiovascular, and diarrhoea cases, and district-level demographic data from Census 2011 and NFHS projections. The merged panel covers 10,800 district-month observations.
One distributional fact changes how you interpret everything downstream: PM2.5 is wildly right-skewed. The mode sits around 30–45 µg/m³ but the tail stretches past 200. This means linear models will systematically underfit the upper end — the part where health consequences are most severe. We addressed this by fitting log-linear models alongside linear ones in the dose-response analysis (the log-linear fit is tighter above 80 µg/m³).
PM2.5 and respiratory disease — the direct link
The simplest version of our central question: at the district-month level, does higher PM2.5 co-occur with more respiratory cases? If the association is spurious, the scatter is a round cloud. If it's real, you see structure.
At district-month level, Pearson r = 0.38. Meaningful for data this granular and noisy, but the district-month unit mixes signal with weather variance, reporting lags, and seasonal cycles. When we aggregate to district means and control for confounders, the picture sharpens considerably.
At district level, the raw Pearson r rises to 0.80. After controlling for urbanisation, literacy, and population (partial correlation), it holds at 0.79. The confounders barely shift the number. The PM2.5–respiratory correlation is not an artefact of richer or more urban areas having both worse monitoring and more hospitals — it survives conditioning on all of those.
The heatmap shows the correlation structure across all variables. PM2.5, PM10, and NO₂ travel together (r > 0.7 between them) — they share emission sources, so disentangling individual effects requires the controlled analysis we run later. All three correlate positively with respiratory and cardiovascular cases. Literacy and urban percentage are mildly negatively correlated with disease — wealthier, more literate districts seek care earlier and report lower raw case counts. This is a known healthcare-access bias that we control for throughout.
Dose-response: more pollution, more disease, every single step
A correlation coefficient tells you two things move together. A dose-response curve tells you something stronger: as the exposure increases by any amount, the outcome consistently increases too — across the entire range, with no plateau and no reversal. That monotonic relationship is one of the classical epidemiological signatures of a genuine causal exposure rather than a statistical artefact.
The curve is unambiguous. At PM2.5 around 17 µg/m³ — still above the WHO guideline but near the low end of our data — the average district-month sees roughly 22 respiratory cases per 100,000. By PM2.5 of 177 µg/m³ (the highest bin), that number is 129 — nearly 6× higher. Every single bin step is above the previous one. There is no plateau, no U-shape, no safe range.
The log-linear fit (dotted line) tracks the data slightly better than a linear model above 80 µg/m³, suggesting the health effect accelerates at very high concentrations rather than levelling off. This matters for policy: the marginal health cost of the last few µg/m³ in a heavily polluted district is higher than the marginal cost in a moderately polluted one.
From 17 to 177 µg/m³, respiratory cases per 100k increase nearly 6-fold without a single reversal. There is no PM2.5 level in this data that is safe — only less unsafe.
Where in India is worst — and why it's structural
The red bars exceed 40 µg/m³ — the NAAQS guideline for annual PM2.5. Every Indo-Gangetic Plain state sits in this zone. The reasons are partly topographic: the Plain is flanked by the Himalayas to the north, limiting vertical dispersion and trapping emissions at breathing level. Partly industrial: high coal combustion, brick kilns, vehicular density. Partly agricultural: paddy stubble burning in Punjab and Haryana creates an identifiable October–November spike that repeats every year.
None of the southern states exceed NAAQS on annual average. This isn't primarily better enforcement — it's the Western and Eastern Ghats providing natural ventilation, a longer monsoon flushing period, and far less agricultural burning. Geography has done the work that regulation hasn't managed in the North. Whether that continues as southern India industrialises is an open question this dataset can't answer but raises urgently.
Why winter is the killing season
National PM2.5 has tracked between 55–70 µg/m³ annually from 2013 to 2023 with no sustained downward trend. Respiratory cases follow the same flat-to-rising trajectory. The co-movement is the dominant pattern across all years: when pollution rises, so does disease; when pollution holds flat, so does disease burden.
Within each year the mechanism is seasonal: temperature inversions in winter (November–February) cap the atmospheric boundary layer and trap pollutants near ground level. Add stubble burning in Punjab and Haryana each October–November, reduced wind speeds, and higher combustion from heating, and the winter spike returns predictably every year. The annual data here represents each year's mean — the underlying intra-year variation is 1.5–2× between winter peaks and monsoon troughs.
Pollution leads, disease follows — in every single district
Correlation is symmetric. Establishing temporal direction requires exploiting the time structure of the data: does PM2.5 in month T predict respiratory cases in month T+k better than cases predict future PM2.5? Cross-correlation at lags answers this at the national level; Granger causality answers it at the individual district level.
The cross-correlation peaks at a positive lag of 1–2 months: today's PM2.5 best predicts respiratory cases 1–2 months later. The reverse direction (today's disease predicting future PM2.5) is weaker across all lags. The signal runs forward in time.
The unanimity is unusual. In most health datasets, 50–70% of units pass this test due to local idiosyncrasies, data gaps, and heterogeneous lag structures. Getting 150 out of 150 means the PM2.5 → respiratory pathway is not being driven by a few extreme outlier districts pulling up the national average. It operates uniformly across every region, pollution level, and demographic profile in the panel. That uniformity is itself strong evidence that the underlying mechanism is biological and physical, not statistical.
Quantifying the epidemiological burden
Relative risk (RR) and population attributable fraction (PAF) are the standard epidemiological tools for translating a statistical association into a claim about burden. RR answers: how much more likely is disease in exposed vs unexposed? PAF answers: what fraction of all cases could be eliminated if the exposure were removed?
Against the NAAQS threshold: exposed districts show 2.26× the respiratory rate of unexposed. A PAF of 32.8% means roughly one-third of all respiratory cases in our panel are attributable to above-NAAQS exposure — and would not occur at that rate if those districts met the standard.
Against the WHO guideline: RR rises to 3.96× and PAF to 74.6%. Only 110 district-months in the entire panel stayed below 15 µg/m³ — practically the whole dataset is "exposed" in WHO terms. If India achieved WHO levels, three-quarters of the respiratory burden would theoretically be preventable. That is not a near-term target. But it calibrates the gap between where things are and where the evidence says they should be.
Machine learning confirms the same pattern
Panel fixed-effects regression answers a precise question: within a single district, holding all its fixed characteristics constant, does PM2.5 over time predict respiratory cases over time? This removes any confounding from geography, infrastructure, or baseline health — and looks only at variation within each district.
PM2.5 has a within-district coefficient of 0.655 at p = 0.00002. For the same district, a 1 µg/m³ rise in monthly PM2.5 is associated with approximately 0.65 additional respiratory cases per district-month, with all fixed district characteristics held constant. PM10 falls short of significance when PM2.5 is included — expected, since the two pollutants are highly collinear. The within-R² of 0.50 means the model explains roughly half of all within-district variation in monthly respiratory cases. For monthly health data, that is strong.
A Random Forest trained on the same features reaches an out-of-sample R² of 0.81 — higher because it captures the log-linear dose-response shape and pollutant interaction effects that linear models approximate poorly. Variable importance from the Random Forest ranks PM2.5 first across all metrics: permutation importance, SHAP, and mean decrease in impurity all agree. The ML result and the econometric result point to the same variable.
The socioeconomic multiplier — same air, worse outcomes
K-Means clustering on all standardised pollution and demographic features surfaces four natural groups. Three are intuitive. The fourth disrupts the narrative.
The critical-pollution cluster (Delhi NCR, mean PM2.5 ≈ 110 µg/m³) has the worst air. The moderate cluster (south India) has the best. In between sits the at-risk middle group. And then there is the fourth: predominantly UP and Bihar districts, with PM2.5 around 70 µg/m³ — well below Delhi. Their respiratory case rates are 2.7× higher than even Delhi's — not because their air is worse, but because the healthcare infrastructure to absorb the damage is far weaker.
The districts with the worst health outcomes don't have the worst air. They have bad air and no healthcare infrastructure to absorb the consequences.
Lower literacy, more people working outdoors in hazardous conditions, weaker primary care networks, less ability to seek treatment early — these factors compound with pollution rather than operating in parallel. A clean-air policy that ignores socioeconomic vulnerability will leave the most-affected populations behind even if it succeeds on its emissions target. Both levers are necessary. The populations who need both the most are the ones getting neither.
Pollution doesn't respect state borders
We built a similarity graph of all 150 districts: two districts are connected if their pollution-health profiles are statistically similar. Community detection on this graph finds natural clusters — and the result cuts across state boundaries in a way that directly challenges how pollution policy is currently organised.
Moran's I of 0.669 for PM2.5 — and 0.836 for respiratory rate — says that a district's pollution and disease burden are both strongly predicted by its neighbours', not by its own state's policies. Western UP behaves like Punjab. Coastal Tamil Nadu and Kerala move together. The Bihar-Jharkhand border districts share more with each other than with their respective capitals. And yet pollution control in India is administered almost entirely at the state level. The Indo-Gangetic airshed does not recognise the UP–Haryana border. The policy framework does.
The knowledge graph encodes these relationships as structured triples: District A shares-pollution-profile-with District B, District X has-spatial-proximity-to District Y, and so on across five relationship types. The practical use: when an intervention succeeds in one district, the graph tells you precisely which other districts are most structurally similar — and therefore most likely to benefit from the same approach without additional pilot testing.
How pollution causes disease — mediation analysis
Establishing that PM2.5 causes respiratory disease is the first step. Understanding how it does so matters for designing targeted interventions. Mediation analysis decomposes the total PM2.5 effect into two paths: a direct path (PM2.5 → disease, through inhalation and direct biological damage) and an indirect path (PM2.5 → intermediate variable → disease).
97% of the PM2.5 effect is direct. The indirect path runs through urbanisation(urban_percentage) — and it is slightly suppressive (−0.031, 95% CI: −0.035 to −0.027). The suppression mechanism is counterintuitive but well-defined: high-PM2.5 areas tend to be more urban (path a > 0), and more urban areas have better healthcare access and therefore lower observed respiratory case rates (path b < 0). Urbanisation is acting as a partial buffer — it partially absorbs the damage in high-pollution districts by providing more hospital beds and earlier treatment. This is not protection from exposure; it is better management of its consequences. The underlying cellular damage from fine-particle inhalation is not reduced.
What this means for intervention: the 3% suppression from urbanisation is not a reason for confidence — it shows that rural high-pollution districts, which lack that buffer, absorb the full unmitigated biological effect. Source-specific PM2.5 reductions produce health gains directly regardless of urbanisation level. The most rural, most polluted districts will see the largest absolute gains from emissions reductions precisely because they have the least healthcare infrastructure to offset the exposure.
Three independent methods, one direction
Correlation plus temporal precedence plus a monotonic dose-response is strong circumstantial evidence of causation. But fully establishing it requires a quasi-experimental design that breaks the symmetry more forcefully. We ran three, each with different assumptions and different failure modes.
Method 1: Synthetic Control — "What would Delhi look like with cleaner air?"
We built a "synthetic Delhi" — a weighted combination of other districts that matched Delhi's historical respiratory trajectory before the analysis period on everything except pollution level. The gap between real Delhi and its counterfactual is +76.2 respiratory cases per 100,000 people, or63% above the counterfactual mean. This is our best estimate of the respiratory burden Delhi carries specifically because of its pollution — over and above what a less-polluted version of Delhi would experience.
Method 2: Propensity Score Matching — a clean comparison
We split districts into treated (PM2.5 above 51 µg/m³) and control, then matched each treated district to the most similar control on urbanisation, literacy, population, and seasonality. The balance chart shows the standardised mean difference (SMD) dropping from 0.47 to 0.13 after matching — well below the conventional 0.2 adequacy threshold. The matched estimate: +41.5 extra cases per 100k in the high-pollution group (95% CI: [40.3, 42.7]).
Method 3: Regression Discontinuity — is there a safe threshold?
If NAAQS were a genuine causal threshold, you would see a visible jump in disease rates exactly at the 60 µg/m³ cutoff — as if crossing that line switches a mechanism on. There is no such jump. The scatter is smooth through the threshold with a small, borderline discontinuity. This null result is the most policy-relevant finding in the entire analysis: there is no safe PM2.5 level, even below the NAAQS standard. Risk rises continuously with exposure from the lowest concentrations we observe. Compliance with the current standard is not the finish line — it is a waypoint on a curve that doesn't flatten.
Synthetic control: +76 cases/100k for Delhi. PSM: +41.5 cases/100k across all high-pollution districts. RDD: no safe floor, anywhere. Three methods, different assumptions, one direction.
The geography of harm — same pollution, different damage
Standard regression assumes the PM2.5 effect is uniform across space. Geographically Weighted Regression (GWR) relaxes that by fitting separate local coefficients for each spatial zone, revealing where a unit increase in pollution produces the most health damage.
The reason is structural: East India has lower baseline healthcare coverage, more agricultural workers with extended outdoor exposure, and weaker hospital referral networks. When pollution hits, there is far less capacity to absorb the consequences.
This heterogeneity has a direct implication for standard-setting. A uniform national PM2.5 standard produces structurally unequal health outcomes when local health-system capacity varies this much. Pollution thresholds calibrated to health outcomes rather than emission feasibility would be substantially lower in East India than in Delhi — not because Bihar is biologically different, but because the system that manages the damage is far weaker.
The findings we are confident in
- 1The pollution-disease correlation survives every robustness check.r ≈ 0.38 at district-month level; r ≈ 0.79 at district means after controlling confounders. Robust to outlier removal, non-parametric tests, 15-state ANOVA, and partial-correlation conditioning on urbanisation, literacy, and population.
- 2Temporal precedence established: pollution leads disease in every district.Cross-correlation peaks at positive lag. Granger causality: 150/150 districts reject the null at p < 0.05. The reverse direction is substantially weaker across all tests.
- 3Dose-response is monotonic, log-linear, and has no safe floor.Cases rise from 22 to 129 per 100k across PM2.5 bins from 17 to 177 µg/m³ without a single reversal. RDD finds no discontinuity at the NAAQS threshold.
- 4Relative risk is 2.26× at NAAQS; PAF is 33%. At WHO levels, PAF is 75%.A third of respiratory cases are attributable to above-NAAQS exposure. Against the WHO guideline — which 99% of the panel fails — three-quarters would be attributable.
- 5The causal effect is 97% direct; the 3% suppression is through urbanisation.Mediation analysis: the indirect path through urbanisation is −0.031 (95% CI: −0.035 to −0.027), statistically significant and slightly suppressive. More urban areas buffer the damage through better healthcare access — rural high-pollution districts absorb the full biological effect.
- 6Within-district panel FE confirms the relationship over time.PM2.5 coefficient 0.655, p < 0.0001, within-R² = 0.50. The association holds even when each district is compared only to its own past.
- 7Socioeconomic conditions multiply harm: UP/Bihar has lower PM2.5 but worse outcomes.K-Means Cluster 3 has ~70 µg/m³ PM2.5 but 2.7× Delhi's respiratory rate. GWR confirms: East India PM2.5 coefficient (59.1) is ~3× the South (19.6), exposing a threefold spatial heterogeneity driven by healthcare infrastructure gaps, not pollution level.
- 8Causal estimates converge on 41–76 extra cases/100k from pollution exposure.Synthetic control: +76.2 cases/100k for New Delhi (63% above counterfactual mean). PSM: +41.5 cases/100k, 95% CI [40.3, 42.7], after achieving covariate balance.
- 9Six years of policy and national PM2.5 hasn't moved.2018: ~59 µg/m³. 2023: ~59 µg/m³. BS-VI, GRAP, odd-even, crop-burning bans — none of it shifted the national average. The levers being pulled are not the ones that matter.
- 10Pollution travels in airsheds; policy is organised around state lines.PM2.5 Moran's I = 0.669; Resp Rate Moran's I = 0.836. 6 of 6 Louvain communities span multiple states. The administrative and the physical unit of analysis are structurally mismatched.
Cluster 3: the same air causes more harm
Every analysis so far has treated India as a single system with one PM2.5 → respiratory relationship. But the cluster scatter — where Cluster 3 districts sit well above the regression line at moderate pollution levels — suggests that assumption is wrong. The question is not just how much pollution exists, but how much harm each unit of pollution actually causes in a given place.
To test this, we fit a national OLS model predicting respiratory cases from PM2.5 and log-population, then measure each district's residual: cases above or below what the model expects given its pollution and size. Stratifying by cluster produces the clearest finding in the entire analysis: Cluster 3 is not just unlucky — its excess is structural. At the same PM2.5 concentration and the same population, Cluster 3 districts generate far more cases than any other cluster. The interaction regression (with cluster × PM2.5 interaction terms) pinpoints why: the slope doesn't differ — the level does.
The policy implication is direct. A uniform NAAQS standard of 40 µg/m³ applied equally across all districts implicitly assumes each district experiences the same health outcome per unit of pollution. The sensitivity analysis refutes that assumption. Cluster 3 districts need a lower effective threshold — not because their pollution is highest (it isn't), but because their communities absorb more harm per µg/m³ than anywhere else in the panel. A standard calibrated to Delhi's pollution-disease relationship is structurally inadequate for Lucknow, Patna, or Varanasi.
Policy recommendations
The analysis points clearly at what the data would prefer. We frame these by horizon: some are actionable this winter, others are multi-year structural investments.
- Pre-position health surge capacity in Cluster 1 and 3 districts before October.The CUSUM changepoint fires every October. Health response can be planned three months in advance, not announced reactively after PM2.5 is already hazardous.
- Real-time district-level AQI feeds directly to health officers.The monitoring data exists. The operational link to healthcare planning does not. This is an administrative gap, not a technical one.
- Hospital capacity plans in UP/Bihar for November–February.These districts already show the highest respiratory case rates. The winter surge is predictable. Surge capacity should be pre-planned, not improvised.
- Extend GRAP-equivalent frameworks to all districts where PM2.5 routinely exceeds 60 µg/m³.Multiple Cluster 1 districts outside Delhi have equivalent pollution but no staged-response framework. The concentration of attention on Delhi leaves most of the exposed population uncovered.
- Scale crop-residue management in Punjab, Haryana, and western UP.The October changepoint is partly stubble-driven. The technology for in-situ residue management exists; deployment at scale is the bottleneck, not the science.
- Move HMIS reporting to weekly in high-risk districts.Monthly granularity hides surge events. Weekly reporting enables dynamic resource reallocation before crises develop.
- Calibrate regional pollution standards to health-system capacity, not just emission feasibility.GWR shows East India absorbs 2× the health damage per µg/m³. A uniform national standard produces structurally unequal outcomes.
- Fund primary care infrastructure in Cluster 3 districts.Cutting pollution alone will not close the outcome gap between Delhi and UP/Bihar. The compounding factor is healthcare access. Without parallel health investment, the most-burdened districts remain most-burdened.
- Restructure pollution governance around airshed boundaries, not state lines.An Indo-Gangetic airshed authority with cross-state regulatory mandate would match the policy unit to the physical system. State-level management of an inter-state atmospheric problem cannot, by structure, work.
- Expand CPCB monitoring into network-identified coverage gaps.The link-prediction model flags which unmonitored districts are most likely to have high pollution based on neighbourhood structure. New stations should go where the model says signal is probably missing.
- Institutionalise annual CPCB × HMIS × Census joint reviews with public reporting.We assembled this dataset as a research project. It should be a standing government function, with published results and policy targets measured against outcomes annually.
What this whole project actually says
Six years. 150 districts. 250,008 air-quality readings. 7,817 real HMIS district-year records (10,800 calibrated monthly panel observations). Ten statistical tests, four causal methods, two independent data systems. The answer to the original question is consistent throughout: ambient air pollution is a significant, quantifiable, and apparently stagnant driver of respiratory disease in India. It is not the only driver — socioeconomic vulnerability amplifies it and healthcare access determines how much damage becomes permanent — but the pollution signal is independently real, robust to every test, and uniform across every region and demographic in the panel.
The national PM2.5 average hasn't moved in six years. The dose-response curve has no safe floor. 150 out of 150 districts pass Granger causality. The relative risk against NAAQS is 2.26×. If India met the WHO guideline, three-quarters of the respiratory burden in the panel would theoretically not exist. These are not modelling artefacts — they're what the raw data says before any assumptions are imposed.
Air pollution in India is not an environmental problem with a health side-effect. It is a public health emergency administered by environmental institutions, with the most-affected populations the least protected by either system.
The response has to be airshed-shaped, season-aware, and twinned with healthcare investment in the most vulnerable districts. Anything narrower — a vehicle-only scheme, a Delhi-centric framework, a clean-air plan that doesn't fund East India hospitals — will leave most of the harm intact while generating the appearance of action.
Analysis based on 2018–2023 CPCB/NDAP air-quality and HMIS health data · 150 districts · 15 states · 250,008 AQ records · 7,817 HMIS district-year records · all figures computed from primary data