| Your data situation | Test/technique | Why |
|---|---|---|
| Secondary analysis of NFHS-5, DHS-style or other complex-sample national survey data | Survey-weighted logistic or linear regression, with cluster and strata terms specified | These surveys use multi-stage cluster sampling with sampling weights; ignoring the design produces standard errors that are wrong, usually too small, which overstates significance |
| Own-collected categorical data, two groups, association between two categorical variables | Chi-square test of independence (or Fisher’s exact test if any expected cell count is below 5) | Standard test for association between two categorical variables in simple random or convenience-sampled data |
| Own-collected data, comparing a continuous outcome across two groups | Independent-samples t-test (or Mann-Whitney U if the outcome is not normally distributed) | Standard comparison of means for two independent groups |
| Own-collected data, predicting a binary health outcome (disease present/absent) from several predictors | Binary logistic regression | The standard technique for modelling a binary outcome against multiple predictors, giving odds ratios that are directly interpretable in a public health context |
| HMIS or facility-level administrative data, comparing indicators across facilities or time periods | Time-series or panel methods if repeated over time; ANOVA or Kruskal-Wallis if comparing several facility groups at one time point | Facility-level administrative data is typically aggregated, not individual-level, and needs a technique matched to that structure rather than an individual-level test |
The Decision That Comes Before the Test: What Kind of Data Do You Actually Have?
A public health dissertation in India usually draws on one of three very different data structures, and the test-selection question only makes sense once you know which one you are working with.
- A complex-sample national survey (NFHS-5, DHS-style, NSS/PLFS or a comparable large government survey), where you are the secondary analyst of data someone else collected with a multi-stage cluster design and sampling weights.
- Your own primary data collected via a questionnaire or clinical assessment from a sample you recruited yourself, typically through simple random or convenience sampling.
- Facility or programme administrative data (HMIS-type records, hospital registers), which is usually aggregated at the facility or period level rather than individual-level.
Each of these needs a genuinely different statistical approach, and the single most common Chapter 3 error in Indian public health dissertations is applying a simple-random-sample technique (an ordinary chi-square or unweighted regression) to data from the first category, where it produces a technically wrong answer that happens to look plausible.

Why Complex-Sample Survey Data Cannot Be Analysed Like a Simple Random Sample
NFHS-5 and comparably designed national surveys use a multi-stage stratified cluster sampling design: the country is divided into strata, primary sampling units (typically villages or urban blocks) are selected within strata, and households or individuals are selected within those units. This design is efficient for national data collection, but it means observations within the same cluster are more similar to each other than observations drawn at random from the whole population — a statistical property called intra-cluster correlation. An ordinary chi-square test or an unweighted regression assumes every observation is independent, which understates the true standard error when clusters are ignored, making results look more statistically significant than the data actually supports. The fix is to use survey-weighted analysis (in Stata, the svy command family; in R, the survey package; in SPSS, the Complex Samples module) that explicitly accounts for the sampling weight, the strata and the primary sampling unit — all of which are provided as variables in the public-use NFHS-5 dataset alongside the substantive variables.
A Worked Example: Choosing the Test for an NFHS-5 Secondary Analysis
Suppose your objective is to examine the association between household wealth quintile and full childhood immunisation coverage, using NFHS-5 data.
- Identify the outcome type. Full immunisation status is binary (yes/no) — this points toward logistic regression, not linear regression.
- Identify the sampling design variables in the dataset. NFHS-5’s public-use files include the primary sampling unit, stratum and sampling weight variables needed to declare the survey design in your statistical software before running any test.
- Declare the survey design first, every time. Before running any test on NFHS-5 data, declare the design (weight, strata, PSU) in your software — every subsequent test then automatically accounts for it, rather than needing to be adjusted individually.
- Run survey-weighted logistic regression, with wealth quintile as the key predictor and other objective-relevant covariates (maternal education, place of residence, birth order) included as controls, since a bivariate association alone will not address confounding.
- Report design-adjusted confidence intervals, not just a p-value — a design-adjusted 95% confidence interval on the odds ratio is what a public-health-literate examiner expects to see, and its absence is a common Chapter 4 gap.
Ranked Guidance by Data Source
- NFHS-5 or DHS-style secondary data — survey-weighted methods, no exception. This is not optional rigour; an unweighted analysis of complex-sample survey data is a methodological error an examiner familiar with these datasets will identify immediately. Where to access NFHS-5 and its companion surveys is covered in our comparison of data sources for a public health dissertation in India.
- Your own primary survey data — standard inferential statistics fit, provided your sampling was genuinely random or you disclose that it was not. Chi-square, t-tests, ANOVA and ordinary regression are appropriate here; the caveat is disclosing honestly if your sample was a convenience sample rather than a probability sample, since this affects how your findings should be generalised, not which test you technically run. How large that sample needs to be is a separate question, worked through in our guide to calculating sample size for a thesis.
- HMIS or facility administrative data — match the technique to the level of aggregation. If your unit of analysis is the facility or the month, not the individual patient, treat it as aggregate or time-series data rather than forcing an individual-level test onto facility-level numbers.

Common Mistakes in a Public Health Statistics Chapter
- Running an unweighted chi-square or regression on NFHS-5 or similar complex-sample data — the single most common and most serious error, because it silently overstates statistical significance.
- Reporting a p-value with no effect size or confidence interval. A public-health-literate examiner wants the odds ratio or prevalence ratio with its confidence interval, not only whether p is below 0.05.
- Treating an association from cross-sectional survey data as causal language in the discussion chapter — NFHS-5 and most Indian public health surveys are cross-sectional, so “associated with” is defensible; “causes” or “leads to” is not, without a design that supports causal inference.
- Applying individual-level statistical tests to facility-aggregated HMIS data without first checking whether the unit of analysis matches the technique.
Adjusting for Confounders: Why a Bivariate Result Is Rarely the Whole Story
A bivariate chi-square or an unadjusted odds ratio answers a narrower question than most public health dissertations actually need to answer. If wealth quintile is associated with immunisation coverage in a simple two-way table, that association could partly or wholly reflect a third variable — maternal education, place of residence, or access to a health facility — that correlates with both wealth and immunisation behaviour. Multivariable regression (logistic, for a binary outcome; linear, for a continuous one) lets you report an adjusted estimate that accounts for the confounders your literature review identifies as plausible, and the choice of which variables to adjust for should be justified from that literature, not selected by which ones happen to be statistically significant in your own data. A results chapter that reports only unadjusted bivariate associations, with no multivariable model, is one of the more common reasons a public health methodology chapter reads as incomplete rather than genuinely wrong.
Reporting the Sample Size Behind a Complex-Sample Estimate
A subtlety that surprises many first-time NFHS-5 analysts: the unweighted sample size (the actual number of respondents in your subset) and the weighted, population-representative estimate are two different numbers serving two different purposes. Report the unweighted sample size in your methods section, since that is what determines your statistical power and what a reader needs to judge precision; report the weighted percentage or prevalence in your results, since that is what is actually representative of the national or state population the survey was designed to describe. Presenting an unweighted percentage as though it were nationally representative, or an unweighted sample size as though it reflected the full survey’s scale, are both errors that a reviewer familiar with these datasets will catch immediately.
When a Non-Parametric Alternative Is the Better Choice
For own-collected data, the standard t-test and ANOVA both assume the outcome variable is approximately normally distributed within each group, an assumption public health data — skewed toward zero for many exposure or symptom-count variables — often fails to meet. Before defaulting to a parametric test, check the distribution of your outcome variable (a histogram, plus a formal normality test such as Shapiro-Wilk on samples where it is appropriate); where normality is clearly violated and your sample size is not large enough for the central limit theorem to rescue you, the non-parametric alternative — Mann-Whitney U in place of the independent t-test, Kruskal-Wallis in place of one-way ANOVA — tests a similar question on ranks rather than raw values and does not require the normality assumption. State in your methodology chapter which check you ran and why you chose the test you did; “the data were not normally distributed, so a Mann-Whitney U test was used” is a complete, defensible sentence that pre-empts the question an examiner would otherwise ask. If your regression model instead returns a significant overall fit with no individually significant predictors, that is a different, specific diagnosis, covered in our guide to handling multicollinearity with VIF.
Frequently Asked Questions
Do I need survey-weighted analysis if I am only using a small subset of NFHS-5 variables?
Yes — the sampling design applies to how the data was collected, not to how many variables you use from it. Any analysis of NFHS-5 or similarly designed survey data should account for the weight, strata and cluster variables regardless of how narrow your research question is.
What software actually handles complex-sample survey analysis?
Stata’s svy commands, R’s survey package, and SPSS’s Complex Samples module (a separate module from base SPSS, not always included in a standard licence) all handle this; confirm which your department’s software access actually includes before committing to one.
Can I use chi-square on my own primary-collected public health survey data?
Yes, provided your own data collection was not itself a complex multi-stage design — most student-collected primary surveys use simple random or convenience sampling, where standard chi-square and regression techniques are appropriate without a design adjustment.
How do I know if my dataset needs survey weights?
Check the dataset’s own documentation for sampling-weight, strata and primary-sampling-unit variables — if the data custodian (such as NFHS/DHS) provides these variables explicitly, that is the signal that weighted analysis is expected, not optional.
Is a logistic regression always the right choice for a binary health outcome?
It is the standard choice for most binary public health outcomes, but a rare-outcome design or a case-control structure may call for a related but distinct technique — confirm your specific design with your guide or a biostatistics reference before defaulting to standard logistic regression.
Tesify helps you match the statistical technique to the specific data structure your public health dissertation actually uses — survey-weighted analysis for a national dataset, standard inferential tests for your own primary data — before you write a methodology chapter around the wrong assumption.
