How to Calculate Sample Size for a Thesis: G*Power vs the Krejcie-Morgan Table vs Cochran (2026)

Most sample-size arguments fail not because the arithmetic is wrong but because the scholar used a method built for a different question. The comparison below sorts that out first, then walks each method, then gives you the paragraph to put in your methodology chapter.

Method Answers the question Needs you to know Best for Where it falls short
G*Power How many participants do I need to detect an effect of a given size? The test you will run, expected effect size, alpha, desired power Any thesis whose main claim rests on a hypothesis test — experiments, group comparisons, regression, correlation Requires an effect size estimate you must justify from prior literature or a pilot
Krejcie-Morgan table How many respondents represent a known, finite population? The population size N Surveys of a bounded population: employees of one firm, students of one university, members of one association Fixed at 95% confidence, 5% margin and maximum variability — it cannot be tuned, and it says nothing about power
Cochran’s formula How many respondents represent a large or unknown population? Confidence level, margin of error, expected proportion Surveys where the population is very large, unlisted or effectively infinite Same limitation as above — it is a precision calculation, not a power calculation

The recommendation, stated plainly: if your thesis tests hypotheses, use G*Power. Use Krejcie-Morgan or Cochran only when your study is genuinely descriptive — estimating a proportion or a mean in a population — or as a secondary check alongside a power analysis. The most common error in Indian management, education and commerce theses is a Krejcie-Morgan number attached to a study that then runs regression and ANOVA, which the table was never designed to size.

1. G*Power — the power analysis

G*Power is a free program from the Allgemeine Psychologie und Arbeitspsychologie group at Heinrich-Heine-Universität Düsseldorf. Its own terms of use state that it is free for everyone, including people working in commercial environments, with only commercial redistribution prohibited. It computes power analyses for t tests, F tests, χ² tests, z tests and some exact tests, and it will also compute effect sizes and draw power plots.

The calculation you want is an a priori analysis: you supply the test family, the effect size you expect, your alpha level (conventionally .05) and your desired power (conventionally .80, sometimes .90), and it returns the required N. Four inputs, one output, and the only one that requires judgement is the effect size.

That is also where scholars get stuck, so here is the order of preference for justifying it. Best: an effect size reported in a comparable published study in your area — this is the single strongest justification and it is one more reason to build a proper synthesis matrix while reading, as described in our guide to writing a literature review for a PhD thesis in India. Second best: an effect size from your own pilot study. Acceptable, with a stated rationale: a conventional benchmark for a small, medium or large effect in your field, chosen because you are powering the study to detect the smallest effect that would be practically meaningful. Not acceptable: choosing the effect size that produces the N you can afford.

Abstract illustration of a power curve rising with sample size and flattening
Power rises steeply and then flattens: past a point, more participants buy very little.

When you cite G*Power in your thesis, the program’s authors ask you to reference Faul, Erdfelder, Lang and Buchner (2007) in Behavior Research Methods, volume 39, pages 175-191, and for correlation and regression analyses, Faul, Erdfelder, Buchner and Lang (2009) in the same journal, volume 41, pages 1149-1160. Citing the software without citing the papers is a small omission that a methodologically alert examiner will notice.

2. The Krejcie-Morgan table — the finite-population shortcut

This is the table taped to the wall of every research methodology classroom in the country, and it is genuinely useful when applied to the right design. It answers one question: given a population of N, how many respondents do I need to estimate a proportion at 95% confidence with a 5% margin of error, assuming maximum variability?

It is worth knowing that it is not a lookup table handed down from nowhere — it is the output of a formula you can compute yourself:

s = χ²NP(1−P) / [ d²(N−1) + χ²P(1−P) ]

where χ² = 3.841 (the chi-square value for one degree of freedom at the 95% confidence level), P = 0.50 (the population proportion, set at 0.5 because that maximises the required sample), and d = 0.05 (the margin of error). Computing it directly for a range of population sizes reproduces the published table exactly:

Population (N) Required sample (s)
100 80
200 132
500 217
1,000 278
2,000 322
5,000 357
10,000 370
100,000 383
1,000,000 384

Two things become visible the moment you see the numbers in a column rather than as a single lookup. First, the curve flattens hard: going from a population of 10,000 to a population of a million adds fourteen respondents. Second, everything above roughly 100,000 converges on 384, which is why that number appears in so many dissertations — it is simply the large-population limit of this formula at these settings.

That also exposes the limitation. Those settings are baked in. You cannot use the table to work at 99% confidence, or with a 3% margin, or with a known population proportion far from 0.5, without going back to the formula. And it tells you nothing at all about whether your sample can detect the relationship you intend to test.

3. Cochran’s formula — when the population is unknown

When your population is very large, unlisted or genuinely unbounded — “consumers in Tamil Nadu”, “users of digital payment apps” — you cannot look up N because you do not have one. Cochran’s formula handles that case, computing the required sample from your chosen confidence level, margin of error and expected proportion, and it converges on the same familiar territory: at 95% confidence, 5% margin and maximum variability, it returns approximately 385.

Where a finite population estimate does exist, Cochran also provides a finite population correction that reduces the requirement, which is worth applying when your population is small enough for it to matter — it can meaningfully reduce the burden on a study of a single institution.

4. Whichever method you use, plan for attrition

The calculators return the number of usable responses you need, not the number of questionnaires to distribute. Between distribution and analysis you will lose responses to non-response, incomplete forms, straight-lining, failed attention checks and cases excluded during data cleaning.

Abstract illustration of a planned sample with an added buffer for non-response
The calculator gives you the usable N; the distribution plan has to add the buffer.

Build a buffer explicitly and state the assumed response rate in your methodology chapter. Basing that rate on something real — a comparable published study in your setting, or your own pilot — is far more defensible than a round guess, and it turns an awkward gap into an evidenced planning decision. Then report the actual response rate in your results chapter, where it belongs among the data-preparation figures described in our guide to writing the results chapter.

5. What if your design is qualitative?

None of the above applies. Statistical power and margin of error are properties of numerical estimation and have no meaning for interview or case-study work; importing a Krejcie-Morgan number into a qualitative chapter is a category error that examiners flag immediately. Qualitative adequacy is argued through sampling strategy and saturation instead, which is covered in how many interviews are enough for a qualitative thesis. For a mixed-methods design, size the two strands separately and say so.

6. The justification paragraph to put in your methodology chapter

A complete justification names the method, the inputs, the output and the buffer, in four sentences. For a hypothesis-testing design:

“The required sample size was determined through an a priori power analysis conducted in G*Power (Faul et al., 2007). Assuming a medium effect size, alpha of .05 and power of .80 for a multiple regression with four predictors, the analysis indicated a minimum of 85 participants. The effect size was based on the values reported by [author, year] in a comparable study of [setting]. Anticipating a response rate of approximately 60% on the basis of the pilot, 150 questionnaires were distributed, yielding 97 usable responses.”

For a descriptive survey of a bounded population: “The target population comprised 1,240 registered members. Applying the Krejcie and Morgan formula at 95% confidence with a 5% margin of error and maximum variability (P = .5), the required sample was 294. Allowing for non-response, 420 questionnaires were distributed, of which 311 usable responses were returned.”

Both paragraphs share the property that matters: every number in them can be traced to a decision the reader can evaluate.

Getting the justification into the chapter while it is still fresh

Sample-size reasoning is decided in month two and written up in month fourteen, which is exactly why so many methodology chapters contain a bare number with no argument behind it. Capture the inputs, the source of the effect size and the response-rate assumption at the moment you make them, and the paragraph above writes itself later.

Tesify keeps your chapters structured and every citation attached to a real source as you draft, so the study you designed is the study your methodology chapter describes. The design decisions stay yours.

Start writing your methodology chapter in Tesify

Frequently asked questions

Which sample size method should I use for my thesis?

G*Power if your study tests hypotheses; Krejcie-Morgan or Cochran if it is descriptive estimation from a population. Using a precision formula to size a hypothesis-testing study is the most common mismatch.

Is G*Power free?

Yes. Its terms of use state it is free for everyone, including commercial environments, with only commercial redistribution prohibited.

Why does 384 appear in so many theses?

Because it is the large-population limit of the Krejcie-Morgan formula at 95% confidence, a 5% margin and maximum variability. Above a population of roughly 100,000 the required sample barely moves.

What effect size should I assume if there is no prior study?

Use the smallest effect that would be practically meaningful in your field and say so explicitly, or run a pilot. Never reverse-engineer the effect size from the sample you can afford.

What power level should I aim for?

.80 is the widely used convention and .90 is common where the cost of missing an effect is high. State the level you chose and treat it as a design decision, not a default.

Does the Krejcie-Morgan table account for statistical power?

No. It is a precision calculation for estimating a proportion. A sample adequate for estimating a proportion may be badly underpowered for detecting a small interaction effect.

How much buffer should I add for non-response?

Enough to reach your usable N at a response rate you can defend, ideally taken from a comparable study or your own pilot rather than assumed. State the assumption in the methodology chapter and the actual rate in the results.

Can I reduce my sample size if my population is small?

Yes — that is precisely what the finite population correction does, and Krejcie-Morgan already builds it in. A study of one department with 60 eligible staff does not need 384 respondents.

What if I cannot reach the calculated sample size?

Report the shortfall, report the power your achieved sample provides, and treat it as a stated limitation. An honest underpowered study is defensible; a concealed one is not.

Do I need a sample size calculation for a pilot study?

Not a power calculation. Pilots are sized for practical purposes such as checking item comprehension and estimating reliability, and their role is to inform the main study’s calculation.

Should the sample size go in the proposal or only the thesis?

Both, with the same reasoning. Committees examine the calculation at the proposal stage precisely because it is expensive to fix after data collection has begun.

Where can I download G*Power and read the citations?

From the Allgemeine Psychologie und Arbeitspsychologie pages at Heinrich-Heine-Universität Düsseldorf, which also host the manual and name the two Faul et al. papers the authors ask you to cite.