Where to Find Data for Your Indian Thesis: Official Sources by Discipline

Most stalled Indian doctorates stall at data access, not at ideas. The sources below were all reachable on 12 August 2026 and cover the disciplines where Indian scholars most often need secondary data. Confirm your data exists and covers your period before you commit to a topic, not after your synopsis is approved.

Start here regardless of discipline

Source Holds Access
Open Government Data Platform India (data.gov.in) Datasets published by ministries and departments, searchable by catalogue Open, no account needed
Ministry of Statistics and Programme Implementation (mospi.gov.in) National statistical releases and survey reports Open
Census of India (censusindia.gov.in) Population, household and demographic data down to small administrative units Open
Shodhganga (shodhganga.inflibnet.ac.in) Completed Indian theses Open

Begin at data.gov.in even when you already know which ministry you need. The catalogue surfaces datasets from technical departments you would not think to check, and it is faster than navigating individual ministry sites one at a time.

A practical point that saves weeks: search Shodhganga for theses close to your topic and read their methodology chapters specifically to see where they obtained data. A completed Indian thesis is a proof of access, showing you a route someone actually got through rather than one that merely ought to work.

Commerce, management, economics and finance

  • Reserve Bank of India (rbi.org.in). Monetary data, banking statistics, interest and exchange rates, and regular statistical publications on the Indian economy.
  • MoSPI. National accounts, price indices, industrial production and household survey results.
  • Company annual reports and stock exchange filings. The standard source for firm-level studies of financial performance, disclosure and governance.

This discipline moves fastest precisely because so much is public. If you are early in topic selection and worried about feasibility, a design built on published financial disclosures removes the permission problem entirely.

When you use listed-company data, define your sample selection explicitly: sector, period, and the requirement that a company filed complete reports throughout. Those criteria are what you write up as purposive sampling in your methodology chapter.

Shelves of labelled statistical volumes in an archive
Confirm the dataset covers your years, your geography and your variables before committing to the topic.

Law

  • India Code (indiacode.nic.in). The repository of central legislation, useful for confirming the current text of an Act and tracing amendments.
  • Court judgment portals. The Supreme Court and the High Courts publish judgments through their official websites, and these are the primary material for doctrinal work.
  • Ministry and regulator websites. Rules, notifications and consultation papers that sit beneath primary legislation.

For doctrinal research these materials are your data, which makes law one of the most feasible disciplines for a scholar without institutional access to proprietary databases. Two cautions apply. Always verify whether a provision is still in force and whether it has been amended, because citing a repealed provision is immediately visible to an examiner. And when you cite a judgment, cite it from the court’s own record rather than from a secondary summary.

Education

  • University Grants Commission (ugc.gov.in). Regulations, public notices and policy documents governing higher education.
  • The government’s all-India higher education survey releases. The authoritative national figures on institutions, enrolment and teaching staff.
  • Individual university annual reports. Far more detailed than national aggregates when your study focuses on one institution.

If your study involves schools, note that institutional permission is the gating item and is routinely underestimated. Begin that process before your synopsis is finalised rather than after.

Health and social sciences

  • Census of India. Demographic and household characteristics at fine geographic granularity, valuable as denominators and control variables.
  • MoSPI survey reports. Health, consumption, employment and living-condition indicators from national surveys.
  • data.gov.in. Health and social sector datasets published by the relevant ministries.

Clinical records held by a hospital are a different category entirely. They are not open data, and using them requires institutional permission and ethics clearance, which cannot be obtained retrospectively. Establish that route early or choose a design that does not depend on it.

Engineering, environment and agriculture

  • data.gov.in. Datasets on infrastructure, energy, water, transport and agricultural production published by technical ministries.
  • Departmental monitoring data. Environmental and meteorological agencies publish measurement series that support studies of quality, load and variability over time.
  • Institutional laboratory records, where your own department holds monitoring or testing data. Confirm access in writing before building a design around it.

Five questions to ask of any dataset

  1. Who collected it, and is the method published? An agency that documents its methodology can be evaluated; one that does not, cannot.
  2. What is the reference period, as opposed to the publication date? These differ frequently, and only the reference period belongs in your analysis.
  3. What is the unit and geographic level? It must match your level of analysis. Mixing state-level and district-level figures in one analysis is a serious error.
  4. Have the definitions changed across the series? This is the trap that catches most time-series work. A redefined indicator breaks comparability, and the change is usually disclosed only in a footnote.
  5. Is the coverage complete for your period and region? Gaps are common and must be reported rather than quietly interpolated.

How to cite a dataset

Include the organisation, the dataset or report title, the reference period of the data, the source, and the date you accessed it. A pattern you can adapt for your methodology chapter:

“This study uses secondary data on [indicator] for the period 2019 to 2025, published by [organisation] and obtained from [source], accessed on 12 August 2026. Units with incomplete records across the study period were excluded, leaving a final sample of [n].”

Save a dated copy of every file you download. Portals revise datasets, and your saved copy is the evidence that the figures in your thesis were correct at the time of collection.

What to do when the data does not exist

Three legitimate routes, all of which must be explained rather than concealed:

  • Request it formally. Public bodies have mechanisms for information requests. Start early, because responses take time and may be refused.
  • Substitute a proxy variable that is available, and justify the substitution theoretically rather than by convenience.
  • Change the scope to a region, sector or period where data does exist. This is far better than persisting with a design that cannot be executed.

Discovering that data is unavailable at the topic stage is a minor inconvenience. Discovering it in your third year is a crisis, which is why access belongs in the feasibility test set out in how to choose a PhD research topic in India by discipline.

Related reading

For which national figures could and could not be verified at source, along with how to cite online statistics, see Indian research and higher education data 2026. When you reach the publication stage, the current position on journal verification is covered in how to verify a journal in India in 2026, which matters because the UGC-CARE list is no longer being updated.

Turning a dataset into a defensible chapter

Secondary-data research has a distinctive writing problem. You must justify the source, the selection criteria, the treatment of missing values and the limitations, without the instrument and reliability sections that anchor questionnaire-based methodology chapters. Most published examples scholars encounter are survey-based, which leaves this route underdocumented.

Tesify helps you build a methodology chapter matched to the kind of data you are actually using, keeping your source description, sample criteria and analysis plan consistent as the chapter develops. You choose and verify every source, and you remain responsible for every figure.

Build your methodology chapter in Tesify

Frequently asked questions

Is secondary-data research considered weaker for a PhD?

No. Many strong doctoral theses use exclusively secondary data. What matters is the analytical contribution, not who collected the raw numbers. The demands simply shift from fieldwork to justification and careful treatment of limitations.

Do I need permission to use open government data?

Generally no, since it is published for public use, but attribution is required and individual datasets may carry specific terms. Check the licence note attached to anything you build analysis on.

How many years of data do I need?

It depends on the technique. Panel analysis commonly uses at least three to five years so there is meaningful variation over time. Consistency of definitions across those years matters more than the count.

Can I combine data from two different sources?

Only if the units, definitions and reference periods align. Where they do not, combining them produces invalid conclusions. If you do merge sources, describe the matching procedure and any adjustments explicitly.

What if two official sources disagree?

This happens because coverage and methods differ. Choose one as primary, apply it consistently, and explain the discrepancy in a note rather than switching between them.

Are paid commercial databases worth it?

Check first whether your university library already subscribes, since many scholars pay for access they already have. Where a commercial database is genuinely required, factor the cost and the access period into your timeline.

How do I handle missing values?

Decide and state your rule before analysis: exclude affected units, or apply a documented technique. Report how many observations were affected. Silently filling gaps is a serious integrity problem.

Can I use data from a previous student’s thesis?

Not their raw dataset without explicit permission and acknowledgement. You can and should cite their published findings as prior research, and their methodology chapter is a legitimate guide to where they obtained data.

Does using public data require ethics clearance?

Usually not where the data is aggregate and contains no identifiable individuals. Policies vary by institution, so confirm with your ethics committee rather than assuming, particularly if any dataset contains unit-level records.

How current must my data be?

Use the most recent official release and state its reference period. Official statistics are published with a lag, which is entirely acceptable provided you are transparent about the period rather than implying currency you do not have.