Every one of the 30 M.Tech thesis topics below can be started today from a dataset that is free to download, and 21 of them run on data built in India: ten datasets from AI4Bharat, IIIT Hyderabad, IIT Bombay, ISRO, the NCRB and the Open Government Data platform, checked at source in August 2026, together with five international benchmarks and one classic dataset you should no longer plan around.
The commonest reason an M.Tech dissertation in computer science slips a semester is not the model but the data: a topic chosen in July turns out in October to need a dataset that is paid, gated behind a credentialing process, withdrawn by its host, or reachable only through a browser that blocks scripts. This article inverts the usual order. The datasets come first, with their size, custodian and access condition as published, and each topic is stated as a research question that the named dataset can answer inside a two-semester project. Confirm the scope with your guide; the dataset facts are the custodians’ own, with the year we read them.
The datasets, as published by their custodians
| Dataset | Custodian | Size as published | Access | Read |
|---|---|---|---|---|
| Samanantar | AI4Bharat, IIT Madras | Parallel sentences for 11 Indian languages with English; 10.1 million rows for Hindi, 8.6 million for Bengali, 5.92 million for Malayalam, 5.26 million for Tamil, 998,000 for Odia, 141,000 for Assamese | Free, Hugging Face | 2026 |
| IndicVoices | AI4Bharat | 23,700 hours of speech from 51,000 speakers across 400-plus districts and 22 languages; 8 per cent read, 76 per cent extempore, 15 per cent conversational; 11,200 hours transcribed as of December 2025 | Free, gated: accept conditions on Hugging Face | 2026 |
| Naamapadam | AI4Bharat | Named-entity data for 11 Indian languages; about 1 million rows for Hindi, 967,000 for Bengali, 10,400 for Assamese | Free, Hugging Face | 2026 |
| IIT Bombay English-Hindi Parallel Corpus | CFILT, IIT Bombay | Version 3.1; used at the Workshop on Asian Language Translation since 2016; mirrored on Hugging Face | Free | 2026 |
| IndicCorp v2 | AI4Bharat | Monolingual web text for Indian languages | Free, Hugging Face | 2026 |
| Indian Driving Dataset (IDD) | IIIT Hyderabad | 10,003 images, 34 classes, 182 drive sequences around Hyderabad and Bengaluru; 6,993 train, 981 validation, 2,029 test; mostly 1080p | Free for research, via the dataset site | 2026 |
| RoadSocial | CVIT, IIIT Hyderabad | Video question-answering benchmark for road events from social-media video, CVPR 2025 | Free | 2026 |
| Cityscapes | Cityscapes consortium | 5,000 finely and 20,000 coarsely annotated images, 30 classes, 50 cities | Free for research, registration | 2026 |
| CIC-IDS2017 | Canadian Institute for Cybersecurity, UNB | Benign and current attack traffic as PCAPs plus CICFlowMeter-labelled flow CSVs | Free | 2026 |
| UNSW-NB15 | UNSW Canberra | 100 GB of raw traffic generated with IXIA PerfectStorm, with CSV feature files | Free | 2026 |
| MIMIC-IV v3.1 | MIT and Beth Israel Deaconess, via PhysioNet | De-identified intensive-care records | Credentialed access: training and data-use agreement | 2026 |
| UCI Machine Learning Repository | UC Irvine | 689 datasets maintained | Free | 2026 |
| Open Government Data platform (data.gov.in) | Government of India | Counters on the portal show over 3.2 lakh resources from central ministries and 28,865 from states | Free, some via API key | 2026 |
| Crime in India | National Crime Records Bureau | Annual tables since the 1953 issue; latest issue for 2024 | Free PDFs and tables | 2026 |
| Bhuvan Open EO Data Archive | NRSC, ISRO | AWiFS and LISS-III imagery, digital elevation models and thematic layers | Free, registration | 2026 |
Sources: each custodian’s own dataset page, read in August 2026. Sizes are the custodian’s headline figures, not our counts.
Natural language processing for Indian languages: eight topics
- Low-resource translation with Samanantar. How much does translation quality for Assamese (141,000 pairs) improve when a model is pre-trained on the ten larger Samanantar languages and fine-tuned on Assamese alone?
- Domain shift in Hindi-English translation. How does a model trained on the IIT Bombay corpus degrade on government, medical and legal sentences drawn from IndicCorp v2, and which domain adaptation method recovers the most?
- Named-entity recognition transfer. Does a model trained on Naamapadam’s 1 million Hindi rows transfer to Odia (199,000 rows) better through script transliteration or through a shared multilingual encoder?
- Code-mixed Hinglish detection. Which token-level features separate code-mixed from monolingual sentences in IndicCorp v2, and how does a classifier’s error rate change with sentence length?
- Speech recognition across districts. Using IndicVoices’ district metadata, how much does word error rate vary across districts within one language, and does district-balanced fine-tuning reduce the spread?
- Extempore versus read speech. Since 76 per cent of IndicVoices is extempore, how much accuracy does a model trained only on read speech lose on spontaneous speech, and what proportion of extempore data closes the gap?
- Transliteration quality for names. How well do current transliteration tools handle the person and place names in Naamapadam, and which error classes dominate?
- Sentence alignment audit. What proportion of Samanantar pairs in one mid-size language are misaligned, and does filtering them change downstream translation quality more than adding data?

Computer vision on Indian roads: six topics
- Segmentation under unstructured traffic. How much does a segmentation model trained on Cityscapes’ 30 classes lose on IDD’s 34 classes, and which of IDD’s added classes cause the largest drop?
- Small-object detection. Across IDD’s 182 drive sequences, how does detection recall for two-wheelers and pedestrians change with distance from the camera, and does multi-scale training help?
- Night and monsoon robustness. Using synthetic rain and low-light augmentation on IDD’s daytime images, which augmentation policy best predicts performance on real adverse-weather frames?
- Domain adaptation from Europe to India. Which unsupervised domain adaptation method transfers most from Cityscapes to IDD when only unlabelled IDD images are used?
- Road-event question answering. On RoadSocial, which categories of road event are answered worst by a current video-language model, and does adding Indian-road pre-training change the ranking?
- Label-efficiency study. How does IDD validation accuracy scale as the training set grows from 500 to 6,993 images, and where does the curve flatten?
Cybersecurity and intrusion detection: five topics
- Cross-dataset generalisation. Does an intrusion detector trained on CIC-IDS2017 flows detect the attack families in UNSW-NB15, and which feature set survives the transfer?
- Flow versus packet features. Using CIC-IDS2017’s PCAPs and flow CSVs together, how much detection accuracy is lost when a model sees only flow statistics rather than payload-derived features?
- Class imbalance in rare attacks. For the least frequent attack classes in UNSW-NB15, which resampling or cost-sensitive method gives the best precision-recall trade-off?
- Concept drift. When a detector trained on the first days of CIC-IDS2017’s capture is tested on the later days, how quickly does performance decay and does online learning arrest it?
- Explainability of alerts. Which feature-attribution method produces explanations that a security analyst rates as most useful for UNSW-NB15 alerts, measured in a small user study?
Health and public data: six topics
- Length-of-stay prediction. With MIMIC-IV v3.1, once credentialed, how well do the first 24 hours of vital signs predict ICU length of stay, and does adding free-text notes help?
- Tabular baselines that are hard to beat. Across ten UCI classification datasets, how often does a tuned gradient-boosting model beat a deep tabular model, and by how much?
- District crime forecasting. Using NCRB’s district-wise tables from successive Crime in India issues, which time-series model best forecasts next year’s counts for one offence class, and how far out does it stay useful?
- Open data quality audit. Of a random sample of resources on data.gov.in in one ministry, what proportion are machine-readable, complete and current, and what does that imply for reuse?
- Air quality nowcasting. Using air-quality monitoring data published on the Open Government Data platform, how accurately can PM2.5 be predicted six hours ahead for one city, and which exogenous variables matter most?
- Automated table extraction. How accurately can a document-understanding model extract the tables in NCRB PDFs into machine-readable form, and which table layouts fail?
Earth observation and classic machine learning: six topics
- Crop-type mapping. From Bhuvan’s LISS-III imagery for one district, how accurately can a model separate three major crops across a season, validated against ground truth from a departmental survey?
- Flood-extent detection. Using AWiFS scenes before and after a flood event, which change-detection method best delineates inundated area?
- Urban growth. From multi-year Bhuvan imagery of one city’s outskirts, how has built-up area changed, and does a segmentation model trained on one year transfer to another?
- Feature selection at scale. Across a stratified sample of UCI datasets, which feature-selection method gives the most consistent gain, and when does it hurt?
- Calibration of classifiers. On UCI’s heart disease and similar clinical datasets, how miscalibrated are common classifiers, and which post-hoc calibration method is most reliable at small sample sizes?
- Reproducibility study. Can the published results of five recent papers on a UCI benchmark be reproduced from the released code, and what accounts for the gaps?
The datasets to stop planning around, and the ones to budget for
Three access conditions caught out students we have worked with, and each is visible on the custodian’s page if you look before choosing the topic.
- NSL-KDD is no longer distributed by its host. The Canadian Institute for Cybersecurity’s page states that the dataset is no longer available. Copies circulate, but a dissertation built on a withdrawn benchmark invites the examiner’s first question. Use CIC-IDS2017 or UNSW-NB15 instead.
- MIMIC-IV needs credentialing. PhysioNet grants access after training and a signed data-use agreement. Start the application in the first month; a topic that depends on it cannot be swapped in October.
- IndicVoices is gated. The files are public but you must accept the conditions on Hugging Face before download, and the transcribed subset, 11,200 hours as of December 2025, is smaller than the 23,700-hour headline.
- IMD historical weather data is paid. The Data Supply Portal of the National Data Centre, Pune, in operation since March 2019, quotes data charges before supply; budget for it or use a free alternative.
- Kaggle is browser-only. Its dataset pages sit behind a bot check, so a pipeline that scripts downloads breaks; download once by hand and version the file.
How to turn a topic into a proposal
Every topic above has the same anatomy: one dataset, one measurable outcome and one comparison, which is what a two-semester M.Tech project can finish. The report structure your department expects, chapter by chapter, is in our guide to writing an M.Tech dissertation report, and the choice of evaluation metric, where a 99 per cent accuracy figure is the fastest route to a hard viva question, is worked through in our comparison of accuracy, F1 and AUC for a machine learning project. The same feasibility logic applied to a different discipline, and the six tests a topic must pass, is set out in our list of feasible MD and MS thesis topics.
Two wider maps help when none of the thirty fits. Official Indian data by discipline, beyond computer science, is catalogued in our guide to data sources for an Indian thesis, and the process of narrowing a field to a question, whatever the discipline, is in our guide to choosing a research topic in India. When the topic is fixed, Tesify drafts the problem statement, the dataset description and the evaluation plan in the structure your institute’s synopsis template uses, so the first guide meeting starts from a document; the free plan is enough for a synopsis, and the experiments remain yours.
Draft your M.Tech synopsis around one of these datasets with Tesify
Frequently asked questions
Which Indian datasets are free for an M.Tech thesis?
Samanantar, Naamapadam and IndicCorp v2 from AI4Bharat are free on Hugging Face; IndicVoices is free after accepting conditions; the Indian Driving Dataset and RoadSocial from IIIT Hyderabad are free for research; the IIT Bombay English-Hindi corpus, NCRB tables, data.gov.in resources and Bhuvan’s open archive are free.
How big is the Indian Driving Dataset?
As published by IIIT Hyderabad in 2026: 10,003 images with 34 classes from 182 drive sequences around Hyderabad and Bengaluru, split into 6,993 training, 981 validation and 2,029 test images, mostly at 1080p.
Is NSL-KDD still available?
Not from its host. The Canadian Institute for Cybersecurity’s page states that the dataset is no longer available. Use CIC-IDS2017 or UNSW-NB15, both of which are current and free.
Can an M.Tech student get MIMIC-IV?
Yes, through PhysioNet’s credentialed access, which requires completing the required training and signing a data-use agreement. Apply at the start of the project, because the process takes time and the topic depends on approval.
How many datasets does the UCI repository hold?
The repository stated in August 2026 that it maintains 689 datasets, including classic ones such as Iris and the four-database heart disease collection.
How much speech does IndicVoices contain?
The custodian reports 23,700 hours from 51,000 speakers across 22 languages and more than 400 districts, of which 11,200 hours were transcribed as of December 2025. Plan the topic around the transcribed subset.
Is Kaggle a citable source for a thesis dataset?
Cite the original custodian wherever one exists, because Kaggle mirrors are unversioned copies. Where a dataset exists only on Kaggle, cite the Kaggle page with the version and download date, and keep a copy of the exact file.
How many of these topics need a GPU?
The tabular, forecasting and audit topics run on a laptop. The translation, speech and vision topics need a GPU for fine-tuning, which most institutes provide through a departmental server or a cloud credit; scope the experiments to what you can queue.
