Building your own tool is the most defensible thing you can do and the easiest thing to do badly. A self-constructed questionnaire with a blueprint, an expert panel and a try-out behind it is stronger evidence than a borrowed scale used without permission on a population it was never normed for. A self-constructed questionnaire assembled the week before data collection is the weakest instrument in Indian research.
This guide is the construction sequence, written out. Before you start, though, settle whether you need to build at all. Selecting from published tools is a different job and usually the right one: our guide to standardised tools for an M.Ed dissertation covers permission, adaptation and translation, and our comparison of validated scales for a psychology dissertation in India covers the main published instruments for psychological constructs.
Build only when no available tool measures your construct in your setting, and say that in one sentence in Chapter 3.
Step 1: Write the blueprint before you write a single item
The blueprint is a table that fixes what the tool measures, in what proportion, before any wording exists. It is what makes the content-validity claim checkable, and it is the document your expert panel actually reviews.
| Dimension | What it covers | Weight | Items planned | Item numbers |
|---|---|---|---|---|
| [Dimension 1] | [Sub-topics from the operational definition] | [%] | [n] | [list] |
| [Dimension 2] | [Sub-topics] | [%] | [n] | [list] |
| [Dimension 3] | [Sub-topics] | [%] | [n] | [list] |
| Total | 100% | [N] |
Two rules govern the weights. They come from your operational definition and your review of literature, not from how easy each dimension is to write items for. And they are decided before item writing, because a blueprint reconstructed afterwards is a description of what you happened to write, which is not content validity.
Draft more items than you need. A common working ratio is to write roughly twice the number you intend to keep, because expert review and item analysis will remove some and you do not want to be rewriting under deadline.
Step 2: Write the items, then fix them
Six rules cover most of what an expert panel will flag. Each is shown with a faulty item and its repair.
| Rule | Faulty item | Repaired |
|---|---|---|
| One idea per item | “My supervisor is available and gives useful feedback.” | Split into two items: availability, and usefulness of feedback. |
| No leading wording | “Do you agree that the new assessment system has improved learning?” | “The new assessment system has improved learning in my class.” (with a full agreement scale) |
| No assumed behaviour | “How often do you use the competency descriptors?” | Add a filter item first: “Do you use the competency descriptors?” Then ask frequency of those who do. |
| No undefined quantifier | “I frequently give written feedback.” | “In the last one month, I gave written feedback on [n] or more occasions.” Or define the frequency scale in the instructions. |
| No negatives inside a negative scale | “I do not find item construction difficult.” (Strongly disagree … Strongly agree) | “I find item construction difficult.” Reverse-score it instead of writing a double negative. |
| Language the respondent uses | “Rate your metacognitive regulation during summative assessment design.” | “When I set a test, I check whether the questions match what I taught.” |
Decide the response format once and hold it. Mixing four-point and five-point scales within one tool creates a scoring problem that surfaces when you compute a total, and mixing agreement scales with frequency scales in the same block confuses respondents even when it is logically fine.
Step 3: Content validation by an expert panel
Content validity is an argument, not a statistic, and the argument is: competent people in this field, shown the blueprint and the items, agreed that the items cover the construct. To make it checkable, give the panel something to rate rather than something to comment on.
The proforma to send each expert
Tool validation proforma
Title of the study: [title] · Investigator: [name], [programme], [department] · Supervisor: [name]
Operational definition of the construct: [one paragraph]
Blueprint: [the table above]
For each item, please rate: Relevance to the dimension (1 = not relevant, 2 = somewhat relevant, 3 = relevant, 4 = highly relevant) and Clarity of language (1–4 on the same scale). Where you rate an item below 3 on either count, please suggest a revision or mark it for deletion.
[Table: Item number | Item text | Dimension | Relevance 1-4 | Clarity 1-4 | Suggested revision]
Name and designation of the expert: ____ · Institution: ____ · Signature and date: ____
Two design choices to make before you send it. Decide the retention criterion in advance and state it in Chapter 3: for example, items rated 3 or 4 on relevance by a stated proportion of the panel are retained, items rated below that are revised or deleted. And keep the returned proformas: they are the evidence, and some Indian departments require them in the appendix.
Where a numeric index is expected, the two in common use both work from the same ratings. A content validity ratio for an item is computed as the number of experts rating it essential minus half the panel, divided by half the panel, which runs from −1 to +1. A content validity index for an item is simply the proportion of experts rating it relevant. The critical value that counts as acceptable depends on how many experts you had, and it should be taken from the published table your department accepts rather than assumed — state the table you used and the panel size alongside the values.

Step 4: The try-out, and what it actually tests
The try-out is not a small version of your study. It answers four questions, none of which is “what are my results”.
- Does the instrument work as an object? Do respondents understand the instructions, find the response format obvious, and finish it? Record the time taken, because Chapter 3 needs that number.
- Do any items fail on the ground? Items everyone answers identically, items many skip, items respondents ask about.
- Is it internally consistent? Compute reliability on the try-out data, for the whole scale and for each sub-scale.
- Does the procedure work? Access, timing, the consent conversation, and how long the whole visit takes.
Run the try-out on respondents drawn from the population but excluded from the main sample. Reusing them contaminates your results, and an examiner who spots it will ask about it. What counts as an acceptable reliability value, and why deleting items to raise it is a trap rather than a fix, is set out in our guide to acceptable Cronbach’s alpha for a thesis.
Step 5: Item analysis for knowledge and achievement tests
If your tool is an achievement or knowledge test with right and wrong answers, the try-out yields two more numbers per item.
| Index | What it is | What it tells you |
|---|---|---|
| Difficulty index | The proportion of the try-out group answering the item correctly | Items almost everyone gets right or wrong carry little information about differences between respondents |
| Discrimination index | The difference in the proportion correct between the higher-scoring and lower-scoring groups | A near-zero or negative value means the item does not separate stronger from weaker respondents, or separates them the wrong way |
Report the retention rule you applied, and report it as a rule fixed in advance. A negative discrimination index is the clearest signal to re-read an item: usually the key is wrong, or the wording misleads the people who know the material.
The Chapter 3 paragraph that reports all of it
Since no available tool measured [construct] among [population] in [setting], an instrument was constructed for this study. A blueprint was prepared from the operational definition and the review of literature, distributing [N] items across [n] dimensions in the proportions shown in Table [n]. An initial pool of [n] items was drafted.
The pool, together with the blueprint and the operational definition, was submitted to [n] experts in [field] from [type of institutions]. Experts rated each item for relevance and clarity on a four-point scale using the proforma at Appendix [n]. Items rated [criterion] were retained; [n] items were revised on expert suggestion and [n] were deleted, yielding a [n]-item instrument.
The revised instrument was tried out on [n] respondents drawn from the population but excluded from the main sample. The mean completion time was [n] minutes. Cronbach’s alpha was [value] for the total scale and ranged from [value] to [value] across the sub-scales. [For tests: item analysis yielded difficulty indices from [value] to [value] and discrimination indices from [value] to [value]; [n] items falling below the criterion of [value] were removed.] The final instrument comprises [n] items and is at Appendix [n].
Three paragraphs, every number traceable, and nothing asserted that an appendix does not support. Where this sits among the other eight sections is set out in our written-out methodology chapter, and the definitions the blueprint is built from are in our guide to operational definitions across six disciplines.
Translating and adapting
If your respondents work in a regional language, the instrument has to exist in that language, and a single translation is not enough. The accepted practice is forward translation by one bilingual, back translation to the original by a second who has not seen the source, and reconciliation of the differences by a panel including the investigator. Keep all three versions, and describe the procedure in Chapter 3.
Two things to watch. Response scales are harder to translate than items: the distance between “agree” and “strongly agree” does not survive translation automatically, so pilot the scale labels with real respondents. And any item referring to an institutional practice may need adaptation rather than translation, because the practice itself may differ.
Five faults examiners flag in a constructed tool
- A blueprint written after the items. Dimensions with wildly uneven item counts and no stated weights is the usual tell.
- Experts who rated nothing. “The tool was validated by five experts” with no proforma, no criterion and no record of changes.
- A try-out inside the main sample. The pilot respondents appear again in Chapter 4.
- Reliability reported for the wrong version. Alpha computed on the pre-revision pool and reported for the final instrument.
- Items that no objective needs. The instrument collects data that appear nowhere in the results, which invites the question of why respondents were asked.
The last point is worth checking against your objectives before printing: the verb table in our guide to objectives and the design each verb commits you to is the fastest way to see whether every item earns its place. And who the instrument is administered to is a separate decision, worked through in our guide to population and sample with the paragraph you have to write.
Building the instrument while the chapter takes shape
Tool construction fails on sequence: the items get written before the blueprint, the panel is approached before the criterion is fixed, and Chapter 3 is reconstructed from memory afterwards. Tesify holds the operational definitions, the blueprint and the methodology section together as the document grows, so the instrument and the chapter that reports it stay in step.
Build your Chapter 3 in Tesify
Frequently asked questions
Is a self-constructed questionnaire acceptable in an Indian dissertation?
Yes, where no available tool measures your construct in your setting, and provided the construction is documented: a blueprint, expert validation with a stated criterion, a try-out and reliability evidence. State the reason for constructing in one sentence in Chapter 3.
What is a blueprint in tool construction?
A table fixing the dimensions of the construct, the weight of each, and the number of items allotted to each, prepared before any item is written. It is what the expert panel reviews and what makes the content-validity claim checkable.
How many experts should validate the tool?
There is no national rule, and departments differ. What matters to an examiner is that you state how many, their field, their type of institution, and the retention criterion you fixed before the ratings came back.
How many items should I draft initially?
More than you intend to keep, because expert review and item analysis will remove some. A working ratio of roughly two drafted for each one retained leaves room without creating an unmanageable pool.
What is the difference between a difficulty index and a discrimination index?
The difficulty index is the proportion answering an item correctly. The discrimination index is the difference in that proportion between higher and lower scorers. An item can be at a good difficulty level and still fail to discriminate.
Can I use the pilot respondents in my main sample?
No. They have seen the instrument, and in an intervention study they may have received part of the intervention. Draw the try-out group from the population and exclude them from the main sample.
How many respondents does a try-out need?
Enough for the reliability statistic to mean something and for item-level problems to appear. Departments give different guidance, so follow your supervisor’s and state the number you used.
Do I need to report reliability for each sub-scale?
Yes, if you report sub-scale scores in your results. A good total-scale alpha can conceal a weak sub-scale, and an examiner who sees sub-scale tables will ask for the sub-scale values.
Do I need permission for a tool I built myself?
Not from anyone else, but you do need your ethics committee’s approval for the instrument as part of the protocol, and your department may require the final version to be filed with the supervisor before data collection.
How do I validate a translated version?
Forward translation, independent back translation by someone who has not seen the original, and reconciliation by a panel. Keep all versions, pilot the response-scale labels specifically, and describe the whole procedure in Chapter 3.
Should the tool appear in the appendix?
In most Indian dissertations yes, along with the expert proformas and the permission letters. Check your ordinance, since some specify exactly what the appendices must contain.
What if my alpha is low after the try-out?
Read the items before touching the statistic. A low value usually means the dimension is not coherent, or items are double-barrelled, or the sample was too homogeneous to show variance. Deleting items until the number rises is the fix that fails in the viva.
