The model runs. The F test is significant, the R squared is respectable — and not one predictor comes out significant. Or a coefficient has flipped sign against everything your literature review said it would do. You have three weeks to submission and Chapter 4 will not write itself around output that contradicts itself.
That pattern is multicollinearity, and it is the single most common reason an Indian commerce, management or economics thesis stalls at the results stage. It is also fixable in an afternoon, provided you do not reach for either of the two remedies most scholars reach for first.
Draft your results chapter in Tesify — free to start
What collinearity is actually doing to your output
Multicollinearity is a high degree of linear intercorrelation between the explanatory variables in a multiple regression. The mechanism is precise. For any predictor, regress it on all the other predictors and take the resulting coefficient of determination, R²ₕ. The variance inflation factor for that predictor is 1 ÷ (1 − R²ₕ), and the variance of its regression coefficient is proportional to that VIF.
So as one predictor becomes more predictable from the others, R²ₕ rises, the VIF rises, and the variance of the coefficient rises with it. A larger variance means a larger standard error, and a larger standard error produces exactly what you are looking at: unreliable p-values and confidence intervals wide enough to contain both a positive and a negative effect.
The important consequence is that your model is not broken and your data is not wrong. The overall fit is fine, which is why the F test still passes. What has collapsed is your ability to attribute the outcome to individual predictors — and individual predictors are usually what your hypotheses are about.

The three symptoms, in the order you notice them
- A significant F test with no significant t tests. The clearest signal, and the one that sends most scholars to a forum at midnight.
- A coefficient with the wrong sign. A predictor the literature says raises the outcome comes out negative. This is not a discovery; it is usually collinearity redistributing shared variance.
- Coefficients that move wildly when you add or remove one variable. If dropping a single control changes another predictor’s estimate by a large margin, the two are sharing variance.
How to diagnose it properly
| Diagnostic | Threshold indicating a problem | What it tells you | What it cannot tell you |
|---|---|---|---|
| Variance inflation factor (VIF) | Above 5 to 10 | That collinearity is present, and for which predictor the variance is inflated | Which other variables it is colliding with |
| Tolerance (1 ÷ VIF) | Below 0.1 to 0.2 | The same information as VIF, inverted; equivalent to R²ₕ of 0.8 to 0.9 | The same limitation as VIF |
| Condition index | Above 10 to 30 | That a near-dependency exists among the standardised predictors | Which predictors form it, on its own |
| Variance decomposition proportions (VDP) | Two or more above 0.8 to 0.9 on a shared condition index above 10 to 30 | Which variables are collinear with each other | What to do about them |
Those thresholds are the ones stated in the published methodological literature — Kim (2019), writing in the Korean Journal of Anesthesiology, sets out VIF above 5 to 10 and tolerance below 0.1 to 0.2 as the working criteria, and the condition index and VDP procedure as the way to identify the offending variables.
The point most scholars miss is in the fourth column. VIF flags that a predictor’s variance is inflated. It does not tell you what it is inflated by. Two variables can only be identified as collinear with each other through the variance decomposition proportions: when two or more VDPs sitting on the same high condition index both exceed 0.8 to 0.9, those are your colliding variables. Every major statistics package will produce collinearity diagnostics including condition indices and VDPs if you ask for them; the reason most theses do not report them is that most scholars stop at the VIF column.
The two things not to do
Do not run stepwise regression to “solve” it. This is the standard advice passed down in Indian coursework and it is wrong. Writing in the International Journal of Surgery in 2024, Xi, Jiang and Yang of Peking Union Medical College Hospital titled their correspondence precisely that: using stepwise regression to address multicollinearity is not appropriate. Stepwise selects on the same unstable standard errors that collinearity has already corrupted, so it makes an arbitrary choice between two collinear variables and presents it as a finding. An examiner who knows this will ask you to justify the selection procedure, and there is no good answer.
Do not simply delete the variable with the highest VIF. If that variable is theoretically relevant to your model — if your hypotheses or your literature say it belongs there — excluding it produces biased regression coefficients, which is a more serious problem than the collinearity you were trying to remove. Omitting a relevant predictor is not a neutral simplification; it changes what every remaining coefficient means.
A third, quieter error: screening only pairwise correlations. Two variables can each correlate modestly with every other variable and still form a near-dependency together. Pairwise correlation has strict applicability conditions and does not substitute for a collinearity diagnostic run on the fitted model.

Four remedies you can defend in a viva
1. Find the duplicate you did not know you had
Start here, because it is free and it is often the whole answer. The unintentional inclusion of the same construct twice produces a collinear model by construction: total income and total expenditure, a summated scale and one of its own items, age and years of experience in a workforce sample, a variable and a transformation of that same variable. Sometimes a near-exact relationship can be written as an equation linking the two, in which case including that equation in the model removes one of them outright.
If you find a duplicate, removing it is not a compromise. It is a data-handling correction, and you say so.
2. Replace the variable with a better-measured one
Where two predictors collide because both are noisy proxies for the same underlying thing, the fix may be to substitute a single variable that measures the construct more accurately. This is a design decision, not an analysis trick, and it is defensible precisely because you can explain what the new variable measures.
3. Combine the colliding variables into one
Principal component analysis or factor analysis can generate a single variable that carries the shared information from several collinear ones. This works and is widely accepted — with one explicit cost that you must state: once combined, it is no longer possible to assess the effect of the individual variables. If your hypotheses are about those individual effects, this remedy answers a different question from the one you asked.
If your predictors are summated scales, this is also the moment to confirm each one is actually unidimensional and reliable, which is the argument our guide to an acceptable Cronbach’s alpha for a thesis works through.
4. Use ridge regression to keep every variable
Where all the collinear variables are theoretically necessary and none can be dropped, ridge regression is the recognised alternative: it retains every predictor in the model and accepts a small bias in exchange for far more stable estimates. It is more demanding to report than ordinary least squares and your supervisor may not have used it, so raise it early rather than presenting it as a surprise in Chapter 4.
One remedy that sounds appealing and mostly is not: increasing the sample size. In theory a larger N reduces the standard errors of the coefficients and therefore the practical impact of collinearity. In practice, under strong collinearity the standard errors are not reliably reduced, and in any case you are unlikely to be collecting more data three weeks before submission.
How to write it up so it strengthens the thesis
A collinearity problem that you diagnosed, acted on and reported reads as competence. The same problem left in the output and not mentioned reads as an oversight the examiner found for you. Report four things, in this order:
- The diagnostic and the threshold you applied, named explicitly, with a citation.
- The values for every predictor, in the regression table, not in prose. A VIF column beside the coefficients costs one column and answers the question before it is asked.
- Which variables were identified as colliding, and by which diagnostic.
- The action taken and the reason for it, including what you chose not to do.
A worked sentence for your results chapter:
Collinearity diagnostics were examined before interpreting the coefficients. Variance inflation factors for all predictors were below 5 except for perceived value (VIF = 8.42) and perceived quality (VIF = 7.96), which shared a condition index of 24.6 with variance decomposition proportions of 0.91 and 0.88 respectively, indicating a near-dependency between them. As both constructs were required by the hypothesised model, they were retained and the model re-estimated using ridge regression; the ordinary least squares estimates are reported in Appendix C for comparison.
The table conventions for reporting that output — where the note goes, how many decimal places, how to report the statistic in running text — are covered in our guide to writing the results chapter with APA tables. Which package you run any of this in makes no difference to the diagnosis; our comparison of statistical tools for an Indian thesis covers the options and their costs.
The real cost of leaving it
An unaddressed collinearity problem does not usually fail a thesis. It does something slower and more expensive: the external examiner queries the coefficient interpretation, the report comes back with corrections, and a submission that should have closed in six months acquires another cycle. In a system where the whole evaluation is meant to conclude within six months of submission, a corrections round is the difference between graduating this year and graduating next.
Ten minutes of diagnostics now avoids that. The analysis plan you set out in your synopsis should have named the assumption checks in advance — if yours did not, our guide to writing a PhD synopsis shows where they belong for next time.
Write the chapter, keep the analysis yours
Tesify does not run your regression and does not decide which variable to drop. What it does is hold the structure of your results chapter while you work — the assumption checks you promised in Chapter 3 sitting beside the output you are reporting in Chapter 4, every citation attached to the source you actually read, and the methodological reference for your VIF threshold still findable in week eleven. There is a free tier, so you can put this chapter into it today and decide afterwards.
Start your results chapter in Tesify
Frequently asked questions
What is an acceptable VIF value for a thesis?
The published working criterion is that multicollinearity is present when VIF exceeds 5 to 10, equivalently when tolerance falls below 0.1 to 0.2. Name the threshold you applied and cite it, because the range is a convention rather than a constant.
Is VIF above 5 or above 10 the cut-off?
Both appear in the literature and neither is definitive. Applying 5 is the more conservative choice and the easier one to defend. What matters more than the number is that you stated it in advance and applied it consistently.
Can I just remove the variable with the highest VIF?
Not if it is theoretically relevant. Excluding a relevant explanatory variable produces biased coefficients, which is a more serious problem than the collinearity. Remove a variable only when it is a duplicate or a derived version of another.
Why is stepwise regression not the answer?
Because it selects variables using the same standard errors that collinearity has already made unstable, so its choice between two collinear predictors is arbitrary. Published methodological correspondence states the point directly: using stepwise regression to address multicollinearity is not appropriate.
Does VIF tell me which variables are collinear?
No, and this is its central limitation. VIF identifies that a predictor’s variance is inflated. To identify which variables are colliding with each other you need variance decomposition proportions: two or more above 0.8 to 0.9 sharing a condition index above 10 to 30.
Will a bigger sample fix multicollinearity?
In theory a larger sample reduces the standard errors of the coefficients. In practice, under strong collinearity those standard errors are not always reduced, so treat this as a design consideration for your next study rather than a remedy for this one.
Is it a problem if my dummy variables have high VIF?
Sets of dummy variables representing one categorical predictor are related by construction, and high VIF among them is expected rather than diagnostic. Interpret collinearity across distinct constructs, not within a single dummy set.
Do I have to report VIF even if there is no problem?
Report it. A regression table with a clean VIF column tells the examiner you checked; a table with no VIF column invites the question of whether you did. The check is cheap and the omission is expensive.
Is using an AI tool for my results chapter against academic integrity norms?
Writing assistance is not the same as fabricated analysis. Your institution’s integrity policy is the binding document, and the safe line is the same one that applies to a human editor: the analysis, the judgements and the interpretation must be yours, and any assistance should be declared where your policy requires it.
What does Tesify cost?
There is a free tier you can start on without paying, which is enough to structure a chapter and see whether it fits how you work. Check the current plans on the site before committing to anything paid — and never pay for a tool you have not used on your own draft first.
Is my unpublished thesis data safe in a writing tool?
A fair question to ask of any tool that holds unpublished work. Read the privacy terms of whatever you use, check whether your content is used to train models, and where your data is genuinely sensitive, keep raw participant data in your own analysis files and put only the written chapter into a drafting tool.
My supervisor has never used ridge regression. Should I still propose it?
Raise it early, with the diagnostics that motivate it, and offer the ordinary least squares estimates alongside as a comparison. A remedy introduced in Chapter 4 without warning invites resistance; the same remedy discussed in a supervision meeting usually does not.
