,

Accuracy, F1 or AUC? Choosing the Evaluation Metric for an MCA Machine Learning Project (2026)

Verdict up front: report accuracy alongside macro-averaged precision, recall and F1, with the confusion matrix, for any classification project — and replace ROC-AUC with PR-AUC the moment your positive class falls below roughly 10% of the data.

Metric What it measures Range Breaks down when Use it for
Accuracy Proportion of all predictions that were correct 0 to 1 Classes are imbalanced — a 95:5 split gives 95% accuracy for predicting the majority class every time Balanced multi-class problems; always report it, never report only it
Precision Of the items predicted positive, how many were positive 0 to 1 Positive predictions are very few — the value becomes unstable Projects where a false positive is costly: spam filters, fraud alerts
Recall (sensitivity) Of the actual positives, how many were found 0 to 1 Read alone — trivially maximised by predicting everything positive Projects where a missed positive is costly: disease screening, intrusion detection
F1 score Harmonic mean of precision and recall 0 to 1 The two error types have genuinely different costs, which F1 hides by weighting them equally The default single-number summary for imbalanced binary classification
Macro-F1 F1 computed per class, then averaged unweighted 0 to 1 Some classes have very few test instances, making their F1 noisy Multi-class problems where small classes matter as much as large ones
ROC-AUC Probability the model ranks a random positive above a random negative 0.5 (chance) to 1 Severe imbalance — it stays optimistically high because true negatives dominate Roughly balanced binary problems; comparing ranking quality across models
PR-AUC (average precision) Area under the precision-recall curve Baseline equals the positive class rate Nothing much — but the baseline moves with class balance, so it is not comparable across datasets Rare-event detection: fraud, defect prediction, rare disease
Matthews correlation coefficient Correlation between predicted and actual labels using all four confusion-matrix cells -1 to +1 Rarely; it is the most conservative single number available Binary classification where you want one honest headline figure
Cohen’s kappa Agreement corrected for agreement expected by chance -1 to +1 Marginal distributions are very skewed Multi-class classification; comparing to human labelling
RMSE / MAE Average size of prediction error, in the units of the target 0 upwards Compared across datasets with different target scales Regression projects: price prediction, demand forecasting
R² (coefficient of determination) Proportion of target variance explained Up to 1, can be negative Used alone — a high R² can coexist with unusable absolute error Regression, reported alongside RMSE, never instead of it

The ranked shortlist: which metric for which project

1. Balanced binary or multi-class classification — accuracy plus macro-F1

Sentiment classification on a balanced review corpus, handwritten digit recognition, image classification across similarly sized classes. Report accuracy as the headline, macro-F1 to show that performance is not concentrated in the easy classes, and the confusion matrix so the examiner can see which classes the model confuses.

Where it falls short: accuracy tells you nothing about which class is failing, which is exactly what the external examiner will ask about.

2. Imbalanced binary classification — F1 and PR-AUC

Credit card fraud, churn prediction, defect detection, disease prediction on a real clinical dataset. The positive class is rare, so accuracy is uninformative and ROC-AUC is flattering. Report precision, recall and F1 for the positive class specifically, plus average precision.

Where it falls short: F1 assumes false positives and false negatives cost the same. In a screening application they do not. Say which error your application cares about and justify the threshold you chose.

3. Ranking and recommendation — precision@k, recall@k, NDCG

Recommender-system projects are the single most common MCA final-semester topic and the most commonly mis-evaluated. Users see the top few items, so evaluate the top few items: precision at k and recall at k for a stated k, and NDCG when the position within the list matters. RMSE on predicted ratings answers a question no user asks.

Where it falls short: offline ranking metrics do not measure whether recommendations are diverse or novel, which is worth a sentence in your limitations.

4. Regression — RMSE with MAE, and R² as context

House price prediction, sales forecasting, crop yield estimation. RMSE penalises large errors more heavily and is in the units of your target, so it is interpretable; MAE alongside it reveals whether a few large errors are driving the figure. R² is context, not conclusion.

Where it falls short: both are scale-dependent. Comparing your RMSE to a figure from a published paper on a different dataset proves nothing.

5. Clustering — silhouette score, Davies-Bouldin, and a stated k rationale

Customer segmentation projects need internal validation indices because there are no labels. Report the silhouette score across a range of k values and show the elbow or silhouette plot that justified your chosen k. Choosing k = 3 because it produced a tidy chart is the answer that fails a viva.

The four numbers that actually survive the viva

Whatever your project type, the results section should carry a table of four things for every model you compare:

  1. A baseline. The majority-class classifier, or a simple logistic regression, or the mean predictor for regression. Without a baseline, no metric value means anything — 82% accuracy is impressive against a 50% baseline and embarrassing against a 79% one.
  2. Your metric with its variability. A single number from a single train-test split is a sample of one. Report the mean and standard deviation across k-fold cross-validation, stating k and the random seed.
  3. The confusion matrix for the final model, so per-class behaviour is visible.
  4. Training and inference time, if you are claiming a model is practical. A deep model that beats a random forest by half a percentage point while taking forty times as long is a finding worth reporting honestly.

Why a 99% accuracy figure invites the hardest question

Near-perfect results in a student project almost always have one of four causes, and Indian external examiners in computer applications departments know all four:

  • Data leakage. A feature that encodes the target, or scaling and imputation fitted on the whole dataset before the split rather than on the training fold alone.
  • Duplicate rows appearing in both training and test sets, common in scraped or augmented datasets.
  • Testing on the training set, or tuning hyperparameters against the test set until it looks good — which converts the test set into a second training set.
  • A trivially separable benchmark dataset that has been solved a thousand times, where 99% is the known ceiling and your contribution is unclear.

The defence is procedural and belongs in your methodology: split first, fit preprocessing on the training fold only, keep a held-out test set untouched until the final run, and report the split ratio and seed. Write that paragraph before the examiner asks for it.

Choosing the metric before you choose the model

The order matters. Decide what a correct prediction is worth and what each kind of error costs in your application, and the metric follows from that; pick the model first and you will be tempted to pick whichever metric flatters it. This is the same discipline that applies to selecting an inferential test in a quantitative dissertation, which our decision table for choosing a statistical test lays out for hypothesis-driven work — and if you are comparing two models’ scores and want to claim one is genuinely better, you need a statistical test, not a bigger decimal.

Where the metrics go in an Indian project report

MCA and M.Tech project reports in Indian universities generally follow a five- or six-chapter structure, with implementation and testing separated from results and discussion. Metrics appear in three places: the evaluation methodology in the methods chapter, stating which metrics you will report and why; the results chapter, containing the comparison tables and plots; and the conclusion, restating the headline figure with its baseline.

Whether your submission is called a project report, a dissertation or a thesis is not cosmetic in India — the three terms map to three degree levels with different requirements, as our explainer on the difference between a thesis, a dissertation and a project report sets out. For the chapter-by-chapter structure of a technical postgraduate report, our guide to writing an M.Tech dissertation report transfers almost directly to MCA work, and the layout conventions for the results tables themselves are covered in our guide to writing the results chapter with tables and figures.

One caution specific to computer applications projects: code and text generated with AI assistance must be declared according to your institution’s policy, and the model outputs you report must be your own runs. Our summary of what the Indian academic integrity rules require of AI-assisted work covers what disclosure normally looks like.

Write the evaluation chapter around your own numbers

Once your experiments are logged, the evaluation chapter is a writing problem, not a research one: the same tables, the same explanations of why each metric was chosen, the same careful sentences about baselines and splits. Tesify turns your results log into a structured chapter in your department’s format, with the references handled as you write, so the report is ready before the demo.

Build your project report with Tesify and get the evaluation chapter done this week.

Frequently asked questions

Is accuracy ever enough on its own for an MCA project?

Only when the classes are close to balanced and every error type costs the same, and even then you should show the confusion matrix. Reporting accuracy alone on an imbalanced dataset is the single most common evaluation criticism in project vivas, because a trivial majority-class predictor would score almost as well.

What is the difference between macro, micro and weighted averaging?

Macro averaging computes the metric for each class and averages them equally, so a rare class counts as much as a common one. Micro averaging pools all predictions before computing, so large classes dominate — for single-label classification, micro-F1 equals accuracy. Weighted averaging weights each class’s score by its support. State which one you used; an unlabelled F1 score is ambiguous.

Should I use ROC-AUC or PR-AUC for an imbalanced dataset?

PR-AUC. ROC-AUC includes true negatives in its calculation, and when negatives vastly outnumber positives that inflates the score even for a model with poor precision. The precision-recall curve ignores true negatives, so it reflects performance on the class you care about. Note that the PR-AUC baseline is the positive class rate, so report that rate alongside the score.

How many folds should I use for cross-validation?

Five or ten folds are the common choices, with stratification for classification so that class proportions are preserved in every fold. Ten folds give a slightly less biased estimate at higher computational cost. For small datasets, repeated stratified k-fold gives a more stable estimate than a single pass. Report the number of folds, whether it was stratified, and the random seed.

Can I claim my model is better than the published baseline?

Only if you evaluated both on the same dataset with the same split and the same metric. Comparing your score against a number quoted from a paper that used a different partition, different preprocessing or a different version of the dataset is not a comparison. Re-run the baseline yourself; that re-run is itself a legitimate contribution in a postgraduate project.

What threshold should I use to convert probabilities into class labels?

The default 0.5 is a convention, not a rule, and it is rarely optimal for imbalanced data. Choose the threshold on validation data according to what your application values — maximising F1, or fixing recall at a required level and reporting the resulting precision. Whatever you choose, state the threshold in the report; a metric quoted without its threshold cannot be reproduced.

Do I need a validation set as well as a test set?

If you tuned anything — hyperparameters, features, thresholds — then yes. Tuning against the test set makes the reported test score optimistic, and an examiner who reads your methodology carefully will spot it. Use a three-way split, or nested cross-validation on a small dataset, and touch the test set once.

How do I evaluate a project that has no labelled data at all?

Use internal validation indices for clustering, and supplement them with a small manually labelled sample — even 100 hand-checked items lets you report precision on a subset, which is far more convincing than a silhouette score alone. Describe how you drew and labelled that sample so the estimate is reproducible.