Direct answer
With only 1 percent of examples labeled, a semi-supervised approach is worth trying, since there's likely a large amount of untapped structure in the 99 percent unlabeled data that a purely supervised model trained on the 1 percent alone would never see. But whether it actually helps has to be checked, not assumed, because a poorly-chosen semi-supervised technique can just as easily add confident noise as it can add useful signal.
Structured elaboration
- Why semi-supervised is worth trying here: a model trained on only 1 percent of the data has very little to learn from directly, and is likely to be both high-variance (unstable across different random 1 percent samples) and missing patterns that only show up with more examples. Techniques such as self-training (using the model's own confident predictions on unlabeled data as if they were labels, then retraining on the enlarged set) or consistency-regularization let the model extract additional signal from the unlabeled 99 percent, without needing new human labels, which is exactly the constraint here (getting more labels is expensive).
- What would make you confident it's actually helping: the honest test is a controlled comparison, not just checking whether the semi-supervised model's training-time metrics look good (a self-training approach in particular can look confidently good on the pseudo-labels it generated for itself, which is a weak signal since it's grading its own homework, while actually being wrong). Train a baseline model on only the labeled 1 percent, train the semi-supervised model using the same labeled 1 percent plus the unlabeled 99 percent, and compare both on the same held-out labeled evaluation set that neither model trained on. If the semi-supervised model beats the baseline there, by a margin larger than the run-to-run noise you'd see just from re-training either model with a different random seed, that's real evidence it helped.
- Setting up a fair comparison: use the identical held-out evaluation set for both models, sourced entirely from labeled data that neither model touched during training. Since 1 percent of the data is a small sample, also check that the comparison isn't just noise from that particular small labeled split, for example by repeating the comparison across a few different random 1 percent labeled subsets, or by looking at whether the improvement is consistent rather than a one-off.
Worked example
Suppose the full dataset has 500,000 examples, so 1 percent labeled means roughly 5,000 labeled examples. You'd hold out a separate labeled evaluation set, say 1,000 of those 5,000, and train the supervised baseline on the remaining 4,000. The semi-supervised model would also train on those same 4,000 labeled examples, plus the roughly 495,000 unlabeled ones, using a technique like self-training. Both models get evaluated on the same 1,000 held-out labeled examples, and only if the semi-supervised model's score there is meaningfully and consistently higher than the baseline's would you conclude the extra unlabeled data was actually worth using.
Trade-offs and pitfalls
The main trap is judging "it's helping" from the semi-supervised model's own training process rather than an independent, held-out comparison; a self-training loop, in particular, can report improving performance on its own pseudo-labels while actually drifting away from correct behavior, because it's partly grading its own homework. Getting more labels, if it's feasible even at a slower pace, is also worth weighing against the engineering cost and added complexity of a semi-supervised pipeline; sometimes a modest amount of additional targeted labeling (focused on the examples the current model is least confident about) is a more sample-efficient investment than a more elaborate semi-supervised setup.