The idea
A baseline is a simple comparison method, such as predicting the most common class. Keep evaluation examples separate from training and threshold tuning. Duplicate or future-derived information can leak into the test and inflate scores. Accuracy can be misleading when one class is rare, so record the class distribution and the kinds of mistakes that matter.
Worked example
Only 10 of 100 fictional messages are urgent. A system that labels everything non-urgent has 90% accuracy but misses every urgent message. That baseline reveals why an accuracy-only target is insufficient. Use labeled examples not used to develop the classifier to assess whether a new system improves the actual decision.
Try it
Create ten fictional messages with two marked urgent. Calculate the accuracy of always predicting non-urgent and count missed urgent cases. Define a separate development and test set. Identify one way that copying the same example into both sets would distort the evaluation.
