AI & ML 2025 07 / 14
Expression Model Benchmark
Five models, one problem — measuring what actually moves accuracy.
Built with
Language
- Python
Skills & tools
- PyTorch
- torchvision
- scikit-learn
- CNNs
- Model evaluation
A controlled comparison of a baseline, linear and logistic regression, a deeper MLP and an augmented CNN on 48×48 facial-expression images, built in PyTorch and scikit-learn.
Question
How much does model choice really matter for facial-expression recognition? To answer honestly, every model gets the same data split and the same metrics.
Results (validation)
| Model | Accuracy | Precision | Recall |
|---|---|---|---|
| Baseline (majority class) | 0.250 | 0.036 | 0.143 |
| Linear regression | 0.221 | 0.221 | 0.154 |
| Logistic regression | 0.352 | 0.329 | 0.303 |
| Deeper MLP (256→128, dropout) | 0.455 | 0.472 | 0.416 |
| CNN + augmentation | 0.560 | 0.570 | 0.480 |
Takeaways
- Treating classification as regression is a trap — linear regression does worse than guessing the majority class.
- Non-linearity helps a lot, but convolutions and augmentation are what unlock the jump.
- A strong baseline keeps everyone honest.