From raw audio to screening result in under 3 seconds
SickNote is a binary cough classifier that distinguishes healthy coughs from abnormal ones. It converts audio recordings into mel spectrograms — visual representations of sound frequencies — and feeds them through an ensemble of three convolutional neural networks. Each prediction includes a Grad-CAM heatmap showing which spectrogram regions drove the result.
The COUGHVID dataset contains ~34,400 crowdsourced cough recordings. Of these, 2,841 were reviewed by four expert physicians who provided diagnostic labels.
34,400
Total recordings
2,841
Expert-labeled
2,267
After filtering
Filtering criteria
Class distribution
A small convolutional neural network trained from scratch, 71,793 parameters per model. Three models are trained with different random seeds and their predictions are averaged. Conv features are pooled over the time axis only — never over frequency, since 200 Hz and 2 kHz are genuinely different sounds while a cough at 2s and a cough at 5s are the same cough. Each result includes a Grad-CAM heatmap showing which spectrogram regions influenced the classification.
Input: (1, 64, T) mel spectrogram
↓
Conv2d(8) + BatchNorm + ReLU + MaxPool
Conv2d(16) + BatchNorm + ReLU + MaxPool
Conv2d(32) + BatchNorm + ReLU + MaxPool
↓
Flatten
Linear(128) + ReLU + Dropout(0.5)
Linear(1) → logit
Loss function
BCEWithLogitsLoss + pos_weight
Optimizer
Adam (lr=3e-4, wd=1e-4)
Ensemble
3 models, different seeds, averaged probabilities
Explainability
Grad-CAM heatmaps on last conv block
| Metric | Actual | Reference |
|---|---|---|
| AUC-ROC | 0.7238 | 0.6556 from clip duration and loudness alone |
| Improvement over previous | +0.0308 | 95% CI [+0.0137, +0.0494], better on 5 of 5 folds |
| Sensitivity | 0.8029 | policy floor is 0.80 |
| Specificity | 0.5054 | at the same operating point |
| Rater agreement | κ 0.181 | four physicians on the same audio — the ceiling |
Classification threshold is 0.46, selected on the validation split under a stated policy: catch at least 80% of abnormal coughs, then take the best specificity available. An earlier version shipped 0.52, which was the argmax of Youden's J computed on the test set — a value fitted to the data it was then reported on. Figures above come from nested 5-fold cross-validation, in which all 2,267 clips are scored by a model that never trained on them.
The engineering choices behind SickNote
Architecture
Multi-class classification (COVID, URTI, LRTI, obstructive disease) dropped accuracy sharply because pathological coughs occupy overlapping feature space in the spectrogram domain. Binary classification — healthy vs. abnormal — is medically honest and performs significantly better with limited data. The output is "something sounds off," not a specific diagnosis.
Architecture
With only 2,267 expert-labeled samples, a small 3-layer CNN with aggressive dropout (0.5) and reduced channel widths [8, 16, 32] was the right fit. Larger architectures overfit before learning useful features. We tuned every hyperparameter against real dataset statistics from explore.py before writing a single training loop.
Training
We attempted to replace our CNN with a pretrained ResNet18 backbone to leverage features learned from millions of ImageNet images. The plan: freeze pretrained layers, train only a classifier head on our cough spectrograms. With ~1,600 training samples, the model's higher capacity worked against us — it overfit faster than our from-scratch CNN. The pretrained features were too general for the narrow spectrogram patterns that distinguish healthy from abnormal coughs. We reverted to the original architecture.
Data
Standard audio augmentation (noise injection, time stretching) on a dataset this small amplified noise rather than adding signal, and we dropped it. SpecAugment is different and we kept it: it blanks random frequency bands and time windows in the spectrogram rather than distorting the audio, so whatever survives the mask is untouched real recording. It cut seed-to-seed standard deviation from 0.0078 to 0.0025, making single models roughly three times more reproducible.
Training
We train three models with different random seeds and average them, which smooths the variance of a small dataset. Three rather than five because we measured the size curve: members four and five are worth +0.0013 combined, well inside noise, for 40% more inference cost. The threshold is 0.46, chosen on the validation split under a stated policy — catch at least 80% of abnormal coughs, then take the best specificity available. Maximising Youden's J would weight a missed illness the same as a false alarm, which is the wrong tradeoff for a screening tool.