Back to screening

How SickNote works

From raw audio to screening result in under 3 seconds

Overview

SickNote is a binary cough classifier that distinguishes healthy coughs from abnormal ones. It converts audio recordings into mel spectrograms — visual representations of sound frequencies — and feeds them through an ensemble of three convolutional neural networks. Each prediction includes a Grad-CAM heatmap showing which spectrogram regions drove the result.

Dataset

The COUGHVID dataset contains ~34,400 crowdsourced cough recordings. Of these, 2,841 were reviewed by four expert physicians who provided diagnostic labels.

34,400

Total recordings

2,841

Expert-labeled

2,267

After filtering

Filtering criteria

  • Cough detection confidence > 0.8
  • At least one expert diagnosis present
  • Quality rated acceptable by majority of experts

Class distribution

78% abnormal
22%

Processing pipeline

Raw Audio
Filter
Label
Spectrogram
Normalize
Split

Model architecture

A small convolutional neural network trained from scratch, 71,793 parameters per model. Three models are trained with different random seeds and their predictions are averaged. Conv features are pooled over the time axis only — never over frequency, since 200 Hz and 2 kHz are genuinely different sounds while a cough at 2s and a cough at 5s are the same cough. Each result includes a Grad-CAM heatmap showing which spectrogram regions influenced the classification.

Input: (1, 64, T) mel spectrogram

  ↓

  Conv2d(8) + BatchNorm + ReLU + MaxPool

  Conv2d(16) + BatchNorm + ReLU + MaxPool

  Conv2d(32) + BatchNorm + ReLU + MaxPool

  ↓

  Flatten

  Linear(128) + ReLU + Dropout(0.5)

  Linear(1) → logit

Loss function

BCEWithLogitsLoss + pos_weight

Optimizer

Adam (lr=3e-4, wd=1e-4)

Ensemble

3 models, different seeds, averaged probabilities

Explainability

Grad-CAM heatmaps on last conv block

Evaluation metrics

MetricActualReference
AUC-ROC0.72380.6556 from clip duration and loudness alone
Improvement over previous+0.030895% CI [+0.0137, +0.0494], better on 5 of 5 folds
Sensitivity0.8029policy floor is 0.80
Specificity0.5054at the same operating point
Rater agreementκ 0.181four physicians on the same audio — the ceiling

Classification threshold is 0.46, selected on the validation split under a stated policy: catch at least 80% of abnormal coughs, then take the best specificity available. An earlier version shipped 0.52, which was the argmax of Youden's J computed on the test set — a value fitted to the data it was then reported on. Figures above come from nested 5-fold cross-validation, in which all 2,267 clips are scored by a model that never trained on them.

Known limitations

  • All COUGHVID recordings are voluntary intentional coughs
  • 2,267 expert-labeled samples, split into ~1,600 for training — small by production ML standards
  • Class imbalance: ~78% abnormal / ~22% healthy after expert filtering
  • No external validation dataset — generalization to new devices unknown
  • Binary only — does not distinguish COVID vs URTI vs LRTI vs other
  • COUGHVID was collected during the COVID pandemic — label distribution reflects that context
  • Screening tool only — not a diagnostic

Design decisions

The engineering choices behind SickNote

Architecture

Why binary classification

Multi-class classification (COVID, URTI, LRTI, obstructive disease) dropped accuracy sharply because pathological coughs occupy overlapping feature space in the spectrogram domain. Binary classification — healthy vs. abnormal — is medically honest and performs significantly better with limited data. The output is "something sounds off," not a specific diagnosis.

Architecture

Why we built a CNN from scratch

With only 2,267 expert-labeled samples, a small 3-layer CNN with aggressive dropout (0.5) and reduced channel widths [8, 16, 32] was the right fit. Larger architectures overfit before learning useful features. We tuned every hyperparameter against real dataset statistics from explore.py before writing a single training loop.

Training

Why we tried transfer learning and reverted

We attempted to replace our CNN with a pretrained ResNet18 backbone to leverage features learned from millions of ImageNet images. The plan: freeze pretrained layers, train only a classifier head on our cough spectrograms. With ~1,600 training samples, the model's higher capacity worked against us — it overfit faster than our from-scratch CNN. The pretrained features were too general for the narrow spectrogram patterns that distinguish healthy from abnormal coughs. We reverted to the original architecture.

Data

Why waveform augmentation failed but masking works

Standard audio augmentation (noise injection, time stretching) on a dataset this small amplified noise rather than adding signal, and we dropped it. SpecAugment is different and we kept it: it blanks random frequency bands and time windows in the spectrogram rather than distorting the audio, so whatever survives the mask is untouched real recording. It cut seed-to-seed standard deviation from 0.0078 to 0.0025, making single models roughly three times more reproducible.

Training

Why three models, and why this threshold

We train three models with different random seeds and average them, which smooths the variance of a small dataset. Three rather than five because we measured the size curve: members four and five are worth +0.0013 combined, well inside noise, for 40% more inference cost. The threshold is 0.46, chosen on the validation split under a stated policy — catch at least 80% of abnormal coughs, then take the best specificity available. Maximising Youden's J would weight a missed illness the same as a false alarm, which is the wrong tradeoff for a screening tool.