Waste classification on the edge
About half of US municipal waste ended up in landfills in 2018, according to the EPA. We looked at how to train an image classifier that sorts trash into recycling categories and is small enough to run on a cheap device like a Raspberry Pi, and wrote up the practical steps that tend to get skipped.
The problem
The EPA reports 292.4 million short tons of municipal solid waste in the US in 2018, with about 32.1% recycled or composted. Automatic sorting could help, and the most common setup in the literature is a pretrained CNN (a standard image-recognition network) fine-tuned on the TrashNet dataset and deployed on an edge device.
TrashNet has 2,527 photos of worn and damaged items in six classes: glass, paper, cardboard, plastic, metal and trash. We split it 70/15/15 into train, validation and test sets.
What we did
We compared three pretrained architectures: ResNet50, EfficientNet-B0 and EfficientNet-B4, each with two new fully connected layers on top for the six classes.
Then we ran ablation studies (switch one thing off at a time and see what changes): freezing versus unfreezing the pretrained layers, which image augmentations to use, and how much training data you need.
Finally we quantized the trained models, meaning we stored their weights as 8-bit integers (qint8) instead of 32-bit floats. That uses a quarter of the memory, which matters on a small device. We ran the classifier on a Raspberry Pi 5 with a live camera feed.
Results
Unfreezing mattered most. Letting all of ResNet50's weights train raised validation accuracy from 73% to 89% in the ablation. Frozen models got confused by leftovers from ImageNet pretraining: bottle-shaped metal items were often called plastic.
Dropping random rotation and random crop improved accuracy, likely because TrashNet photos are centered on plain backgrounds and those transforms zoom in on background. Accuracy grew roughly logarithmically with data, with small gains between 75% and 100% of the training set.
Final test accuracy: 96.58% for unfrozen ResNet50 and 96.10% for EfficientNet-B0, which has about a fifth of the parameters (5.3M versus 25.6M) and needed 0.925 TFLOP-hours to train against 4.26. Quantization barely changed accuracy: 96.54% for ResNet50 and 95.76% for B0.
What I took away
A smaller, newer architecture can match a bigger one. If you want to freeze the pretrained layers, a larger model like EfficientNet-B4 held up best (91.03% frozen). If you unfreeze, EfficientNet-B0 gave the best mix of accuracy and training cost.
High test accuracy on TrashNet is not the same as working in a cafeteria bin. The photos have a uniform background and lighting, so the next step would be testing on more varied data and measuring inference time on the Pi itself, which we did not get to.