← All work

CMU, 2024

Waste classification on the edge

About half of US municipal waste ended up in landfills in 2018, according to the EPA. We looked at how to train an image classifier that sorts trash into recycling categories and is small enough to run on a cheap device like a Raspberry Pi, and wrote up the practical steps that tend to get skipped.

With Chad Merrill and Nishanth Mohankumar, equal contributors. Carnegie Mellon, 2024.

Final report · 7 pages Download PDF

The problem

The EPA reports 292.4 million short tons of municipal solid waste in the US in 2018, with about 32.1% recycled or composted. Automatic sorting could help, and the most common setup in the literature is a pretrained CNN (a standard image-recognition network) fine-tuned on the TrashNet dataset and deployed on an edge device.

TrashNet has 2,527 photos of worn and damaged items in six classes: glass, paper, cardboard, plastic, metal and trash. We split it 70/15/15 into train, validation and test sets.

Two TrashNet photos: a crushed metal can and a crumpled paper tea packet on white backgrounds.
Two TrashNet examples. Items are photographed on a white background, often worn or damaged.

What we did

We compared three pretrained architectures: ResNet50, EfficientNet-B0 and EfficientNet-B4, each with two new fully connected layers on top for the six classes.

Then we ran ablation studies (switch one thing off at a time and see what changes): freezing versus unfreezing the pretrained layers, which image augmentations to use, and how much training data you need.

Finally we quantized the trained models, meaning we stored their weights as 8-bit integers (qint8) instead of 32-bit floats. That uses a quarter of the memory, which matters on a small device. We ran the classifier on a Raspberry Pi 5 with a live camera feed.

Horizontal bar chart of validation accuracy with each augmentation removed in turn; removing random rotation scores highest.
Augmentation ablation. Leaving out random rotation or random crop gave the best validation accuracy.
Horizontal bar chart of validation accuracy for 25%, 50%, 75% and the full training set.
Dataset-size ablation. More data helps, with diminishing returns past 75%.

Results

Unfreezing mattered most. Letting all of ResNet50's weights train raised validation accuracy from 73% to 89% in the ablation. Frozen models got confused by leftovers from ImageNet pretraining: bottle-shaped metal items were often called plastic.

Dropping random rotation and random crop improved accuracy, likely because TrashNet photos are centered on plain backgrounds and those transforms zoom in on background. Accuracy grew roughly logarithmically with data, with small gains between 75% and 100% of the training set.

Final test accuracy: 96.58% for unfrozen ResNet50 and 96.10% for EfficientNet-B0, which has about a fifth of the parameters (5.3M versus 25.6M) and needed 0.925 TFLOP-hours to train against 4.26. Quantization barely changed accuracy: 96.54% for ResNet50 and 95.76% for B0.

Table of parameters, training compute and test accuracy for ResNet50, EfficientNet-B0 and EfficientNet-B4 in unfrozen, frozen and quantized versions.
Final models compared. EfficientNet-B0 matches ResNet50 with about a fifth of the parameters.
Three webcam frames labeled by the classifier: a plastic bottle, a metal Pepsi can, and a paper food box.
Live inference on the Raspberry Pi 5 camera feed.

What I took away

A smaller, newer architecture can match a bigger one. If you want to freeze the pretrained layers, a larger model like EfficientNet-B4 held up best (91.03% frozen). If you unfreeze, EfficientNet-B0 gave the best mix of accuracy and training cost.

High test accuracy on TrashNet is not the same as working in a cafeteria bin. The photos have a uniform background and lighting, so the next step would be testing on more varied data and measuring inference time on the Pi itself, which we did not get to.

Tools PyTorch, ResNet50, EfficientNet, TrashNet, Raspberry Pi 5

Report

Loading…