Design and Deployment of Multi-Class Image Classification Model on ESP32 Using a Lightweight CNN Approach
DOI:
https://doi.org/10.57041/p6f08552Keywords:
ESP32, Convolutional Neural Network (CNN), TensorFlow Lite, Post-Training Quantization, Multi-class Image Classification, TinyMLAbstract
Tremendous progress has been made in developing intelligent and autonomous systems in recent years, driven by AI software tools such as Chatbots and large language models (LLMs) for various applications. This has led to the adoption of smart technology solutions. Nonetheless, computer vision requires significant computational power, making it difficult to run on devices such as the ESP32. To showcase the development of smart technologies leveraging controller platforms, the research conducted training and deployment of an optimal multi-class image classification model on the ESP32 microcontroller platform, ensuring acceptable predictive performance across four classes: person, fruit, car, and unknown. It is important to clearly distinguish between the proposed system, which performs image classification and object detection. Regarding model training, the images were preprocessed to minimise computational cost. Nonetheless, the trained model achieved a peak validation accuracy of 99.34±1% when evaluated using five-fold cross-validation, with an average validation accuracy of 98±1% without altering the model or the training process, producing similar results. Moreover, TFLite Quantisation, a variant of the INT8 quantisation technique, was used to optimise the model size. The model was converted to C array format while retaining most features. To improve accuracy during training, the training dataset was augmented with large, high-quality images that had been filtered out. The chosen hardware used for this project was the ESP32-S3 N16R8 microcontroller with built-in OV2640 camera and the 3.2'' TFT to visualize the process of computer vision task execution in real time. The model works smoothly and predicts accurately, even with just a few samples per frame, to decrease response time latency. Successfully running the multi-class image classification model on the ESP32 microcontroller is the evidence of full independence of embedded vision solution without having anything to do with cloud technology.Downloads
Published
2026-09-09
Issue
Section
Emerging Trends in Artificial Intelligence, Multidisciplinary Engineering, Health and Smart Technologies
License
Copyright (c) 2026 https://grsh.org/journal1/index.php/ijeet/cr

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Design and Deployment of Multi-Class Image Classification Model on ESP32 Using a Lightweight CNN Approach. (2026). International Journal of Emerging Engineering and Technology, 5(1-2), 16-27. https://doi.org/10.57041/p6f08552