Design and Deployment of Multi-Class Image Classification Model on ESP32 Using a Lightweight CNN Approach

Authors

  • Abdul Hannan Department of Electrical Engineering, Superior University, Lahore, Pakistan Author
  • Rao M. Asif Department of Electrical Engineering, Superior University, Lahore, Pakistan Author

DOI:

https://doi.org/10.57041/p6f08552

Keywords:

ESP32, Convolutional Neural Network (CNN), TensorFlow Lite, Post-Training Quantization, Multi-class Image Classification, TinyML

Abstract

Tremendous progress has been made in developing intelligent and autonomous systems in recent years, driven by AI software tools such as Chatbots and large language models (LLMs) for various applications. This has led to the adoption of smart technology solutions. Nonetheless, computer vision requires significant computational power, making it difficult to run on devices such as the ESP32. To showcase the development of smart technologies leveraging controller platforms, the research conducted training and deployment of an optimal multi-class image classification model on the ESP32 microcontroller platform, ensuring acceptable predictive performance across four classes: person, fruit, car, and unknown. It is important to clearly distinguish between the proposed system, which performs image classification and object detection. Regarding model training, the images were preprocessed to minimise computational cost. Nonetheless, the trained model achieved a peak validation accuracy of 99.34±1% when evaluated using five-fold cross-validation, with an average validation accuracy of 98±1% without altering the model or the training process, producing similar results. Moreover, TFLite Quantisation, a variant of the INT8 quantisation technique, was used to optimise the model size. The model was converted to C array format while retaining most features. To improve accuracy during training, the training dataset was augmented with large, high-quality images that had been filtered out. The chosen hardware used for this project was the ESP32-S3 N16R8 microcontroller with built-in OV2640 camera and the 3.2'' TFT to visualize the process of computer vision task execution in real time. The model works smoothly and predicts accurately, even with just a few samples per frame, to decrease response time latency. Successfully running the multi-class image classification model on the ESP32 microcontroller is the evidence of full independence of embedded vision solution without having anything to do with cloud technology.

Downloads

Published

2026-09-09

Issue

Section

Emerging Trends in Artificial Intelligence, Multidisciplinary Engineering, Health and Smart Technologies

How to Cite

Design and Deployment of Multi-Class Image Classification Model on ESP32 Using a Lightweight CNN Approach. (2026). International Journal of Emerging Engineering and Technology, 5(1-2), 16-27. https://doi.org/10.57041/p6f08552

Similar Articles

1-10 of 34

You may also start an advanced similarity search for this article.