Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Compressing Convolutional Neural Networks for Offline Educational Deployment An Empirical Study of Pruning and INT8 Quantization

Domain:

education

Record type:

paper
Creator:
KarGup
Publisher:
Zenodo
Host:avatar
Access to artificial intelligence tools in rural and low-resource educational settings is fundamentally constrained by hardware: most schools have, at best, low-end smartphones or shared desktop computers, and rarely have internet connectivity reliable enough for cloud-based inference. This paper presents a fully reproducible empirical study of two standard model-compression techniques—magnitude-based weight pruning and post-training 8-bit (INT8) quantization—applied to a compact convolutional neural network (CNN) trained for handwritten digit recognition on the MNIST dataset. Using a controlled experimental design, we measure the accuracy, model size, and compression ratio of four deployment configurations. These results provide a transparent, low-cost, and entirely reproducible reference point for deploying handwriting-recognition tools on inexpensive edge hardware without requiring GPUs or cloud services. Methodology This paper presents a fully reproducible empirical study of two standard model-compression techniques: Magnitude-based weight pruning (targeting 70% structural sparsity) Post-training 8-bit (INT8) quantization These techniques were applied to a compact convolutional neural network (CNN) trained for handwritten digit recognition on the MNIST dataset. Using a strictly controlled experimental design, an independently cloned copy of a trained baseline model was pruned and fine-tuned to isolate the exact effects of compression without shared-state artifacts. Key Findings The empirical results isolate the distinct architectural and physical impacts of both techniques under standard deployment tooling (TensorFlow Lite): The Baseline: The FP32 CNN achieved 98.78% test accuracy with a 224,892-byte physical footprint. The Quantization Impact: INT8 quantization reduced the footprint to 61,800 bytes (a 3.64x reduction) with a negligible 0.09 percentage-point shift in accuracy. The Pruning Impact: Iterative pruning to 70% structural sparsity matched the baseline almost perfectly at 98.77%, but left the physical flatbuffer file size unchanged due to standard dense storage formatting. The Combined Approach: Combining both techniques perfectly restored accuracy to 98.78% while maintaining the 3.64x physical compression. Conclusion These results offer practical, evidence-based guidance for institutions in low-resource settings considering deploying offline AI tools on inexpensive edge hardware, demonstrating that INT8 quantization acts as a near-cost-free first step for deployment.

Visit

doi.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2026 Adityajyoti Karhttp://rightsstatements.org/vocab/InC/1.0/

Similar

Optimizing Weather Forecasting Accuracy via Radial Basis Function Networks, Convolutional Neural Networks and Convolutional Neural NetworksEfficient Convolutional Neural Networks for Diacritic RestorationVisualizing and Comparing Convolutional Neural NetworksTime Gated Convolutional Neural Networks for Crop ClassificationThe Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine TranslationApplication of convolutional neural networks for discriminating mining blasts and earthquakes

Optimizing Weather Forecasting Accuracy via Radial Basis Function Networks, Convolutional Neural Networks and Convolutional Neural Networks

Weather forecasting is crucial for various sectors, including agriculture, disaster management, and

Efficient Convolutional Neural Networks for Diacritic Restoration

Diacritic restoration has gained importance with the growing need for machines to understand written

Visualizing and Comparing Convolutional Neural Networks

Convolutional Neural Networks (CNNs) have achieved comparable error rates to well-trained human on I

Time Gated Convolutional Neural Networks for Crop Classification

This paper presented a state-of-the-art framework, Time Gated Convolutional Neural Network (TGCNN) t

The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation

A “bigger is better” explosion in the number of parameters in deep neural networks has made it increasingly challenging to make state-of-the-art networks accessible in compute-restricted environments. Compression techniques have taken on renewed importance as a way

Application of convolutional neural networks for discriminating mining blasts and earthquakes

Earthquake source is among the key element within the framework of se