This thesis explores the performance analysis and comparison of quantized language models (LLMs) in resource-constrained environments, with a specific focus on supply chain risk data. As large language models continue to expand, there is a growing demand for compact AI solutions that work effectively with limited memory and reduced power consumption, particularly in hardware-limited settings. This research addresses these challenges by aiming to develop lightweight AI models that maintain high accuracy while minimizing size, execution speeds, and energy consumption.
The study investigates the use of WasmEdge, a WebAssembly runtime, as a promising solution for executing lightweight language models beyond traditional browser environments. This approach has the potential to unlock the use of LLMs in resource-constrained environments.
The research presents a novel technique for deploying quantized LLMs to generate inferences using supply chain risk management data. The findings reveal a strong correlation between inference quality, duration, and performance metrics. However, shorter inference times and enhanced performance efficiency do not always result in the best inference quality. This study provides valuable insights for developing efficient and accessible AI models, contributing to advancements in AI-assisted supply chain risk management applications and similar domains.