Università di Genova logo, link al sitoUniRe logo, link alla pagina iniziale
    • English
    • italiano
  • English 
    • English
    • italiano
  • Login
View Item 
  •   DSpace Home
  • Tesi
  • Tesi di Laurea
  • Laurea Magistrale
  • View Item
  •   DSpace Home
  • Tesi
  • Tesi di Laurea
  • Laurea Magistrale
  • View Item
JavaScript is disabled for your browser. Some features of this site may not work without it.

"Compressione e implementazione efficienti di modelli linguistici specifici di dominio su sistemi embedded con risorse limitate: un caso di studio sull'assistente della lavatrice"

Thumbnail
View/Open
tesi38110920.pdf (6.053Mb)
allegato381109201.pdf (417.8Kb)
Author
Khakpour, Amin <1996>
Date
2026-09-04
Data available
2026-09-10
Abstract
Here is the translation of the abstract into Italian, keeping the technical terminology accurate for the computer science and machine learning context: L'implementazione di modelli linguistici generativi su hardware embedded privo di un'unità di elaborazione neurale (NPU) dedicata rimane limitata dalla larghezza di banda della memoria, dal throughput computazionale e dai limiti termici. Questa tesi indaga se, e in quali condizioni architetturali, un assistente conversazionale specifico per un dominio possa essere compresso e implementato su tale piattaforma, utilizzando un agente consulente per lavatrici come caso di studio e il microprocessore STMicroelectronics STM32MP251/253 (Arm Cortex-A35, 4 GB DDR4, nessuna NPU integrata) come target di implementazione. Il lavoro si articola in tre fasi. La distillazione della conoscenza (knowledge distillation) trasferisce il comportamento da un modello "teacher" SmolLM2-1.7B a un modello "student" SmolLM2 da 135 milioni di parametri, stabilendo che la distillazione cross-family tra architetture di tokenizer divergenti non è praticabile. Il fine-tuning supervisionato tramite Low-Rank Adaptation allinea successivamente la struttura conversazionale del modello student rimanendo sotto un limite massimo di 4 GB di VRAM. Infine, il modello unito viene convertito in GGUF, quantizzato a una precisione intera a 8 bit (143 MB) e implementato sulla scheda con un livello di retrieval in Python su un backend di inferenza nativo llama.cpp. Quattro fasi della pipeline sono state valutate su 60 domande specifiche del dominio utilizzando un protocollo "LLM-as-judge". Né la distillazione né il fine-tuning hanno migliorato l'accuratezza fattuale, con la correttezza che è rimasta a 0,018 alla scala dei 135M. L'integrazione del retrieval (retrieval augmentation) ha innalzato la correttezza a 0,301 e ha ridotto il tasso di allucinazione da 0,810 a 0,590 senza alcun riaddestramento; il sistema implementato ha mantenuto 4,91 token al secondo sul
 
Deploying generative language models on embedded hardware without a dedicated neural processing unit remains constrained by memory bandwidth, computational throughput, and thermal limits. This thesis investigates whether, and under what architectural conditions, a domain-specific conversational assistant can be compressed and deployed onto such a platform, using a washing machine advisory agent as the case study and the STMicroelectronics STM32MP251/253 microprocessor (Arm Cortex-A35, 4 GB DDR4, no integrated NPU) as the deployment target. The work proceeds in three phases. Knowledge distillation transfers behaviour from a SmolLM2-1.7B teacher into a 135M-parameter SmolLM2 student, establishing that cross-family distillation across divergent tokenizer architectures is not viable. Supervised fine-tuning with Low-Rank Adaptation then aligns the student's conversational structure under a 4 GB VRAM ceiling. Finally, the merged model is converted to GGUF, quantized to 8-bit integer precision (143 MB), and deployed to the board with a Python retrieval layer over a native llama.cpp inference backend. Four pipeline stages were evaluated across 60 domain-specific questions using an LLM-as-judge protocol. Neither distillation nor fine-tuning improved factual accuracy, with correctness remaining at 0.018 at the 135M scale. Retrieval augmentation raised correctness to 0.301 and reduced the hallucination rate from 0.810 to 0.590 without retraining, and the deployed system sustained 4.91 tokens per second on the target hardware, returning a complete grounded response in approximately 9.7 seconds at a first-token latency of 5.7 seconds while occupying 519 MB (13.8%) of available system memory.
 
Type
info:eu-repo/semantics/masterThesis
Collections
  • Laurea Magistrale [8051]
URI
https://unire.unige.it/handle/123456789/16751
Metadata
Show full item record

UniRe - Università degli studi di Genova | Information and Contacts
 

 

All of DSpaceCommunities & Collections

My Account

Login

UniRe - Università degli studi di Genova | Information and Contacts