HPC technologies for the development of neural networks and analysis of big data in chemistry, biology, agrifood, forestry and ecology
Author(s)
Camerlingo, Armando
Date Issued
July 11, 2025
Type
Doctoral Thesis
Abstract
Advancements in Machine Learning (ML) and the availability of Big Data opened new
opportunities in scientific research but also gave birth to new challenges in how to fully
exploit these new technologies by requiring access to top class computing systems and
algorithms. High Performance Computing (HPC) systems have the capabilities to tackle
these issues by leveraging their infrastructure and computing environment to manage big
datasets and run optimally ML algorithms. New software tools are also being developed to
help researchers create and run modern models on HPC systems by using all the available
resources, such as thousands of nodes, GPUs and high performance/low bandwidth storage
and network subsystems. This thesis can be divided into two main parts. The first part is
dedicated to the activities performed on the HPC cluster at DIBAF (Dipartimento per la
Innovazione nei sistemi Biologici, Agroalimentari e Forestali) with the scope of reorganization
and preparation to host new services and softwares, including Machine
Learning Development Environment (MLDE), a recent tool by Hewlett Packard Enterprise to
create and run ML models. In the second part of the thesis, applications of Neural Networks
on datasets of different origins are described. In many scientific fields, it is fundamental to
obtain results with high precision to have significance in the research, so the ML models
must run in double precision representation at the cost of increasing computing time and
resources. Parameters and hyperparameters’ selection can lessen the effort to train neural
networks while retaining double precision and this approach have been implemented in
models created to anlayse three different kind of datasets of interest to the DIBAF
department. Moreover, in this work it is also highlighted the importance of selecting the
model hyperparameters to improve the accuracy of the prediction. The exploration of the
possible choices of hyperparameters for the created models is performed using MLDE,
achieving results with an accuracy suitable to be meaningful for the applications analysed
during the PhD.
I progressi nel Machine Learning (ML) e la disponibilità di Big Data ha aperto nuove
opportunità nella ricerca scientifica, ma insieme a loro sono nate nuove sfide sul come
sfruttare al massimo queste nuove tecnologie facendo leva sui migliori sistemi di calcolo e
algoritmi. I sistemi di calcolo ad alte prestazioni (HPC) hanno le capacità di risolvere questi
problemi facendo leva sulla loro infrastruttura e ambiente di calcolo per gestire grandi
datasets ed eseguire efficientemente algoritmi di ML. Inoltre si stanno sviluppando nuovi
strumenti software per aiutare i ricercatori a creare e far girare i modelli su sistemi HPC,
usando tutte le risorse disponibili, come migliaia di nodi, gpu e sottosistemi di storage e
network ad alta prestazione/bassa bandwidth. Questa tesi si può dividere in 2 parti
principali. Nella prima parte sono descritte le attività svolte sul cluster HPC del DIBAF
(Dipartimento per la Innovazione nei sistemi Biologici, Agroalimentari e Forestali) con lo
scopo di ri-organizzarlo e prepararlo per ospitare nuovi servizi e software, incluso Machine
Learning Development Environment (MLDE), uno strumento recente di Hewlett Packard
Enterprise per creare e eseguire modelli di ML. Nella seconda parte della tesi sono
presentate delle applicazioni di Reti Neurali su dataset di origini diverse. In molti ambiti
scientifici, è fondamentale ottenere risultati con alta precisione per avere significanza nella
ricerca, quindi i modelli di ML devono essere eseguiti a precisione doppia al costo di
richiedere più tempo di calcolo e risorse computazionali. La selezione di parametri e
iperparamentri può diminuire lo sforzo nel training di una rete neurale, conservando la
precisione doppia. Inoltre, si mette in risalto l’importanza della scelta degli iperparametri
del modello per migliorare l’accuratezza delle predizioni. L’esplorazione delle possibili
combinazioni di iperparametri per i modelli creati è stata eseguita usando MLDE, ottenendo
risultati con un’accuratezza adatta per essere significante per le applicazioni analizzate
durante il dottorato.
Additional information
Dottorato di ricerca in Scienze, Tecnologie e Biotecnologie per la Sostenibilità
File(s)![Thumbnail Image]()
Name
acamerlingo_tesid.pdf
Size
10.23 MB
Format
Adobe PDF
Checksum (MD5)
adf1e1f783401c1d91054f7de4e99344
