Repository logo
Log In(current)
  1. Home
  2. Unitus Open Access
  3. Tesi di Dottorato di Ricerca
  4. Archivio delle tesi di dottorato di ricerca
  5. HPC technologies for the development of neural networks and analysis of big data in chemistry, biology, agrifood, forestry and ecology

HPC technologies for the development of neural networks and analysis of big data in chemistry, biology, agrifood, forestry and ecology

Author(s)
Camerlingo, Armando
Date Issued
July 11, 2025
Type
Doctoral Thesis
Abstract
Advancements in Machine Learning (ML) and the availability of Big Data opened new opportunities in scientific research but also gave birth to new challenges in how to fully exploit these new technologies by requiring access to top class computing systems and algorithms. High Performance Computing (HPC) systems have the capabilities to tackle these issues by leveraging their infrastructure and computing environment to manage big datasets and run optimally ML algorithms. New software tools are also being developed to help researchers create and run modern models on HPC systems by using all the available resources, such as thousands of nodes, GPUs and high performance/low bandwidth storage and network subsystems. This thesis can be divided into two main parts. The first part is dedicated to the activities performed on the HPC cluster at DIBAF (Dipartimento per la Innovazione nei sistemi Biologici, Agroalimentari e Forestali) with the scope of reorganization and preparation to host new services and softwares, including Machine Learning Development Environment (MLDE), a recent tool by Hewlett Packard Enterprise to create and run ML models. In the second part of the thesis, applications of Neural Networks on datasets of different origins are described. In many scientific fields, it is fundamental to obtain results with high precision to have significance in the research, so the ML models must run in double precision representation at the cost of increasing computing time and resources. Parameters and hyperparameters’ selection can lessen the effort to train neural networks while retaining double precision and this approach have been implemented in models created to anlayse three different kind of datasets of interest to the DIBAF department. Moreover, in this work it is also highlighted the importance of selecting the model hyperparameters to improve the accuracy of the prediction. The exploration of the possible choices of hyperparameters for the created models is performed using MLDE, achieving results with an accuracy suitable to be meaningful for the applications analysed during the PhD.
I progressi nel Machine Learning (ML) e la disponibilità di Big Data ha aperto nuove opportunità nella ricerca scientifica, ma insieme a loro sono nate nuove sfide sul come sfruttare al massimo queste nuove tecnologie facendo leva sui migliori sistemi di calcolo e algoritmi. I sistemi di calcolo ad alte prestazioni (HPC) hanno le capacità di risolvere questi problemi facendo leva sulla loro infrastruttura e ambiente di calcolo per gestire grandi datasets ed eseguire efficientemente algoritmi di ML. Inoltre si stanno sviluppando nuovi strumenti software per aiutare i ricercatori a creare e far girare i modelli su sistemi HPC, usando tutte le risorse disponibili, come migliaia di nodi, gpu e sottosistemi di storage e network ad alta prestazione/bassa bandwidth. Questa tesi si può dividere in 2 parti principali. Nella prima parte sono descritte le attività svolte sul cluster HPC del DIBAF (Dipartimento per la Innovazione nei sistemi Biologici, Agroalimentari e Forestali) con lo scopo di ri-organizzarlo e prepararlo per ospitare nuovi servizi e software, incluso Machine Learning Development Environment (MLDE), uno strumento recente di Hewlett Packard Enterprise per creare e eseguire modelli di ML. Nella seconda parte della tesi sono presentate delle applicazioni di Reti Neurali su dataset di origini diverse. In molti ambiti scientifici, è fondamentale ottenere risultati con alta precisione per avere significanza nella ricerca, quindi i modelli di ML devono essere eseguiti a precisione doppia al costo di richiedere più tempo di calcolo e risorse computazionali. La selezione di parametri e iperparamentri può diminuire lo sforzo nel training di una rete neurale, conservando la precisione doppia. Inoltre, si mette in risalto l’importanza della scelta degli iperparametri del modello per migliorare l’accuratezza delle predizioni. L’esplorazione delle possibili combinazioni di iperparametri per i modelli creati è stata eseguita usando MLDE, ottenendo risultati con un’accuratezza adatta per essere significante per le applicazioni analizzate durante il dottorato.
Additional information
Dottorato di ricerca in Scienze, Tecnologie e Biotecnologie per la Sostenibilità
Subjects

HPC

Machine learning

Neural networks

Reti neurali

CHEM-03/A

Handle
https://dspace.unitus.it/handle/2067/71458
File(s)
Thumbnail Image
Name

acamerlingo_tesid.pdf

Size

10.23 MB

Format

Adobe PDF

Checksum (MD5)

adf1e1f783401c1d91054f7de4e99344

Metrics

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify