Marullo Giorgia

Ricercatore TD(A)


Politecnico di Torino
giorgia.marullo@polito.it

Sito istituzionale
SCOPUS ID: 57216456257

Publications
Updated to September 03, 2026

[1] Ulrich L., Innocente C., Marullo G., Audisio A., Aprato A., Massè A., Moos S., Vezzetti E., A mixed reality framework for interpretable and explainable joint replacement assessment. International Journal of Medical Informatics, 213 (2026).
Mostra Abstract

Abstract: Objective: Joint replacement surgery, also known as arthroplasty, is a common procedure that restores mobility and relieves pain in patients with severe joint pathologies. Despite being considered routine, arthroplasties are complex interventions with potential complications and variable clinical outcomes. Accurate evaluation of replaced joint mobility to ensure implant stability within the patient's functional range of motion (ROM) is a major challenge in postoperative care. However, the reliability of current assessment methods is limited due to their lack of standardized and quantitative tools. This study presents a patient-specific Mixed Reality (MR) framework designed to enhance postoperative evaluation in joint replacement with a focus on total hip arthroplasty (THA). Methods: The proposed system enables objective quantification and MR visualization of prosthesis biomechanics by integrating ROM simulation and 3D modeling, promoting explainability and interpretability of surgery outcomes. A retrospective analysis of 67 THAs was performed to compare simulated ROM results with clinical assessments and literature benchmarks. Additionally, surgeons evaluated the system's clinical relevance and usability through a preliminary study, including completion of the System Usability Scale (SUS). Results: Simulated ROM measurements showed good agreement with both clinical assessments and established literature reference values across ten movements commonly examined in orthopedic practice. The MR tool demonstrated high accuracy, repeatability, and potential to support postoperative decision-making, with usability testing yielding a favorable median SUS score of 82.5, indicating strong acceptance among clinicians. Conclusion: The patient-specific MR framework provides a reliable, quantitative, and interpretable method for assessing prosthetic joint performance after replacement, supporting its integration into postoperative workflows for improved surgical outcome assessment.

Keywords: Computer-aided surgery | Human-computer interaction | Joint replacement | Mixed reality | Orthopedic surgery | Total Hip arthroplasty

[2] Salerno F., Ulrich L., Marullo G., Moos S., Vezzetti E., Advancing surgical cutting guide flexibility: A hybrid physical and Augmented Reality solution for cranio-maxillofacial surgery. Journal of Cranio Maxillofacial Surgery, 54(6) (2026).
Mostra Abstract

Abstract: Cutting guides are widely used in cranio-maxillofacial surgery by providing mechanical references for precise bone resections and reducing intraoperative variability. Nevertheless, their rigid and patient-specific design requires dedicated CAD modeling and fabrication, making them time-consuming to produce and difficult to adapt when anatomical conditions or bone surfaces change. This work presents a hybrid surgical cutting guide that combines a physically adjustable device with Augmented Reality (AR) feedback to support intraoperative alignment in maxillofacial osteotomies. The concept merges the tactile reliability of conventional guides with the adaptability of digital visualization, enabling surgeons to fine-tune the cutting plane directly through AR tracking. Registration was performed using cephalometric landmarks and an inside-out tracking approach with HoloLens 2, allowing precise superimposition of virtual cutting planes onto 3D-printed mandibular models. The system was evaluated by both expert and novice operators under three feedback conditions: no AR, holographic overlay, and real-time distance guidance. Results showed that AR feedback was associated with improved positional accuracy, with mean linear deviations of 1.27±0.71mm and angular errors of 4.46±3.27°. Operator experience influenced overall performance, yet enhanced feedback compensated part of this variability. Combining physical and digital guidance can yield more adaptable, precise, and reusable osteotomy tools, paving the way for flexible surgical assistance in clinical settings.

Keywords: Adjustable surgical guide | Augmented Reality | Cranio-maxillofacial surgery | Hololens 2 | Hybrid surgical guide | Pose estimation

[3] Ulrich L., De Luca A., Miraglia R., Mulassano E., Quattrocchio S., Marullo G., Innocente C., Salerno F., Vezzetti E., A 3D Camera-Based Approach for Real-Time Hand Configuration Recognition in Italian Sign Language. Sensors, 26(3) (2026).
Mostra Abstract

Abstract: Deafness poses significant challenges to effective communication, particularly in contexts where access to sign language interpreters is limited. Hand configuration recognition represents a fundamental component of sign language understanding, as configurations constitute a core cheremic element in many sign languages, including Italian Sign Language (LIS). In this work, we address configuration-level recognition as an independent classification task and propose a machine vision framework based on RGB-D sensing. The proposed approach combines MediaPipe-based hand landmark extraction with normalized three-dimensional geometric features and a Support Vector Machine classifier. The first contribution of this study is the formulation of LIS hand configuration recognition as a standalone, configuration-level problem, decoupled from temporal gesture modeling. The second contribution is the integration of sensor-acquired RGB-D depth measurements into the landmark-based feature representation, enabling a direct comparison with estimated depth obtained from monocular data. The third contribution consists of a systematic experimental evaluation on two LIS configuration sets (6 and 16 classes), demonstrating that the use of real depth significantly improves classification performance and class separability, particularly for geometrically similar configurations. The results highlight the critical role of depth quality in configuration-level recognition and provide insights into the design of robust vision-based systems for LIS analysis.

Keywords: Italian Sign Language (LIS) | machine learning | MediaPipe | RGB-D cameras | sign-language recognition

[4] Ruggieri R., Marullo G., Grandvalet Y., Moos S., Vezzetti E., Ulrich L., Assessing Physical Ergonomics in Industry 5.0: A Preliminary Deep Learning-Based Approach. Lecture Notes in Mechanical Engineering, 130-141 (2026).
Mostra Abstract

Abstract: Musculoskeletal disorders are frequent workplace injuries, especially during manual lifting activities. They are influenced by posture, lifting technique, and repetitive movements. Various ergonomic assessment methods exist, but each has limitations: observational methods can be slow and prone to error, while contact sensor-based methods, although more accurate, tend to be invasive and expensive. Recent developments have focused on non-contact sensors, such as RGB and RGB-D cameras, combined with Deep Learning algorithms and observational methods, to improve efficiency and reliability. This study proposes a solution combining a skeleton-based Deep Learning algorithm for Human Pose Estimation with an observational method for postural assessment. Using an RGB camera, four lifting techniques (stoop, squat, semi-squat, and weightlifter) were analyzed, evaluating their impact on worker posture through the REBA score. Among handle-assisted lifts, the stoop and weightlifter techniques showed the lowest average maximum REBA scores (5.375 and 6.125), while the squat and semi-squat techniques scored highest at 7. The semi-squat without handles showed the greatest postural risk (7.875). Future work will integrate 3D data and validate the approach with a larger, more diverse population.

Keywords: Deep learning | Ergonomics | Industry 5.0 | Postural analysis

[5] Marullo G., Amati D., Barbarino S., D’Onofrio M., Sechi T., Innocente C., Vezzetti E., Ulrich L., Three-Dimensional Vision-Based Recognition of Guitar Chords. Computer Music Journal (2026).
Mostra Abstract

Abstract: Artificial intelligence-powered assistants are revolutionizing the music industry by transforming how music is produced and experienced, opening new frontiers for creativity and innovation. Their application in supporting learners to play a musical instrument, however, has yet to be fully explored. Most current applications primarily focus on sound data to distinguish between different notes. This approach excludes the correct hand form, which is essential for learning to play an instrument. Furthermore, sound can be subject to background noise or compression. The current article presents a pipeline for guitar chord recognition based on 3-D data acquisition and color images. The developed method utilizes MediaPipe to estimate hand landmarks from color images and leverages depth images to retrieve real-depth information. In this study, various machine learning (ML) algorithms were compared to perform chord recognition from hand landmark information. The number of chords was limited to four, plus one class for unknown gestures. The classifier was trained with an RGB-D video data set (red, green, and blue plus depth) that features the hands of 18 subjects who performed the selected chords. The random forest (RF) classifier demonstrated remarkable performance in the classification task, achieving a balanced accuracy (BA) that exceeded 85% on unseen data under the same acquisition conditions and improving the state-of-the-art performance. Future developments will focus on expanding the number of supported chords so that the proposed approach could be used in real-world applications. Such applications might include not only guitar education but also transcription, control of sound synthesis, identification in ensembles, and other contexts where audio input alone is insufficient.

Keywords: Audio acoustics | Color | Color image processing | Computer music | Computer vision | Data acquisition | Learning algorithms | Learning systems | Palmprint recognition

[6] Innocente C., Iaconinoto L., Notarangelo D., Scalcione A., Sergi R., Velardi A., Marullo G., Vezzetti E., Ulrich L., An End-to-End Radiomic Framework for Automatic Vertebral Lesion Classification and 3D Visualization. Eng, 7(1) (2026).
Mostra Abstract

Abstract: Early and reliable identification of vertebral metastases on computed tomography remains a major challenge in oncologic imaging due to the morphological complexity of metastatic lesions and the high inter-patient variability of spinal anatomy. In this study, an end-to-end interpretable radiomic-based framework was developed to automatically distinguish healthy from metastatic vertebrae using segmented DICOM data, coupled with an interactive virtual reality (VR) visualization module implemented in Unity 3D. The proposed framework integrates radiomic feature extraction and selection, informed undersampling to address class imbalance, and automatic machine learning-based classification. To facilitate interpretation, patient-specific 3D models with overlapped classifier outputs were integrated into a VR desktop application, enabling advanced exploration of patient-specific spinal models, with color-coded visualization of algorithmic predictions and expert-defined suspicious lesions. The final classification model, trained using a Random Forest algorithm and optimized via stratified 5-fold cross-validation, achieved an overall accuracy of 0.86, an Area Under the Receiver Operating Characteristic Curve of 0.91, and an F1-score of 0.81 for the metastatic class on the independent test set, achieving competitive diagnostic performance while preserving transparency and clinical interpretability. This study represents a foundational step toward intelligent, interactive, and clinically interpretable tools for the diagnosis and follow-up of spinal metastatic disease.

Keywords: automatic lesion classification | bone lesions | machine learning | medical 3D visualization | spinal metastases | virtual reality

[7] Antonaci F., Ciaramella P., Marullo G., Ulrich L., Papa V., Marra W., Depaoli A., Miraglia R., Moos S., Vezzetti E., CardioSmartAssist: A customisable AI framework for echocardiography-based cardiac assessment. Biomedical Signal Processing and Control, 122 (2026).
Mostra Abstract

Abstract: Background and Objective: Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, accounting for approximately 20.5 million deaths annually, nearly one-third of all global deaths. Despite its central role in cardiology, echocardiographic assessment remains subject to significant inter- and intra-observer variability, particularly based on manual frame selection and segmentation. This limitation has driven increasing interest in Deep Learning (DL) solutions capable of enabling more objective, reproducible, and efficient analyses. Methods: This paper introduces CardioSmartAssist, a deep learning framework for automatic left ventricular segmentation and multi-cycle ejection fraction estimation relying solely on echocardiographic videos. The system integrates frame-by-frame segmentation visualisation, anomaly detection, and volume tracking to enhance clinical usability. Moreover, a key feature is its continuous learning mechanism, which allows clinician-corrected segmentations to be stored and used for progressive model refinement. Results: The framework is based on a MultiResUNet architecture, trained on public (EchoNet-Dynamic) and proprietary (CardioSmartSet) datasets, achieving Dice Coefficient Scores of 0.9328 and 0.9189, respectively. On the held-out test set, the EF estimated by the system showed a mean absolute difference of 10% compared with clinically reported EF values, which is lower than the typical inter-operator variability of approximately 13%. Conclusion: CardioSmartAssist resulted to be a promising tool for consistent cardiac evaluations, improving access to diagnostics, and enhancing clinical decision-making through smart assistance.

Keywords: AI-assisted diagnosis | Artificial intelligence | Clinical decision support systems | Computer-aided diagnosis | Deep learning | Left ventricular ejection fraction

[8] Innocente C., Di Pisa G., Lionetti I., Mamoli A., Vitulano M., Marullo G., Maffei S., Vezzetti E., Ulrich L., Learning Italian Hand Gesture Culture Through an Automatic Gesture Recognition Approach. Future Internet, 18(4) (2026).
Mostra Abstract

Abstract: Italian hand gestures constitute a distinctive and widely recognized form of nonverbal communication, deeply embedded in everyday interaction and cultural identity. Despite their prominence, these gestures are rarely formalized or systematically taught, posing challenges for foreign speakers and visitors seeking to interpret their meaning and pragmatic use. Moreover, their ephemeral and embodied nature complicates traditional preservation and transmission approaches, positioning them within the broader domain of intangible cultural heritage. This paper introduces a machine learning–based framework for recognizing iconic Italian hand gestures, designed to support cultural learning and engagement among foreign speakers and visitors. The approach combines RGB–D sensing with depth-enhanced geometric feature extraction, employing interpretable classification models trained on a purpose-built dataset. The recognition system is integrated into a non-immersive virtual reality application simulating an interactive digital totem conceived for public arrival spaces, providing tutorial content, real-time gesture recognition, and immediate feedback within a playful and accessible learning environment. Three supervised machine learning pipelines were evaluated, and Random Forest achieved the best overall performance. Its integration with an Isolation Forest module was further considered for deployment, achieving a macro-averaged accuracy and F1-score of 0.82 under a 5-fold cross-validation protocol. An experimental user study was conducted with 25 subjects to evaluate the proposed interactive system in terms of usability, user engagement, and learning effectiveness, obtaining favorable results and demonstrating its potential as a practical tool for cultural education and intercultural communication.

Keywords: gesture recognition | human–computer interaction | intangible cultural heritage | Italian culture | machine learning | virtual reality

[9] Lo Faro A., Grandvalet Y., Ulrich L., Moos S., Vezzetti E., Marullo G., Whitening black boxes: Interpretable and explainable DL-based systems for trustworthy healthcare. Artificial Intelligence in Medicine, 180 (2026).
Mostra Abstract

Abstract: Artificial intelligence (AI), and specifically deep learning (DL) models, are rapidly gaining traction in healthcare to analyze complex medical images and support clinical decision-making. However, DL models are often considered black boxes due to the lack of a clear explanation when providing predictions. Explainable artificial intelligence (XAI) methods are emerging as an effective way to make models explainable for developers and provide interpretable outputs for clinicians. This review presents a taxonomy of the most widely used XAI methods for image classification, with related benefits and drawbacks. Furthermore, it examines whether the type of classifier affects the choice of an explainability technique and investigates the impact of black boxes on the healthcare environment. The analysis considered papers published between January 2020 and July 2025 in Scopus and Google Scholar, utilizing the PRISMA guidelines to enhance reporting. Sixty-nine papers were identified as suitable for classifying XAI methods in four categories based on backpropagation, perturbation, attention, and concept. The results show increased use of backpropagation-based techniques, which offer simple and intuitive heatmaps. Perturbation-based methods are frequently employed to validate model robustness, but they are computationally expensive. Finally, concept-based and attention-based approaches are less widespread but represent a promising solution towards explanations that align with human semantics and reflect the intrinsic model behavior. Future research should focus on combined approaches and concept methods that generate explanations in the same semantic field as clinicians and are computationally suitable for healthcare environments, paving the way for transparent and clinically reliable DL systems.

Keywords: Black boxes | Deep learning | Explainability | Explainable Artificial Intelligence (XAI) | Healthcare | Interpretability

[10] Salerno F., Contenti A., Ulrich L., Marullo G., Moos S., Vezzetti E., MOTT modular optical tool tracking framework enabling efficient benchmarking. Scientific Reports, 15(1) (2025).
Mostra Abstract

Abstract: Optical tool tracking is the process of determining the 6DoF pose of an object in real time using visual sensor streams and image processing algorithms. It enables spatial localization in applications such as robotics, medical imaging, augmented reality, and precision manufacturing. However, existing solutions often involve tight coupling between hardware and software, complicating the management and benchmarking of different tracking systems. This paper presents the Modular Optical Tool Tracking (MOTT) framework, a unified platform for implementing, integrating, and benchmarking optical tracking solutions. A requirement-based design approach was adopted, using Quality Function Deployment (QFD) to systematically derive technical features and identify design drivers that guided the architecture of the framework. The resulting software framework standardizes the concept of an optical tracking method, featuring a flexible and extensible architecture based on object-oriented principles. Two marker-based tracking methods using an RGB camera as video source were evaluated through the developed framework. The presented case studies showcased the use of the framework for method implementation and comparison within a unified pipeline, reporting computational metrics such as frame rate, CPU and memory usage, and providing pose visualizations. The proposed framework enables standardized evaluation and benchmarking of optical tracking systems, and provides a foundation for future extensions involving non-optical tracking modalities and large-scale comparative studies. The implementation is openly available at https://github.com/tooltip-optical-tracking/mott-framework.

Keywords: Benchmarking | Modular framework | Optical tracking | Pose estimation | QFD | Tool tracking

[11] Rasetto S., Marullo G., Adamo L., Bordin F., Pavesi F., Innocente C., Vezzetti E., Ulrich L., ReHAb Playground: A DL-Based Framework for Game-Based Hand Rehabilitation. Future Internet, 17(11) (2025).
Mostra Abstract

Abstract: Hand rehabilitation requires consistent, repetitive exercises that can often reduce patient motivation, especially in home-based therapy. This study introduces ReHAb Playground, a deep learning-based system that merges real-time gesture recognition with 3D hand tracking to create an engaging and adaptable rehabilitation experience built in the Unity Game Engine. The system utilizes a YOLOv10n model for hand gesture classification and MediaPipe Hands for 3D hand landmark extraction. Three mini-games were developed to target specific motor and cognitive functions: Cube Grab, Coin Collection, and Simon Says. Key gameplay parameters, namely repetitions, time limits, and gestures, can be tuned according to therapeutic protocols. Experiments with healthy participants were conducted to establish reference performance ranges based on average completion times and standard deviations. The results showed a consistent decrease in both task completion and gesture times across trials, indicating learning effects and improved control of gesture-based interactions. The most pronounced improvement was observed in the more complex Coin Collection task, confirming the system’s ability to support skill acquisition and engagement in rehabilitation-oriented activities. ReHAb Playground was conceived with modularity and scalability at its core, enabling the seamless integration of additional exercises, gesture libraries, and adaptive difficulty mechanisms. While preliminary, the findings highlight its promise as an accessible, low-cost rehabilitation platform suitable for home use, capable of monitoring motor progress over time and enhancing patient adherence through engaging, game-based interactions. Future developments will focus on clinical validation with patient populations and the implementation of adaptive feedback strategies to further personalize the rehabilitation process.

Keywords: gesture recognition | hand rehabilitation | home-based rehabilitation | serious game

[12] Marullo G., Innocente C., Ulrich L., Lo Faro A., Porcelli A., Ruggieri R., Vecchio B., Vezzetti E., Home-based mirror therapy in phantom limb pain treatment: the augmented humans framework. Multimedia Tools and Applications, 84(28), 34145-34177 (2025).
Mostra Abstract

Abstract: The “Augmented Humans” term refers to the opportunity to improve human possibilities by using innovative technologies such as Artificial Intelligence (AI) and Extended Reality (XR). Digital therapies, particularly suitable for those treatments requiring multiple sessions, are increasingly being adopted for home-based treatment, enabling continuous monitoring and rehabilitation for patients, thus alleviating the burden on healthcare facilities by facilitating remote therapy sessions and follow-up visits. Among these, the Mirror Therapy (MT) for patients suffering from Phantom Limb Pain (PLP) could benefit greatly. This paper proposes a novel “Augmented Humans” framework for the treatment of PLP through home-based MT; the framework is designed to consider the activities carried on by the therapy center, the patient, and the system supporting the treatment. Moreover, an XR-based solution that integrates a Deep Learning (DL) approach has been developed to provide patients with a self-testing and self-assessment tool for conducting at-home rehabilitation sessions independently, even in the absence of physical medical staff. The DL algorithm enables real-time monitoring of rehabilitation exercises and automatic provision of personalized feedback on the gesture’s performance, supporting the progressive improvement of the patient’s movements and his ability to adhere to the treatment plan. The technical feasibility and usability of the proposed framework have been evaluated with 23 healthy subjects, highlighting an overall positive user experience. Remarkable results were obtained in terms of automatic gesture evaluation, with macro averaged accuracy and F1-score of 95%, paving the way for the adoption of the “Augmented Humans” approach in the healthcare domain.

Keywords: Augmented humans | Augmented therapy | Deep learning | Digital therapy | Extended reality | Home-based treatment | Phantom limb pain

[13] Innocente C., Boemio M., Lorenzetti G., Pulito I., Romagnoli D., Saponaro V., Marullo G., Ulrich L., Vezzetti E., Deep Learning-Based Lip-Reading for Vocal Impaired Patient Rehabilitation. CMES Computer Modeling in Engineering and Sciences, 143(2), 1355-1379 (2025).
Mostra Abstract

Abstract: Lip-reading technology, based on visual speech decoding and automatic speech recognition, offers a promising solution to overcoming communication barriers, particularly for individuals with temporary or permanent speech impairments. However, most Visual Speech Recognition (VSR) research has primarily focused on the English language and general-purpose applications, limiting its practical applicability in medical and rehabilitative settings. This study introduces the first Deep Learning (DL) based lip-reading system for the Italian language designed to assist individuals with vocal cord pathologies in daily interactions, facilitating communication for patients recovering from vocal cord surgeries, whether temporarily or permanently impaired. To ensure relevance and effectiveness in real-world scenarios, a carefully curated vocabulary of twenty-five Italian words was selected, encompassing critical semantic fields such as Needs, Questions, Answers, Emergencies, Greetings, Requests, and Body Parts. These words were chosen to address both essential daily communication and urgent medical assistance requests. Our approach combines a spatiotemporal Convolutional Neural Network (CNN) with a bidirectional Long Short-Term Memory (BiLSTM) recurrent network, and a Connectionist Temporal Classification (CTC) loss function to recognize individual words, without requiring predefined words boundaries. The experimental results demonstrate the system’s robust performance in recognizing target words, reaching an average accuracy of 96.4% in individual word recognition, suggesting that the system is particularly well-suited for offering support in constrained clinical and caregiving environments, where quick and reliable communication is critical. In conclusion, the study highlights the importance of developing language-specific, application-driven VSR solutions, particularly for non-English languages with limited linguistic resources. By bridging the gap between deep learning-based lip-reading and real-world clinical needs, this research advances assistive communication technologies, paving the way for more inclusive and medically relevant applications of VSR in rehabilitation and healthcare.

Keywords: 3D convolutional neural network | automatic speech recognition | deep learning | Lip-reading | visual speech decoding

[14] Cavazzana R., Faccia A., Cavallaro A., Giuranno M., Becchi S., Innocente C., Marullo G., Ricci E., Secco J., Vezzetti E., Ulrich L., Enhancing Clinical Assessment of Skin Ulcers with Automated and Objective Convolutional Neural Network-Based Segmentation and 3D Analysis. Applied Sciences Switzerland, 15(2) (2025).
Mostra Abstract

Abstract: Skin ulcers are open wounds on the skin characterized by the loss of epidermal tissue. Skin ulcers can be acute or chronic, with chronic ulcers persisting for over six weeks and often being difficult to heal. Treating chronic wounds involves periodic visual inspections to control infection and maintain moisture balance, with edge and size analysis used to track wound evolution. This condition mostly affects individuals over 65 years old and is often associated with chronic conditions such as diabetes, vascular issues, heart diseases, and obesity. Early detection, assessment, and treatment are crucial for recovery. This study introduces a method for automatically detecting and segmenting skin ulcers using a Convolutional Neural Network and two-dimensional images. Additionally, a three-dimensional image analysis is employed to extract key clinical parameters for patient assessment. The developed system aims to equip specialists and healthcare providers with an objective tool for assessing and monitoring skin ulcers. An interactive graphical interface, implemented in Unity3D, allows healthcare operators to interact with the system and visualize the extracted parameters of the ulcer. This approach seeks to address the need for precise and efficient monitoring tools in managing chronic wounds, providing a significant advancement in the field by automating and improving the accuracy of ulcer assessment.

Keywords: 3D analysis | automatic segmentation | chronic wound | clinical parameters | convolutional neural network | edge detection | interactive interface

[15] Ulrich L., Carmassi G., Garelli P., Lo Presti G., Ramondetti G., Marullo G., Innocente C., Vezzetti E., SIGNIFY: Leveraging Machine Learning and Gesture Recognition for Sign Language Teaching Through a Serious Game. Future Internet, 16(12) (2024).
Mostra Abstract

Abstract: Italian Sign Language (LIS) is the primary form of communication for many members of the Italian deaf community. Despite being recognized as a fully fledged language with its own grammar and syntax, LIS still faces challenges in gaining widespread recognition and integration into public services, education, and media. In recent years, advancements in technology, including artificial intelligence and machine learning, have opened up new opportunities to bridge communication gaps between the deaf and hearing communities. This paper presents a novel educational tool designed to teach LIS through SIGNIFY, a Machine Learning-based interactive serious game. The game incorporates a tutorial section, guiding users to learn the sign alphabet, and a classic hangman game that reinforces learning through practice. The developed system employs advanced hand gesture recognition techniques for learning and perfecting sign language gestures. The proposed solution detects and overlays 21 hand landmarks and a bounding box on live camera feeds, making use of an open-source framework to provide real-time visual feedback. Moreover, the study compares the effectiveness of two camera systems: the Azure Kinect, which provides RGB-D information, and a standard RGB laptop camera. Results highlight both systems’ feasibility and educational potential, showcasing their respective advantages and limitations. Evaluations with primary school children demonstrate the tool’s ability to make sign language education more accessible and engaging. This article emphasizes the work’s contribution to inclusive education, highlighting the integration of technology to enhance learning experiences for deaf and hard-of-hearing individuals.

Keywords: gamification | gesture recognition | hand sign alphabet | machine learning | serious game | social inclusion

[16] Marullo G., Ulrich L., Antonaci F., Audisio A., Aprato A., Massè A., Vezzetti E., Classification of AO/OTA 31A/B femur fractures in X-ray images using YOLOv8 and advanced data augmentation techniques. Bone Reports, 22 (2024).
Mostra Abstract

Abstract: Femur fractures are a significant worldwide public health concern that affects patients as well as their families because of their high frequency, morbidity, and mortality. When employing computer-aided diagnostic (CAD) technologies, promising results have been shown in the efficiency and accuracy of fracture classification, particularly with the growing use of Deep Learning (DL) approaches. Nevertheless, the complexity is further increased by the need to collect enough input data to train these algorithms and the challenge of interpreting the findings. By improving on the results of the most recent deep learning-based Arbeitsgemeinschaft für Osteosynthesefragen and Orthopaedic Trauma Association (AO/OTA) system classification of femur fractures, this study intends to support physicians in making correct and timely decisions regarding patient care. A state-of-the-art architecture, YOLOv8, was used and refined while paying close attention to the interpretability of the model. Furthermore, data augmentation techniques were involved during preprocessing, increasing the dataset samples through image processing alterations. The fine-tuned YOLOv8 model achieved remarkable results, with 0.9 accuracy, 0.85 precision, 0.85 recall, and 0.85 F1-score, computed by averaging the values among all the individual classes for each metric. This study shows the proposed architecture\'s effectiveness in enhancing the AO/OTA system\'s classification of femur fractures, assisting physicians in making prompt and accurate diagnoses.

Keywords: Computer assisted diagnosis | Convolutional neural networks | Deep learning | Femur fracture

[17] Checcucci E., Piazzolla P., Marullo G., Innocente C., Salerno F., Ulrich L., Moos S., Quarà A., Volpi G., Amparore D., Piramide F., Turcan A., Garzena V., Garino D., De Cillis S., Sica M., Verri P., Piana A., Castellino L., Alba S., Di Dio M., Fiori C., Alladio E., Vezzetti E., Porpiglia F., Development of Bleeding Artificial Intelligence Detector (BLAIR) System for Robotic Radical Prostatectomy. Journal of Clinical Medicine, 12(23) (2023).
Mostra Abstract

Abstract: Background: Addressing intraoperative bleeding remains a significant challenge in the field of robotic surgery. This research endeavors to pioneer a groundbreaking solution utilizing convolutional neural networks (CNNs). The objective is to establish a system capable of forecasting instances of intraoperative bleeding during robot-assisted radical prostatectomy (RARP) and promptly notify the surgeon about bleeding risks. Methods: To achieve this, a multi-task learning (MTL) CNN was introduced, leveraging a modified version of the U-Net architecture. The aim was to categorize video input as either “absence of blood accumulation” (0) or “presence of blood accumulation” (1). To facilitate seamless interaction with the neural networks, the Bleeding Artificial Intelligence-based Detector (BLAIR) software was created using the Python Keras API and built upon the PyQT framework. A subsequent clinical assessment of BLAIR’s efficacy was performed, comparing its bleeding identification performance against that of a urologist. Various perioperative variables were also gathered. For optimal MTL-CNN training parameterization, a multi-task loss function was adopted to enhance the accuracy of event detection by taking advantage of surgical tools’ semantic segmentation. Additionally, the Multiple Correspondence Analysis (MCA) approach was employed to assess software performance. Results: The MTL-CNN demonstrated a remarkable event recognition accuracy of 90.63%. When evaluating BLAIR’s predictive ability and its capacity to pre-warn surgeons of potential bleeding incidents, the density plot highlighted a striking similarity between BLAIR and human assessments. In fact, BLAIR exhibited a faster response. Notably, the MCA analysis revealed no discernible distinction between the software and human performance in accurately identifying instances of bleeding. Conclusion: The BLAIR software proved its competence by achieving over 90% accuracy in predicting bleeding events during RARP. This accomplishment underscores the potential of AI to assist surgeons during interventions. This study exemplifies the positive impact AI applications can have on surgical procedures.

Keywords: artificial intelligence | complications | prostate cancer | robotics

[18] Marullo G., Tanzi L., Piazzolla P., Vezzetti E., 6D object position estimation from 2D images: a literature review. Multimedia Tools and Applications, 82(16), 24605-24643 (2023).
Mostra Abstract

Abstract: The 6D pose estimation of an object from an image is a central problem in many domains of Computer Vision (CV) and researchers have struggled with this issue for several years. Traditional pose estimation methods (1) leveraged on geometrical approaches, exploiting manually annotated local features, or (2) relied on 2D object representations from different points of view and their comparisons with the original image. The two methods mentioned above are also known as Feature-based and Template-based, respectively. With the diffusion of Deep Learning (DL), new Learning-based strategies have been introduced to achieve the 6D pose estimation, improving traditional methods by involving Convolutional Neural Networks (CNN). This review analyzed techniques belonging to different research fields and classified them into three main categories: Template-based methods, Feature-based methods, and Learning-Based methods. In recent years, the research mainly focused on Learning-based methods, which allow the training of a neural network tailored for a specific task. For this reason, most of the analyzed methods belong to this category, and they have been in turn classified into three sub-categories: Bounding box prediction and Perspective-n-Point (PnP) algorithm-based methods, Classification-based methods, and Regression-based methods. This review aims to provide a general overview of the latest 6D pose recovery methods to underline the pros and cons and highlight the best-performing techniques for each group. The main goal is to supply the readers with helpful guidelines for the implementation of performing applications even under challenging circumstances such as auto-occlusions, symmetries, occlusions between multiple objects, and bad lighting conditions.

Keywords: 6D position estimation | Computer vision | Deep learning | RGB Input

[19] Marullo G., Tanzi L., Ulrich L., Porpiglia F., Vezzetti E., A Multi-Task Convolutional Neural Network for Semantic Segmentation and Event Detection in Laparoscopic Surgery. Journal of Personalized Medicine, 13(3) (2023).
Mostra Abstract

Abstract: The current study presents a multi-task end-to-end deep learning model for real-time blood accumulation detection and tools semantic segmentation from a laparoscopic surgery video. Intraoperative bleeding is one of the most problematic aspects of laparoscopic surgery. It is challenging to control and limits the visibility of the surgical site. Consequently, prompt treatment is required to avoid undesirable outcomes. This system exploits a shared backbone based on the encoder of the U-Net architecture and two separate branches to classify the blood accumulation event and output the segmentation map, respectively. Our main contribution is an efficient multi-task approach that achieved satisfactory results during the test on surgical videos, although trained with only RGB images and no other additional information. The proposed multi-tasking convolutional neural network did not employ any pre- or postprocessing step. It achieved a Dice Score equal to 81.89% for the semantic segmentation task and an accuracy of 90.63% for the event detection task. The results demonstrated that the concurrent tasks were properly combined since the common backbone extracted features proved beneficial for tool segmentation and event detection. Indeed, active bleeding usually happens when one of the instruments closes or interacts with anatomical tissues, and it decreases when the aspirator begins to remove the accumulated blood. Even if different aspects of the presented methodology could be improved, this work represents a preliminary attempt toward an end-to-end multi-task deep learning model for real-time video understanding.

Keywords: bleeding detection | CNN | laparoscopic surgery | multi-task convolutional neural network | semantic segmentation

[20] Padovan E., Marullo G., Tanzi L., Piazzolla P., Moos S., Porpiglia F., Vezzetti E., A deep learning framework for real-time 3D model registration in robot-assisted laparoscopic surgery. International Journal of Medical Robotics and Computer Assisted Surgery, 18(3) (2022).
Mostra Abstract

Abstract: Introduction: The current study presents a deep learning framework to determine, in real-time, position and rotation of a target organ from an endoscopic video. These inferred data are used to overlay the 3D model of patient's organ over its real counterpart. The resulting augmented video flow is streamed back to the surgeon as a support during laparoscopic robot-assisted procedures. Methods: This framework exploits semantic segmentation and, thereafter, two techniques, based on Convolutional Neural Networks and motion analysis, were used to infer the rotation. Results: The segmentation shows optimal accuracies, with a mean IoU score greater than 80% in all tests. Different performance levels are obtained for rotation, depending on the surgical procedure. Discussion: Even if the presented methodology has various degrees of precision depending on the testing scenario, this work sets the first step for the adoption of deep learning and augmented reality to generalise the automatic registration process.

Keywords: abdominal | Kidney | prostate

[21] Cannavò A. et al. Automatic generation of affective 3D virtual environments from 2D images. Visigrapp 2020 Proceedings of the 15th International Joint Conference on Computer Vision Imaging and Computer Graphics Theory and Applications, 1, 113-124 (2020).
Mostra Abstract

Abstract: Today, a wide range of domains encompassing, e.g., movie and video game production, virtual reality simulations, augmented reality applications, make a massive use of 3D computer generated assets. Although many graphics suites already offer a large set of tools and functionalities to manage the creation of such contents, they are usually characterized by a steep learning curve. This aspect could make it difficult for non-expert users to create 3D scenes for, e.g., sharing their ideas or for prototyping purposes. This paper presents a computer-based system that is able to generate a possible reconstruction of a 3D scene depicted in a 2D image, by inferring objects, materials, textures, lights, and camera required for rendering. The integration of the proposed system into a well-known graphics suite enables further refinements of the generated scene using traditional techniques. Moreover, the system allows the users to explore the scene into an immersive virtual environment for better understanding the current objects\' layout, and provides the possibility to convey emotions through specific aspects of the generated scene. The paper also reports the results of a user study that was carried out to evaluate the usability of the proposed system from different perspectives.

Keywords: Human-computer interaction | Image-based modeling | Scene and object modeling | Virtual reality

Top 25 most frequent keywords in publications
Deep learning7
Machine learning4
Virtual reality3
Gesture recognition3
Human-computer interaction2
Pose estimation2
Computer vision2
Serious game2
Artificial intelligence2
Computer-aided surgery1
Joint replacement1
Mixed reality1
Orthopedic surgery1
Total hip arthroplasty1
Adjustable surgical guide1
Augmented reality1
Cranio-maxillofacial surgery1
Hololens 21
Hybrid surgical guide1
Italian sign language (lis)1
Mediapipe1
Rgb-d cameras1
Sign-language recognition1
Ergonomics1
Industry 5.01

Tieniti in contatto con l'Associazione ADM

Per qualunque informazione non esitare a contattare la Segreteria ADM tramite le modalità previste nella sezione Contatti

Soci ADM 225

N° pubblicazioni censite 7426