Please use this identifier to cite or link to this item: http://localhost:8081/jspui/handle/123456789/21750
Title: Vision-based Deep-learning Models for Gesture Synthesis and Recognition
Authors: Mallika
Issue Date: Jun-2025
Publisher: IIT, Roorkee
Abstract: Over the past few decades, the progress in gesture and sign language recognition systems has made human-computer interaction more intuitive and user-friendly. By enabling computers to understand and interpret human gestures and sign language, we have made significant strides toward making technology more accessible to a broader range of individuals. Moreover, gesture and sign language recognition systems have improved communication and accessibility for individuals with hearing or speech impairments. These systems can convert sign language into text or speech, enabling effective communication between individuals who use sign language and those who are not familiar with it. This breakthrough has expanded opportunities for social interaction, education, employment, and various aspects of daily life for people with hearing or speech impairments. Environmental factors and variations, such as self-occlusion, differences in hand size and shape, background noise, and changes in illumination, can greatly impact the accuracy and reliability of the recognition process. Self-occlusion occurs when parts of the hand or body obstruct each other, making it difficult to determine the boundaries of the gesture accurately. This can lead to errors in recognition and misinterpretation of the intended gesture. Some methods aim to minimize the impact of these noises and improve recognition accuracy. Despite the difficulties posed by environmental noises, ongoing research in gesture recognition continues to explore novel algorithms, machine learning models, and sensor technologies to improve the robustness and accuracy of the recognition process. In this thesis, we take these challenges into account for designing various hand gesture algorithms such as static and dynamic hand gesture recognition systems, multiview gesture generation systems, and gesture-to-gesture systems. Many multiview recognition systems are available in the literature that deals with the problem of occlusion, but it requires the same gesture from different views. Collecting multiview data is a challenging task, as it requires the use of multiple cameras positioned at various angles or repeated performance of the actions to capture sufficient variability. To overcome this difficulty, we focus on generating different views of the input gesture using conditional multiview gesture synthesis. Static gestures are generated using paired data, with the generation algorithm designed to prioritize the perceptual and structural quality of the generated images as a primary objective. The aim is to generate realistic and visually appealing gestures that closely resemble the intended gestures. Once the views are generated, they are given as input into the proposed gesture recognition system for classification. This system employs a multiview approach for gesture recognition, utilizing decision-level fusion for improved accuracy.
URI: http://localhost:8081/jspui/handle/123456789/21750
Research Supervisor/ Guide: Ghosh, Debashis and Pradhan, Pyari Mohan
metadata.dc.type: Thesis
Appears in Collections:DOCTORAL THESES (E & C)

Files in This Item:
File Description SizeFormat 
17915013_MALLIKA_FinalThesis.pdf4.53 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.