Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/1992
Title: Proximal interpolative fusion of heterogeneous representations for affective and secure voice intelligence
Authors: Phukan, Orchid Chetia
Buduru, Arun Balaji (Advisor)
Sharma, Rajesh (Advisor)
Keywords: Voice Intelligence
Pre-Trained Models
Emotion Recognition
Synthetic Speech Detection
Issue Date: 14-Jun-2026
Publisher: IIIT-Delhi
Abstract: The rapid emergence of large-scale Pre-Trained Models (PTMs) trained on massive and diverse corpora has fundamentally reshaped the landscape of numerous AI domains, in- cluding natural language processing, computer vision and voice intelligence. In the sce- nario of Voice Intelligence Systems (VIS), PTMs have completely shifted the paradigm. From usage of handcrafted features and task-specific architectures towards rich, gener- alizable representations that encode acoustic, linguistic, and paralinguistic cues in a unified feature space. As a result, enabling downstream tasks to be learned efficiently even in low-resource or label-scarce settings. In this thesis, we focus on two critical frontiers of VIS: firstly, the affective dimension, which concerns the system’s ability to interpret users’ emotional and expressive states. Secondly, the secure dimension, which requires robustness against synthetic-speech manipulations such as voice cloning and the capability to attribute a fake to its generating model. Both of these tasks are essen- tial for preserving trust and system integrity. These two capabilities are complementary and increasingly necessary across modern applications. Affective understanding sup- ports conversational agents, healthcare monitoring, and human–computer interaction, whereas security underpins voice authentication, forensic analysis, and safety-critical communication. In emerging personalized and identity-aware services, both affective insight and resilience to synthetic speech are essential for trustworthy, user-centric VIS. Both dimensions have seen substantial progress with the rise of large PTMs, and recent research has increasingly leveraged these models to advance VIS. However, PTMs are highly heterogeneous—differing in architectural design, training objectives, linguistic scope, and modality grounding. Consequently, each PTM encodes a distinct and partial view of the speech signal, emphasizing different acoustic, paralinguistic, or semantic cues. Recognizing this, researchers in application such as automatic speech recognition iv have begun exploring fusion of heterogeneous PTMs representations to exploit comple- mentary strengths for better performance, with similar trends emerging in different NLP and computer vision applications. In this thesis, we explore and deepen this direction by investigating how heterogeneous PTM representations can be systematically fused to advance both the affective and secure frontiers of VIS. To enable effective fusion, we introduce Proximal Interpolative Fusion, a novel fusion paradigm. Proximal refers to bringing diverse PTM representations into task-relevant compatibility across geometric, divergence, and interaction spaces. Interpolative denotes learning an intermediate fused manifold rather than collapsing one PTM’s representational space into another, thereby preserving complementary paralinguistic, acoustic, and semantic cues. Building on this paradigm, we develop a set of novel fusion frameworks that are organized into three core families, each capturing a distinct principle for how heterogeneous representa- tions are combined. (i) Geometry Guided Alignment enables fusion by placing PTM representations into shared geometric spaces prior to integration. This family includes HYFuse, which fuses representations through hyperbolic projection and möbius op- erations; MATA, which employs optimal transport for manifold-to-manifold alignment; and MERLINN, which performs dual-geometry fusion across euclidean and hyperbolic spaces. (ii) Interaction Guided Alignment achieves fusion through explicit interaction operators applied between representations. This includes MiO, which integrates PTM representations through bilinear pooling, followed by SCAR, which employs a cascaded cross-attention mechanism. (iii) Divergence Guided Alignment fuses representations by harmonizing their underlying distributional structure. This includes FINDER, which aligns representational distributions using renyi divergence, and COFFE, which lever- ages chernoff distance to jointly optimize separability. Additionally, PARROT serves as a hybrid framework that combines Interaction Guided with Geometry Guided Align- ment through a hadamard interaction branch followed by optimal transport. Across af- fective and secure VIS tasks—including emotion recognition (MATA, HYFuse, PARROT) and synthetic-speech detection and attribution (MiO, MERLINN, FINDER, SCAR, COFFE), these frameworks consistently demonstrate that carefully structured fusion outperforms individual PTMs, underscoring its importance for developing improved affective and secure VIS in the era of PTMs.
URI: http://repository.iiitd.edu.in/xmlui/handle/123456789/1992
Appears in Collections:Year-2026

Files in This Item:
File Description SizeFormat 
OC_Final_PHD_thesis_Final (1).pdf13.7 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.