| dc.description.abstract |
The rapid emergence of large-scale Pre-Trained Models (PTMs) trained on massive and diverse corpora has fundamentally reshaped the landscape of numerous AI domains, in- cluding natural language processing, computer vision and voice intelligence. In the sce- nario of Voice Intelligence Systems (VIS), PTMs have completely shifted the paradigm. From usage of handcrafted features and task-specific architectures towards rich, gener- alizable representations that encode acoustic, linguistic, and paralinguistic cues in a unified feature space. As a result, enabling downstream tasks to be learned efficiently even in low-resource or label-scarce settings. In this thesis, we focus on two critical frontiers of VIS: firstly, the affective dimension, which concerns the system’s ability to interpret users’ emotional and expressive states. Secondly, the secure dimension, which requires robustness against synthetic-speech manipulations such as voice cloning and the capability to attribute a fake to its generating model. Both of these tasks are essen- tial for preserving trust and system integrity. These two capabilities are complementary and increasingly necessary across modern applications. Affective understanding sup- ports conversational agents, healthcare monitoring, and human–computer interaction, whereas security underpins voice authentication, forensic analysis, and safety-critical communication. In emerging personalized and identity-aware services, both affective insight and resilience to synthetic speech are essential for trustworthy, user-centric VIS. Both dimensions have seen substantial progress with the rise of large PTMs, and recent research has increasingly leveraged these models to advance VIS. However, PTMs are highly heterogeneous—differing in architectural design, training objectives, linguistic scope, and modality grounding. Consequently, each PTM encodes a distinct and partial view of the speech signal, emphasizing different acoustic, paralinguistic, or semantic cues. Recognizing this, researchers in application such as automatic speech recognition iv have begun exploring fusion of heterogeneous PTMs representations to exploit comple- mentary strengths for better performance, with similar trends emerging in different NLP and computer vision applications. In this thesis, we explore and deepen this direction by investigating how heterogeneous PTM representations can be systematically fused to advance both the affective and secure frontiers of VIS. To enable effective fusion, we introduce Proximal Interpolative Fusion, a novel fusion paradigm. Proximal refers to bringing diverse PTM representations into task-relevant compatibility across geometric, divergence, and interaction spaces. Interpolative denotes learning an intermediate fused manifold rather than collapsing one PTM’s representational space into another, thereby preserving complementary paralinguistic, acoustic, and semantic cues. Building on this paradigm, we develop a set of novel fusion frameworks that are organized into three core families, each capturing a distinct principle for how heterogeneous representa- tions are combined. (i) Geometry Guided Alignment enables fusion by placing PTM representations into shared geometric spaces prior to integration. This family includes HYFuse, which fuses representations through hyperbolic projection and möbius op- erations; MATA, which employs optimal transport for manifold-to-manifold alignment; and MERLINN, which performs dual-geometry fusion across euclidean and hyperbolic spaces. (ii) Interaction Guided Alignment achieves fusion through explicit interaction operators applied between representations. This includes MiO, which integrates PTM representations through bilinear pooling, followed by SCAR, which employs a cascaded cross-attention mechanism. (iii) Divergence Guided Alignment fuses representations by harmonizing their underlying distributional structure. This includes FINDER, which aligns representational distributions using renyi divergence, and COFFE, which lever- ages chernoff distance to jointly optimize separability. Additionally, PARROT serves as a hybrid framework that combines Interaction Guided with Geometry Guided Align- ment through a hadamard interaction branch followed by optimal transport. Across af- fective and secure VIS tasks—including emotion recognition (MATA, HYFuse, PARROT) and synthetic-speech detection and attribution (MiO, MERLINN, FINDER, SCAR, COFFE), these frameworks consistently demonstrate that carefully structured fusion outperforms individual PTMs, underscoring its importance for developing improved affective and secure VIS in the era of PTMs. |
en_US |