Wish Suharitdamrong
Academic and research departments
Surrey Institute for People-Centred Artificial Intelligence (PAI), Centre for Vision, Speech and Signal Processing (CVSSP).About
My research project
Multi-Modal Foundation ModelsThis research proposal aims to address the limitations of current multimodal learning models, which often overlook fine-grained information in favour of global representations. In domains like multimedia (e.g., videos with images, audio, and transcripts) and healthcare (e.g., medical images and clinical data), multimodal data carry complex, overlapping semantic concepts. This research will develop novel self-supervised learning algorithms that focus on extracting and aligning fine-grained, multi-concept representations across modalities. By designing specialised neural architectures and loss functions, we will enhance the integration of multimodal data, enabling a deeper understanding of complex cross-modal relationships. This approach will have significant implications for fields like multimedia analysis and healthcare informatics, where detailed multimodal interpretation is essential.
Supervisors
This research proposal aims to address the limitations of current multimodal learning models, which often overlook fine-grained information in favour of global representations. In domains like multimedia (e.g., videos with images, audio, and transcripts) and healthcare (e.g., medical images and clinical data), multimodal data carry complex, overlapping semantic concepts. This research will develop novel self-supervised learning algorithms that focus on extracting and aligning fine-grained, multi-concept representations across modalities. By designing specialised neural architectures and loss functions, we will enhance the integration of multimodal data, enabling a deeper understanding of complex cross-modal relationships. This approach will have significant implications for fields like multimedia analysis and healthcare informatics, where detailed multimodal interpretation is essential.
Publications
Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) enable lightweight adaptation, yet they operate in isolation within each modality, limiting their ability in capturing cross-modal interactions. In this paper, we take a step in bridging this gap with Cross-Modal Low-Rank Adaptation (CoLA), a novel PEFT framework that extends LoRA by introducing a dedicated intermodal adaptation pathway alongside the standard intra-modal one. This dual-path design enables CoLA to adapt unimodal foundation models to multimodal tasks effectively, without interference between modality-specific and cross-modal learning. We evaluate CoLA across a range of vision-language (RefCOCO, RefCOCO+, Re-fCOCOg) and audiovisual (AVE, AVS) benchmarks , where it consistently outperforms LORA, achieving a relative gain of around 3% and 2%, respectively, while maintaining parameter efficiency. Notably, CoLA enables the first multi-task PEFT framework for visual grounding, bridging a key gap in efficient multimodal adaptation.