It takes true vision to find meaning in complexity. Our research explores new horizons in visual and aural intelligence to revolutionise a wide range of applications.

Our research

Advanced smart city data analysis: intelligent safety monitoring solutions based on image/video and audio understanding

Our research transforms multimodal urban data into intelligent perception and actionable insights for safer and smarter cities. We develop advanced AI for image, video, and audio understanding, enabling robust detection, recognition, tracking, and behaviour analysis of people, vehicles, objects, and events in complex urban environments. By combining multimodal intelligence with real-time visual and aural analysis, our research delivers scalable solutions for public safety, intelligent transportation, infrastructure monitoring, and next-generation smart city applications.

Multimodal Vision and mmWave Intelligence

Our research develops multimodal vision and mmWave intelligence for privacy-preserving human sensing and health monitoring. We integrate computer vision with millimetre-wave radar to leverage complementary visual and RF information, using vision as cross-modal supervision to enhance the learning and sensing capabilities of radar-based systems. Our research enables contactless and privacy-aware monitoring of human trajectories, activities, and vital signs, including respiration, heart rate, and blood pressure, with support for multiple users simultaneously. This technology addresses critical challenges in elderly care and smart environments, providing robust sensing capabilities without wearable devices and under diverse lighting and privacy-sensitive conditions.

Intelligent Medical Imaging and Diagnosis

Our research develops advanced AI for quantitative medical image analysis and AI-guided diagnosis. We develop deep learning methods for accurate segmentation, identification, and quantitative characterisation of anatomical structures and tissues, with a particular focus on CT-based foot imaging. We investigate unsupervised domain adaptation and cross-modality learning to develop robust and generalisable models that can overcome variations across imaging modalities, datasets, and clinical environments. In collaboration with industry partners, our research is translating these advances into reliable and efficient solutions for automated foot CT analysis, supporting improved clinical assessment, quantitative diagnosis, and personalised healthcare.

Recommendation and Agentic AI

Our research advances intelligent recommendation and agentic AI, developing AI systems that understand users, content, context, and goals to deliver personalised and adaptive recommendations. We explore multimodal recommendation, knowledge-guided reasoning, and large language models to move beyond conventional ranking toward intelligent agents capable of understanding complex user needs, planning, making decisions, and taking actions autonomously. By integrating multimodal understanding with agentic reasoning, our research aims to build personalised, context-aware, and trustworthy AI systems that can continuously learn and assist users in complex real-world scenarios.

Embodied Multimodal Intelligence

Our research advances embodied AI by connecting multimodal perception, reasoning, learning, and physical action. We develop vision-language-action models and efficient learning frameworks that enable robots to adapt to real-world environments, acquire new manipulation skills from limited experience, and learn effectively from human guidance and feedback. By bridging multimodal intelligence with physical interaction, our research aims to build adaptive, efficient, and increasingly autonomous embodied agents for real-world applications.

Advanced Learning for Complex and Data-Limited Applications

Our research develops advanced learning-based solutions for challenging real-world applications where data are noisy, sparse, heterogeneous, or weakly labelled. We investigate graph neural networks, few-shot and unsupervised learning, robust learning from noisy data, and other data-efficient learning paradigms to build AI systems that can learn effectively from limited and imperfect data. These approaches enable reliable and generalisable intelligence across a broad range of multimedia, computer vision, sensing, and real-world applications.