Yu Liu (刘毓)

ELLIS PhD Student · CMVS, University of Oulu & ELLIS Institute Finland

I am an ELLIS PhD student jointly affiliated with the Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, and the ELLIS Institute Finland. I am supervised by Prof. Guoying Zhao, and co-supervised by Prof. Antti Honkela at the ELLIS Institute Finland and Prof. Björn Schuller within the ELLIS PhD Program. I received my M.S. in Artificial Intelligence from the Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, in 2026, supervised by Prof. Taihao Li and Dr. Leyuan Qu.

My research mainly focuses on privacy-aware multimodal affect understanding with multimodal large language models, tracing every affective interpretation back to the audio, video and text evidence it rests on. If you are interested in collaborating, feel free to reach out via email.

News

  • 2026.08 From Coarse to Nuanced was accepted to IEEE Transactions on Affective Computing.
  • 2026.05 Awarded Outstanding Graduate of UCAS.
  • 2026.05 Awarded Outstanding Master's Thesis of HIAS, UCAS.
  • 2026.03 Think-Before-Draw was accepted to Pattern Recognition.

Research Interests

Multimodal Affective Computing

Reading affect from faces, voices and language, from single frames to whole conversations.

Multimodal Large Language Models

Making perception, reasoning and adaptation more reliable when the evidence spans audio, video and text.

Privacy-Aware Multimodal Learning

Protecting the people behind the data, and understanding what that protection costs.

Publications (Selected)

* equal contribution

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Gra... thumbnail

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

IEEE Transactions on Affective Computing

Yu Liu, Leyuan Qu, Hanlei Shi, Di Gao, Yuhua Zheng, Taihao Li

Introduces GRACE for dynamic facial emotion recognition by aligning refined linguistic cues with salient facial dynamics, achieving state of the art on DFEW, FERV39k and MAFW.

HOPE: Hierarchical Fusion for Optimized and Personality-A... thumbnail

HOPE: Hierarchical Fusion for Optimized and Personality-Aware Estimation of Depression

ACM MM'25 1st Place · MPDD

Hanlei Shi*, Yu Liu*, Haoxun Li*, Yuxuan Ding*, Leyuan Qu, Taihao Li

Subject-level depression detection under privacy constraints via hierarchical audio–video fusion with personalized textual cues and cross-task consistency.

EMORL-TTS: Reinforcement Learning for Fine-Grained Emotio... thumbnail

EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS

ICASSP'26

Haoxun Li, Yu Liu, Yuqing Sun, Hanlei Shi, Leyuan Qu, Taihao Li

Unifies global VAD intensity control with local emphasis regulation in LLM-based TTS through supervised fine-tuning and reinforcement learning, improving emotional accuracy and intensity differentiation without sacrificing synthesis quality.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-G... thumbnail

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

Pattern Recognition

Hanlei Shi, Leyuan Qu, Yu Liu, Di Gao, Yuhua Zheng, Taihao Li

A two-stage framework for disentangled, controllable talking-head synthesis; I designed the multimodal fusion module that preserves identity while sharpening affect control.

Follow the Clues, Frame the Truth: Hybrid-evidential Dedu... thumbnail

Follow the Clues, Frame the Truth: Hybrid-evidential Deductive Reasoning in Open-Vocabulary Multimodal Emotion Recognition

Under Review

Yu Liu, Lei Zhang, Haoxun Li, Hanlei Shi, Yuxuan Ding, Leyuan Qu, Taihao Li

Formalizes emotion inference as a Propose–Verify–Decide protocol and uses reinforcement learning with hierarchical reward shaping to internalize abductive reasoning, producing interpretable evidence traces while outperforming strong baselines.

Centering Emotion Hotspots: Multimodal Local-Global Fusio... thumbnail

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

Under Review

Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun, Yuxuan Ding, Linlin Gong, Leyuan Qu, Taihao Li

Extends frame-level recognition to conversation-level tracking by centering multimodal hotspots and fusing them with dialogue context, improving stability and generalization across turns.

Experience

China Southern Airlines 2019.09 - 2023.05

Data Analyst, Airlines Operations Center

  • Built Python/MySQL analytics and monitoring, and automated Tableau/QuickBI reporting, reducing manual report time by over 70% and enabling shared data services across departments.
  • Led a Django + SQL delay/fault analytics platform, cutting report generation time by roughly 85% in support of operational safety and decision-making.
Hangzhou Institute for Advanced Study, UCAS 2024.09 - 2026.06

Lab Coordinator & Administrator

  • Manage onboarding, event organization and GPU resource scheduling; oversee equipment procurement and reimbursement workflows to keep projects moving.

Academic Service

  • Journal Reviewer for IEEE Transactions on Affective Computing, 2026 - present.
  • Journal Reviewer for Pattern Recognition (Elsevier), 2025 - present.

Education

  • University of Chinese Academy of Sciences

    M.S. in Artificial Intelligence, Hangzhou Institute for Advanced Study

    2023.09 - 2026.06
  • Civil Aviation University of China

    B.S. in Transportation

    2015.09 - 2019.06

Acknowledgements

During my master's studies at HIAS, UCAS, I was very fortunate to be supervised by Prof. Taihao Li and Dr. Leyuan Qu, whose guidance has shaped how I think about research.

I also greatly enjoy collaborating with Haoxun Li and Hanlei Shi — I truly value our inspirational discussions and seamless teamwork.