Face Is Not All You Need:
MIME Benchmark for Incomplete Multimodal Emotion Recognition

Yuxin Jia1, Xing Lan1, Wensong Wang1, Jian Xue2, Feiliang Ren1, Ke Lu2
1School of Computer Science and Engineering, Northeastern University, Shenyang, China
2School of Engineering Science, University of Chinese Academy of Sciences, Beijing, China
News: [2026.04] The MIME benchmark repository is created. Mini data and codes are publicly available at our GitHub repository.
Availability

Only sample data are publicly released: 4 videos per subset (28 videos total). For access to the full dataset, please sign the license.pdf and email the signed file to jinj62062@gmail.com with a CC to lanx@cse.neu.edu.cn.

1. Scene Understanding

The target person seated at a table with a patterned booth seat. The scene is a restaurant with other patrons visible. Her upright posture and right hand raised near her face suggest emphasis during a heated interaction. The face is heavily blurred.

2. Emotional Analysis

Although the face is heavily blurred, we can still infer anger from the following cues. The target person says "I felt his tongue actually licked my teeth. I don't get..." with a loud, fast-paced tone and slight tremor. She faces slightly right, hand raised, indicating frustration during a direct confrontation in a public setting.

3. Conclusion

Anger

Abstract

Multimodal Emotion Recognition (MER) is a fundamental task in multimedia understanding, where state-of-the-art methods have achieved prominent success relying on the assumption of complete, well-aligned multimodal inputs. However, real-world unconstrained scenarios often suffer from unpredictable modality degradation and information loss.

To bridge this gap, we formally define the Incomplete Multimodal Emotion Recognition (IMER) task, which benchmarks model generalization and robustness under fine-grained modality degradation and information loss. We construct MIME, a dedicated IMER benchmark with 2000 video segments spanning natural and controlled incompleteness across seven subsets: FM, FDM, FSM, VMM, FDAM, FSAM, and AMM. We further propose the Chain of Emotion (CoE) analysis paradigm with tailored evaluation metrics to examine how Multimodal Large Language Models (MLLMs) perceive context, infer affect, and reach emotion conclusions under incomplete observations.

Key Features

Multi-Grained Scenarios

MIME provides 7 subsets covering one full-modality subset and six incomplete subsets: FM, FDM, FSM, VMM, FDAM, FSAM, and AMM.

Adaptive Data Pipeline

A four-stage pipeline combines large multimodal models, ground-truth quality filtering, and remained-cue verification to reduce hallucination while preserving realistic incomplete cues.

Structured Reasoning (CoE)

Annotations explicitly guide models to decompose observations into three parts: Scene Understanding, Emotional Analysis, and Conclusion.

LLM-as-a-Judge Evaluation

We evaluate reasoning with Context Perception Score (CPS), Affective Inference Score (AIS), Label Consistency Score (LCS), and Holistic CoE Index (HCI), and verify ranking stability with cross-judge, prompt-sensitivity, and HCI-sensitivity checks.

Chain-of-Emotion Pipeline

Data Construction Pipeline

Our adaptive four-stage generation pipeline delivers high-quality tripartite reasoning chains with conditional remained-cue verification. For evaluation, we use an LLM-as-a-Judge protocol and additionally validate robustness with an independent Gemini judge, prompt variants, and alternative HCI weight settings.

Robustness Checks

Cross-judge validation. Re-scoring a 245-sample stratified subset with gemini-3.1-flash preserves the overall ranking with Spearman correlation 0.9643; the top-5 models remain unchanged.

Prompt sensitivity. Three prompt variants produce stable rankings, with Spearman correlation 1.0000 between Prompt 1 and Prompt 2, and 0.9286 between Prompt 1 and Prompt 3.

HCI sensitivity. Reweighting CPS/AIS/LCS yields ranking Spearman correlations from 0.93 to 1.00, indicating that conclusions are not driven by one specific weighting choice.

Natural Degradation Already Present in Source Videos

MIME is not built on purely clean source footage. Before adding controlled blur or modality removal, the in-the-wild source videos already contain natural incompleteness such as low-light concealment, side-facing heads, foreground or object occlusion, and hand or body occlusion.

Our rebuttal-stage supplement now highlights four representative source clips spanning Subsets 2 and 3, supporting our revised framing of MIME as natural + controlled incompleteness.

Dataset Diversity

Dataset Diversity

The benchmark encompasses specific modality missingness types decoupled across diverse in-the-wild scene contexts (e.g., Lifestyle, Movie, Vlog).

Evaluation Metrics

Evaluation Metrics

Performance comparison across various MLLMs on four predefined CoE metrics. The radar charts illustrate the varying robustness and reasoning capabilities under multi-grained missing scenarios.

fig_duration
fig_emotion

7 MIME Subsets

Evaluating models under unpredictable modality degradation and information loss while keeping the original subset naming scheme.

Natural Degradation Examples

MIME is not blur-only. The source videos already contain natural degradation such as low-light concealment, side-facing heads, foreground/object occlusion, and hand/body occlusion.

All four examples below are original source clips with no extra synthetic blur added. They illustrate why MIME should be understood as a benchmark with natural + controlled incompleteness. We also show the corresponding ground-truth CoE annotations.

Neutral | Low-light / low-visibility

Low-light / low-visibility source clip in the original data, before any controlled degradation is applied.

1. Scene Understanding: The man in the grey suit stands motionless in a dim, industrial setting. Despite the blurred face, his upright posture and profile view suggest stillness. Structural pipes loom in the shadowed background, reinforcing the sterile, unemotional atmosphere.

2. Emotional Analysis: Although the face is hard to read under low visibility, his low, steady, slightly hesitant tone and stationary upright posture reinforce calm detachment. The dark industrial scene adds no strong affective cue and remains consistent with neutrality.

3. Conclusion: Neutral

Fear | Side-facing / turning-away head

Side-facing / turning-away head, leaving emotion inference to non-frontal evidence in the original source clip.

1. Scene Understanding: The woman in the light-colored top is in a dim, shadowy indoor setting. Her head tilts forward, hands clutch near her chest, and her body angles away as if reacting to an unseen threat.

2. Emotional Analysis: Even with non-frontal facial evidence, the strongest cues are the high-pitched sustained scream, tense defensive posture, and confined dark environment. These signals jointly support an acute fear reading.

3. Conclusion: Fear

Fear | Foreground / object occlusion

Foreground / object occlusion in the original source clip, with fear inferred from body tension and scene dynamics.

1. Scene Understanding: The woman with dark hair pulled back sits tense in the driver’s seat, hands gripping the steering wheel tightly. The car’s interior is visible, with a blurred urban backdrop.

2. Emotional Analysis: The strongest cues are the tight grip on the steering wheel, rigid posture, and fixed forward gaze. Ambient vehicle motion and the constrained in-car setting reinforce anticipatory fear and acute alertness.

3. Conclusion: Fear

Anger | Hand / body occlusion

Hand / body occlusion with tension preserved through posture and residual facial evidence.

1. Scene Understanding: The man in the vest lies tense on a bed indoors, under warm light with curtains. His body posture suggests emotional strain within a quiet, private setting.

2. Emotional Analysis: The strongest cues are the tense body posture, restrained movement, and suppressed emotional display. Combined with residual lower-face tension, these cues support restrained anger rather than overt expression.

3. Conclusion: Anger

Qualitative CoE Evaluation

A pronounced gap between visual perception and cognitive reasoning:
"Seeing" a scene does not guarantee "understanding" its emotion.

Ground Truth CoE

1. Scene Understanding:

The woman with shoulder-length reddish-brown hair in a brown sweater stands in a cozy kitchen. The face is heavily blurred. Wooden shelves with jars fill the background under warm lighting, suggesting domestic comfort as she tilts her head slightly, relaxed and engaged.

2. Emotional Analysis:

Although the face is heavily blurred, we can still infer Happy from the following cues. Another person says "Dis-moi, est-ce que la reponse est pas tres proche ?" with a light, playful, inquisitive tone. The warm, domestic scene and her relaxed posture reinforce this cheerful, curious mood.

3. Conclusion:

Happy

Qwen3.5-plus (Predicted)

Rating: 7.97
1. Scene Understanding: Score: 6.67

The subject stands in a dimly lit, cozy pantry or kitchen, surrounded by shelves of jars and bottles. She wears a soft pink sweater and maintains a relaxed, upright posture. Her head is slightly bowed with eyes cast downward, suggesting a moment of quiet reflection or internal amusement within a comfortable domestic setting.

2. Emotional Analysis: Score: 7.67

Despite facial blurring, the visible upward curve of the mouth and slight cheek elevation indicate a gentle smile. The downward gaze and lack of tension in the shoulders suggest contentment rather than excitement. The subtitles imply she is responding to a question with a hint of playful confirmation, reinforcing a sense of quiet satisfaction or mild amusement.

3. Conclusion: Score: 10.00

Happy

GPT-5-mini (Predicted)

Rating: 5.33
1. Scene Understanding: Score: 8.33

A young person stands in a dim, warm-toned pantry/kitchen framed by shelves of jars and boxes. She wears a textured sweater and leans slightly forward with relaxed shoulders and a subtle upward head tilt across frames. Subtitles present questions, indicating conversational pacing. Environment is intimate and calm, with no defensive or abrupt body motions.

2. Emotional Analysis: Score: 6.67

The forward lean, head tilt and consistent questioning in the subtitles imply engagement and anticipation rather than anger or fear. Relaxed shoulders and absence of closed-off gestures point to openness. Taken together these cues suggest the subject is reacting to new information with alert interest and mild astonishment, an engaged and expectant state bordering on surprise.

3. Conclusion: Score: 0.00

Surprise

Rebuttal Supplementary Tables

We also place the new rebuttal-stage tables directly on the website for quick inspection without opening the PDF supplement.

Cross-Judge Validation

245-sample stratified subset, re-scored by Gemini-3.1-flash. Top-3 in each metric are highlighted.

Model CPS AIS LCS HCI
Qwen3.5-plus 7.24 5.01 4.29 5.68
Gemini-3.1-flash 6.96 4.84 4.59 5.61
Qwen3.5-27b 7.17 5.01 3.92 5.54
Doubao-seed-2-0 6.80 4.67 4.65 5.52
GPT-5-mini 6.54 4.54 3.22 4.95
GPT-5.4-nano 5.92 4.06 2.82 4.43
GPT-5-nano 5.93 3.85 2.69 4.34

Prompt Sensitivity

Same stratified subset under three Gemini judge prompts, reported with HCI. Top-3 in each column are highlighted.

Model P1 P2 P3
Qwen3.5-plus 5.68 5.98 5.97
Gemini-3.1-flash 5.61 5.95 5.92
Qwen3.5-27b 5.54 5.81 5.73
Doubao-seed-2-0 5.52 5.75 5.89
GPT-5-mini 4.95 5.43 5.43
GPT-5.4-nano 4.43 4.86 4.86
GPT-5-nano 4.34 4.78 4.89

HCI Sensitivity to Alternative Weight Settings

HCI under four (α,β) settings. Spearman correlation is computed against the default ranking under (0.4, 0.3).

Model Default
(0.4, 0.3)
S1
(0.4, 0.2)
S2
(0.3, 0.4)
S3
(0.2, 0.5)
Qwen3.5-plus 5.605 5.516 5.467 5.329
Gemini-3.1-flash 5.592 5.536 5.462 5.332
Qwen3.5-27b 5.463 5.382 5.319 5.175
Doubao-seed-2-0 5.455 5.424 5.308 5.161
GPT-5-mini 4.734 4.640 4.585 4.436
gpt-5-nano 4.315 4.204 4.168 4.021
GPT-5.4-nano 4.234 4.130 4.090 3.946
Spearman (vs. Default) 1.00 0.93 1.00 0.96
table1
table2
table34

Repository Structure

MIME/
      |- data/
      |  |- Subset1_FM/
      |  |- Subset2_FDM/
      |  |- Subset3_FSM/
      |  |- Subset4_VMM/
      |  |- Subset5_FDAM/
      |  |- Subset6_FSAM/
      |  \- Subset7_AMM/
      |- data_list.txt
      |- eval/
      |  |- eval_coe.py
      |  \- predictcoe_evalacc.py/
      |- label.jsonl
      |- README.md
      |- license.pdf
      \- supplementary_material.pdf

BibTeX

@misc{jia2026mime,
  title  = {Face Is Not All You Need: MIME Benchmark for Incomplete Multimodal Emotion Recognition},
  author = {Yuxin Jia and Xing Lan and Wensong Wang and Jian Xue and Feiliang Ren and Ke Lu},
  year   = {2026},
  url    = {https://yuxinokk.github.io/MIME/}
}