A condition-aligned benchmark · 2026

One performance grid.
Four views of human intelligence.

HUG-VIS brings human-centered understanding and generation onto the same coordinate system—aligning affect, motion, speech, identity, and foreground quality within every performance.

Controlled release The dataset page and access instructions are available. A finalized license form and gated dataset files will be released before applications open.

01 Construction & evaluation workflow

From controlled collection and synchronized annotation to four task-specific evaluation tracks.

HUG-VIS

A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

Fei Ma · Zebang Cheng · Minghui Li · Hongbo Xu · Yuyong Tan · Yihua Shao · Hanling Wang · Zhou Liu · Yuqing Gao · Dong Wang · Long Ma · Laizhong Cui · Nicu Sebe · Qi Tian

01 / Overview

Evaluation that keeps the human performance fixed.

Most benchmarks isolate one capability on one dataset. HUG-VIS instead evaluates distinct systems against the same actors, assignments, source clips, and condition identifiers.

This shared coordinate system makes it possible to ask whether a difficult case is specific to one model family, shared across tasks, or tied to an interpretable property of the source performance. Each track keeps its own outputs and metrics—there is no artificial single score.

01

Controlled correspondence

The same 280 emotion–action–prompt assignments recur for every retained actor.

02

Native task protocols

Recognition, synthesis, cloning, and matting retain criteria appropriate to their outputs.

03

Diagnostic intersections

Shared identifiers support paired, source-level comparisons across otherwise distinct tasks.

02 / Dataset

A complete actor-by-assignment grid, not a loose collection.

30actors
×
7conditions
×
4actions
×
10utterances
=
8,400clips
02 Controlled performance examples

Every clip is a self-contained evaluation unit.

Each seated half-body performance begins from a canonical resting pose, follows an assigned Mandarin prompt and action template, and returns to rest. The fixed frontal studio setup prioritizes precise cross-condition diagnosis over in-the-wild diversity.

Capture
1920 × 1080 · 240 FPS source · 30 FPS release
Protocol
Start · perform · return
Quality
Complete-grid retention with professional review
Setting
Fixed frontal view · controlled lighting · green screen

Per-clip assets

  • RGB video
  • WAV audio
  • TXT prompts & transcripts
  • α soft alpha matte
  • ID condition identifiers

03 / Benchmark tracks

Four foundational tasks on one controlled benchmark.

HUG-VIS evaluates perception, generation, speech, and foreground understanding on the same actor-by-assignment grid. Every task follows a zero-shot protocol, while its inputs, outputs, and evaluation criteria remain native to the capability being measured.

01

Perception

Multimodal emotion recognition

Recognize seven instructed emotion conditions from isolated modalities and multimodal combinations.

Inputs
Image · video · audio · text
Output
Seven-class emotion label
Evaluation
Top-1 accuracy
02

Generation

Human video generation

Generate portrait, head, or half-body performances under audio-driven and vision-driven settings.

Inputs
Reference image · audio or motion
Output
Synchronized human video
Evaluation
Fidelity · identity · sync · MOS
03

Speech

Voice cloning

Synthesize expressive Mandarin speech while preserving source-speaker identity and affective delivery.

Inputs
Reference voice · target text
Output
Cloned speech waveform
Evaluation
Quality · speaker similarity · MOS
04

Foreground

Human video matting

Recover temporally coherent alpha mattes from controlled RGB sequences with fine boundary motion.

Input
Green-screen RGB video
Output
Per-frame alpha sequence
Evaluation
Spatial and temporal errors

Qualitative evaluation

Inspect what aggregate scores can hide.

Temporally aligned sequences and enlarged local regions reveal identity drift, incomplete motion, structural hand failures, and unstable alpha boundaries that average metrics can dilute.

01 · Audio-driven generation

Expression and identity across a Happy utterance.

Six ordered time points align the source portrait, transcript, audio waveform, and model outputs. The sequence exposes facial deformation, background artifacts, identity stability, and expression dynamics.

02 · Vision-driven generation

Motion completion across an Angry performance.

The source frame and driving motion sequence are followed by aligned outputs from open- and closed-source systems. Differences appear in arm-trajectory timing, gesture completion, appearance, and background stability.

03 · Fine-grained structure

Small hand regions reveal large structural failures.

Enlarged insets localize fused, extra, blurred, distorted, and missing fingers or hands across six systems. These local defects occupy few pixels and can disappear inside frame-averaged scores.

04 · Human video matting

Boundary errors tracked through time.

Seven frames compare RGB input, reference alpha, and nine model predictions. Red overlays highlight boundary discrepancies, foreground leakage, and temporal inconsistency relative to the reference.

04 / Experimental results

Results organized by task and criterion.

No HUG-VIS sample is used for training, fine-tuning, calibration, or model selection.

  • Common zero-shot protocol
  • Task-specific metrics
  • Objective results + human MOS
MER

Task 01

Emotion recognition

Best emotion recognition model and accuracy for each input setting
InputBest modelAccuracy ↑
Image frameMMA-DFER37.05%
VideoMiniCPM-o 4.548.52%
AudioAudio-Reasoner-7B74.38%
TextDeepSeek-V3.282.93%
Video + AudioQwen2.5-Omni-7B73.70%
Video + TextQwen2.5-Omni-7B83.74%
Video + Audio + TextQwen2.5-Omni-7B83.79%

Text is substantially stronger than visual-only inputs; adding audio to video + text improves the tested trimodal system by only 0.05 points.

ADG

Task 02 · Audio-driven

Human video generation

Audio-driven human video generation metric leaders
EvaluationCriterionBest modelScore
ObjectiveCSIM ↑Ditto0.904
ObjectiveSync-C ↑LatentSync5.43
ObjectiveSync-D ↓Sonic / LatentSync7.73
MOSID similarity ↑Sonic4.44
MOSEmotion ↑Sonic4.16
MOSLip sync ↑LatentSync4.47

Identity and synchronization produce different leaders: Ditto leads CSIM, while Sonic and LatentSync lead perceptual and lip-sync criteria.

VDG

Task 02 · Vision-driven

Human video generation by output scope

Vision-driven human video generation leaders within each output scope
GroupObjective metric leadersMOS leaders
Open · HeadX-NeMo: LPIPS 0.520, PSNR 10.15, FID 150.57
AniPortrait: CSIM 0.884, SSIM 0.383
PersonaLive!: ID 4.41
X-NeMo: Emotion 3.68, Motion 3.82
Open · BodyAnimate-X: LPIPS 0.139, PSNR 18.81, SSIM 0.755
Wan2.2: CSIM 0.783, FID 20.24
Wan2.2: ID 4.82, Emotion 4.61, Motion 4.72
Closed-sourceVidu: LPIPS 0.195, CSIM 0.786, FID 20.05
Kling: PSNR 16.41, SSIM 0.721
Kling: ID 4.76, Motion 4.64
Vidu: Emotion 4.60

Systems are compared only within their reported output scopes. No single model leads reconstruction, identity, motion, and human judgment simultaneously.

VC

Task 03

Voice cloning

Voice cloning leaders for open-source and closed-source systems
GroupUTMOS ↑DNSMOS ↑Speaker sim. ↑MOS ID ↑MOS emotion ↑
Open-sourceOpenAudio S1
2.32
OpenAudio S1
3.01
CosyVoice 3
0.856
IndexTTS2
4.24
IndexTTS2
4.40
Closed-sourceInworld TTS-1.5
2.69
Inworld TTS-1.5
3.22
Eleven Multilingual v2
0.779
Eleven Multilingual v2
3.73
Eleven Multilingual v2
4.23

Real audio provides a 0.990 speaker-similarity reference. Reference-free quality scores should not be treated as upper-bound fidelity measurements.

VM

Task 04

Human video matting

Top two human video matting systems across five error metrics
ModelMAD ↓MSE ↓dtSSD ↓Grad ↓Conn ↓
BiRefNet2.300.822.0410.434.31
MatAnyone 23.860.912.1211.454.53

BiRefNet leads all five criteria. Residual errors concentrate around fine hand boundaries, rapid motion contours, and foreground leakage.

Cross-task analysis

Difficulty depends on the source, capability, and evaluation criterion.

There is no universally difficult emotion.

Afraid is hardest for recognition, Sad for several head-generation criteria, Angry for voice-quality predictors, and Happy for spatial matting error. Difficulty belongs to the combination of source, capability, and evaluation criterion—not to the source alone.

  • MER Afraid is hardest
  • Head synthesis Sad often ranks hardest
  • Voice cloning Angry is hardest
  • Matting Happy has highest MAD
03 Emotion × metric difficulty
Agreement between video-driven and audio-driven source difficulty
Matched generation criteria. Audio- and video-driven source difficulty aligns closely on shared clips (ρ = 0.99).
Relationship between ground-truth motion and temporal matting error
Motion and temporal error. Larger foreground changes are associated with higher dtSSD (ρ = 0.63).

05 / Access & responsibility

Controlled access for responsible academic research.

Controlled release

HUG-VIS contains identifiable recordings of human participants. Access is limited to approved, non-commercial academic research under the recipient data-use agreement. Once applications open, requests will be reviewed manually, and completing the steps below will not guarantee approval.

Request access

Keep applicant details consistent across every step.

  1. Read the finalized license. Confirm that your institution and proposed work satisfy the non-commercial academic-use terms.
  2. Submit the Hugging Face request. When gated access opens, sign in with an individual account, complete every field, and provide the exact username that should receive access.
  3. Complete and sign the finalized agreement. The responsible applicant/signatory and Hugging Face requester must be the same eligible individual, with matching name, institution, position, and official institutional email. They must be faculty, a researcher, or research staff at a university or public/non-profit research institution; students cannot sign as the responsible applicant.
  4. Email the signed agreement. Send it from the same official institutional address to Zebang Cheng, copy Fei Ma, and use [HUG-VIS Access] Full Name | Institution | HF username as the subject.
  5. Wait for manual review. If approved, access is granted only to the individual Hugging Face username named in the application.

The form in the current main GitHub repository still contains provider and citation placeholders and is not final. Check for a finalized revision before signing or submitting it.

Download after approval

Authenticate with the approved account.

hf auth login
hf download GML-MMGroup/HUG-VIS \
  --repo-type dataset \
  --local-dir HUG-VIS

Use the same individual account named in the application. Never share passwords, tokens, or access credentials.

Contact

Direct each question to the right contact.

Dataset access & project
Zebang Cheng
zebang.cheng@gmail.com

Cc: Fei Ma
mafei@gml.ac.cn
Paper correspondence
Long Ma · Qi Tian

Responsible use

The people represented in HUG-VIS remain identifiable.

All actors provided written informed consent before recording for research capture and authorized use of their identifiable likeness, voice, and performed behavior. Public examples released by the project are limited to research uses covered by those authorizations; dataset access does not grant recipients permission to republish identifiable material.

License essentials

  • Use the dataset and derived materials only for scientific, educational, non-commercial academic purposes.
  • Do not redistribute data, annotations, restricted derivatives, or access credentials.
  • Keep research-group use under the responsible applicant's direct supervision, inside the approved research environment, and under the same terms.
  • Do not edit, composite, dub, replace, or republish original video or audio as modified source data; research processing stays inside the approved environment.
  • Do not identify, contact, track, impersonate, surveil, or harm recorded participants.
  • Do not use HUG-VIS for deceptive deepfakes, defamation, discrimination, sexual content, or unlawful activity.
  • Do not publicly release identifiable samples or outputs reproducing a recognizable face or voice without prior written permission from the Data Provider.
  • Store the dataset securely, report loss, leakage, or unauthorized access, and delete every copy when research ends, access is withdrawn, or the Data Provider requests deletion.

This summary does not replace the complete HUG-VIS Dataset Academic Use License.

06 / Citation

Build on the shared coordinate system.

Formal publication metadata and the final BibTeX entry will be added with the paper release.

Follow the release
Draft citationReady to copy
@misc{hugvis2026,
  title  = {HUG-VIS: A Multimodal Benchmark for
            Human-centered Understanding and Generation
            in Visual Intelligence},
  author = {Ma, Fei and Cheng, Zebang and Li, Minghui and
            Xu, Hongbo and Tan, Yuyong and Shao, Yihua and
            Wang, Hanling and Liu, Zhou and Gao, Yuqing and
            Wang, Dong and Ma, Long and Cui, Laizhong and
            Sebe, Nicu and Tian, Qi},
  year   = {2026},
  url    = {https://github.com/GML-MMGroup/HUG-VIS}
}
Expanded research figure