PUBLICATIONS / RESEARCH RECORD

Ideas, methods,
and the evidence behind them.

Browse three connected research areas separately, then refine by year, status, authorship, and publication type. Each entry is written at two levels: a fast plain-language takeaway and the technical record.

New Scholar records awaiting curation Checking Scholar snapshot…

Automatically indexed records are shown with minimal metadata until their venue, authorship, contribution, and public evidence are manually verified.

Newest first · status shown explicitly

Research figure for TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
Preprint 2026 Preprint

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

Yuhao Wang, Mu Qiao, Xindong Zhang, Lei Zhang, Yunzhi Zhuge, Huchuan Lu

Authorship First author

arXiv preprint

A training-free GUI-agent framework for irreversible visual-token admission that remains useful for unknown future targets while preserving coverage of operable regions.

Why this work matters
Contribution

Ranks visual evidence with layout-derived interaction priors, instruction relevance, and feature novelty; repairs spatial coverage into a nested token order; and contracts retired frames with monotone KV contraction.

  • GUI Agents
  • Efficient Agents
  • Evidence Ordering
Research figure for Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation
Accepted 2026 Journal

Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

Weixiang Zhou†, Xingguo Xu†, Yuhao Wang, Cong Wang, Yang Yang, Zhixun Su, Jinshan Pan

Authorship Contributor

IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

Dual semantic guidance and global-local modulation improve part-level alignment and hierarchical multimodal aggregation.

Why this work matters
Contribution

Combines unified text semantics, soft-mask local modulation, and hierarchical mixture-of-experts fusion.

  • Multimodal ReID
  • Semantic Guidance
  • Mixture of Experts
Research figure for Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
Published 2026 Journal

Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

Weixiang Zhou†, Jiabei Zuo†, Yuhao Wang, Cong Wang*, Huchuan Lu, Zhixun Su*

Authorship Contributor

IEEE Transactions on Image Processing (TIP)

PRISM combines Prompt-S6, semantic-driven token pruning, and progressive fusion for efficient tri-modal object ReID.

Why this work matters
Contribution

Uses segmentation-model priors to suppress background tokens while preserving linear-complexity cross-modal interaction.

  • Multimodal ReID
  • State Space Models
  • Token Pruning
Accepted · Poster 2026 Conference

Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification

Weixiang Zhou, Yuhao Wang†, Xingguo Xu, Cong Wang, Weizhen Zhou, Zhixun Su, Jinshan Pan

Authorship Co-first author † Shared first authorship

European Conference on Computer Vision (ECCV)

Accepted ECCV 2026 poster. A technical summary is intentionally omitted until reliable public paper metadata is available.

  • Multimodal ReID
  • Structural Priors
Research figure for ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs
Preprint 2026 Preprint

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

Yuhao Wang, Mu Qiao, Haiwen Diao, Yunzhi Zhuge, Pingping Zhang, Xindong Zhang, Lei Zhang, Huchuan Lu

Authorship First author

arXiv preprint

An entropy-guided, training-free pruning framework that preserves visual evidence by rectifying attention-logit collapse after token reduction.

Why this work matters
Contribution

ERA unifies head-wise entropy pruning, bias-aware token recycling, and logit-preserving attention rectification across single-image, multi-image, video, and vLLM serving settings.

  • Efficient MLLM
  • Visual Token Pruning
  • Attention Rectification
Research figure for STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification
Published 2026 Conference

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

Xingguo Xu†, Zhanyu Liu†, Weixiang Zhou†, Yuansheng Gao, Junjie Cao, Yuhao Wang*, Jixiang Luo, Dell Zhang

Authorship Co-corresponding author * Corresponding author

AAAI Conference on Artificial Intelligence (AAAI)

Segmentation priors modulate foreground tokens while a cross-modal hypergraph captures higher-order RGB/NIR/TIR relations.

Why this work matters
Contribution

Uses SAM-guided token redistribution and unified hypergraph interaction without relying on hard foreground deletion.

  • Multimodal ReID
  • Hypergraph
  • Segmentation Guidance
Research figure for VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
Preprint 2026 Preprint

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Qing’an Liu†, Juntong Feng†, Yuhao Wang†, Xinzhe Han, Yujie Cheng, Yue Zhu, Haiwen Diao, Yunzhi Zhuge, Huchuan Lu

Authorship Co-first author † Shared first authorship

arXiv preprint

A paired benchmark testing whether VLMs understand visualized text as reliably as equivalent pure text.

Why this work matters
Contribution

Builds 1,500 paired items and evaluates more than 30 VLMs, exposing a modality gap that widens with rendering difficulty.

  • Vision-Language Models
  • Benchmark
  • Visualized Text
Research figure for Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
Published 2026 Conference

Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification

Yangyang Liu†, Yuhao Wang†, Pingping Zhang*

Authorship Co-first author † Shared first authorship

AAAI Conference on Artificial Intelligence (AAAI)

Selective patch interaction and joint global-local alignment suppress background interference and improve consistency across three visual modalities.

Why this work matters
Contribution

Introduces selective interaction, Gramian-space global alignment, and shift-aware local alignment for multimodal object ReID.

  • Multimodal ReID
  • Token Selection
  • Global-local Alignment
Research figure for SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
Published 2026 Journal

SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification

Yuhao Wang, Xiang Hu, Lixin Wang, Pingping Zhang*, Huchuan Lu

Authorship First author

IEEE Transactions on Image Processing (TIP)

A view-aware generative framework that uses controllable diffusion priors to improve identity consistency across aerial and ground cameras.

Why this work matters
Contribution

Models view-specific feature distributions with a controllable Stable Diffusion model and a view-refined decoder across five aerial-ground ReID benchmarks.

  • Aerial–Ground ReID
  • Diffusion Models
  • Generative Learning
Research figure for HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation
Published 2026 Journal

HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation

Pingping Zhang, Tianyu Yan, Yuhao Wang, Yang Liu*, Tongdan Tang, Yili Ma, Long Lv, Feng Tian, Weibing Sun, Huchuan Lu

Authorship Contributor

IEEE Transactions on Image Processing (TIP)

Frequency-domain priors and full-view Mamba adapt SAM to fine-grained marine animal segmentation under complex underwater noise.

Why this work matters
Contribution

Combines a frequency-guided adapter, frequency-aware point selection, and linear-complexity contextual modeling.

  • Segmentation
  • Foundation Models
  • Mamba
Research figure for CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking
Published 2026 Conference

CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking

Hao Li, Yuhao Wang, Xiantao Hu, Wenning Hao, Pingping Zhang, Dong Wang, Huchuan Lu

Authorship Contributor

AAAI Conference on Artificial Intelligence (AAAI)

Contextual aggregation and deformable alignment improve robust visual tracking across visible and thermal modalities.

Why this work matters
Contribution

Combines linear-complexity cross-modal interaction, sparse expert aggregation, and deformable temporal alignment.

  • RGBT Tracking
  • Multimodal Perception
  • Deformable Alignment
Research figure for RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
Published 2026 Conference

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

Hao Li, Yuhao Wang, Wenning Hao*, Pingping Zhang*, Dong Wang, Huchuan Lu

Authorship Contributor

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Introduces MLLM-generated language and retrieval-augmented temporal reasoning into visible-thermal object tracking.

Why this work matters
Contribution

Builds text-enhanced RGBT benchmarks and combines adaptive token fusion with a dynamic visual-language knowledge base.

  • RGBT Tracking
  • Vision-Language
  • Retrieval-Augmented Generation
Research figure for SAS-VPReID: A Scale-Adaptive Framework with Shape Priors for Video-based Person Re-Identification at Extreme Far Distances
Published 2026 Workshop

SAS-VPReID: A Scale-Adaptive Framework with Shape Priors for Video-based Person Re-Identification at Extreme Far Distances

Qiwei Yang, Pingping Zhang, Yuhao Wang, Zijing Gong

Authorship Contributor

WACV Workshop on Video Re-Identification at Extreme Far Distances

A scale-adaptive video ReID framework combining low-resolution enhancement, multiscale temporal modeling, and clothing-robust shape priors.

Why this work matters
Contribution

The DLUT challenge solution ranked first in the associated extreme-far-distance video ReID track.

  • Video ReID
  • Extreme Distance
  • Shape Priors
Research figure for VReID-XFD: Video-based Person Re-identification at Extreme Far Distance Challenge Results
Published 2026 Challenge Report

VReID-XFD: Video-based Person Re-identification at Extreme Far Distance Challenge Results

Kailash A. Hambarde, Hugo Proença, Md Rashidunnabi, Pranita Samale, Qiwei Yang, Pingping Zhang, Zijing Gong, Yuhao Wang, et al.

Authorship Contributor

WACV Workshop on Video Re-Identification at Extreme Far Distances

Challenge report for extreme-far-distance video person ReID, including the DLUT SAS-VPReID solution.

  • Video ReID
  • Benchmarking
  • Challenge
Research figure for AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
Published 2025 Challenge Report

AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results

Kien Nguyen, Clinton Fookes, Sridha Sridharan, Huy Nguyen, Feng Liu, Xiaoming Liu, Arun Ross, et al., Yuhao Wang, Xuehu Liu, Pingping Zhang, et al.

Authorship Contributor

IEEE International Joint Conference on Biometrics (IJCB)

Challenge report for large-scale aerial-ground video person ReID; the associated DLUT/WUT team ranked second.

  • Aerial–Ground ReID
  • Benchmarking
Research figure for IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-Modal Object Re-Identification
Published 2025 Conference

IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-Modal Object Re-Identification

Yuhao Wang, Yongfeng Lv, Pingping Zhang*, Huchuan Lu

Authorship First author

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

A text-guided multimodal framework that uses cooperative deformable aggregation to bridge visual and language cues for object re-identification.

Why this work matters
Contribution

Constructs three MLLM-generated text-enhanced benchmarks and introduces inverted semantic guidance with adaptive local feature sampling.

  • Multimodal ReID
  • Cross-modal Retrieval
Research figure for CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-Identification
Published 2025 Conference

CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-Identification

Chenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang, Pingping Zhang*, Huchuan Lu

Authorship Contributor

AAAI Conference on Artificial Intelligence (AAAI)

A hybrid CLIP–Mamba framework for transferring language-aligned visual representations to person re-identification.

Why this work matters
Contribution

Combines CLIP semantics with Mamba sequence modeling for efficient person representation.

  • Person ReID
  • Vision-Language
Research figure for LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification
Preprint 2025 Preprint

LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification

Pingping Zhang*, Xiang Hu, Yuhao Wang, Huchuan Lu

Authorship Contributor

arXiv preprint

A parameter-efficient CLIP prompt-tuning framework that turns stable person attributes and camera viewpoints into structured text guidance.

Why this work matters
Contribution

Introduces attribute-aware image encoding, prompted attribute classification, and coupled text prompts for aerial-ground ReID.

  • Aerial–Ground ReID
  • Vision-Language
Research figure for Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification
Published 2025 Journal

Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification

Yuhao Wang, Pingping Zhang*, Xuehu Liu, Zhengzheng Tu, Huchuan Lu

Authorship First author

IEEE Transactions on Intelligent Transportation Systems (TITS)

A hybrid representation that combines convolutional locality with Transformer context for robust person re-identification.

Why this work matters
Contribution

Unifies local convolutional cues and global Transformer features rather than treating the two architecture families as alternatives.

  • Person ReID
  • Hybrid Architecture
Research figure for DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
Published 2025 Conference

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

Yuhao Wang, Yang Liu, Aihua Zheng, Pingping Zhang*

Authorship First author

AAAI Conference on Artificial Intelligence (AAAI)

A decoupled mixture-of-experts design that routes complementary modality features for more reliable multimodal object matching.

Why this work matters
Contribution

Preserves modality-specific knowledge through hierarchical feature decoupling and replaces static expert gates with attention-triggered routing.

  • Multimodal ReID
  • Mixture of Experts
Research figure for MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
Published 2025 Conference

MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt

Yuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu, Aihua Zheng, Pingping Zhang*, Huchuan Lu

Authorship First author

AAAI Conference on Artificial Intelligence (AAAI)

A Mamba-based aggregation and prompt learning framework for modeling long-range multimodal dependencies in object re-identification.

Why this work matters
Contribution

Adapts CLIP with parallel feed-forward adapters, synergistic residual prompts, and linear-complexity Mamba aggregation.

  • Multimodal ReID
  • Mamba
Research figure for Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation
Published · Oral 2025 Conference

Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

Zifu Wan, Pingping Zhang, Yuhao Wang, Silong Yong, Simon Stepputtis, Katia Sycara, Yaqi Xie

Authorship Contributor

IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

A Siamese Mamba architecture for fusing complementary modalities in dense semantic prediction.

Why this work matters
Contribution

Introduces Siamese state-space modeling for efficient long-range interaction in multimodal semantic segmentation.

  • Semantic Segmentation
  • Mamba
Ongoing 2025

VMambaX: Exploiting Vision Mamba for Abdominal X-ray Image Based NEC Classification

Zhoushan Feng†, Wen Long†, Yuhao Wang†, Shicun Qiao, Jingwen Mei, Yiyu Shi, Pingping Zhang*, Chunhong Jia, Fan Wu

Authorship Contributor

Manuscript

  • Medical Vision
  • Mamba
Research figure for TOP-ReID: Multi-spectral Object Re-Identification with Token Permutation
Published 2024 Conference

TOP-ReID: Multi-spectral Object Re-Identification with Token Permutation

Yuhao Wang, Xuehu Liu, Pingping Zhang*, Hu Lu, Zhengzheng Tu, Huchuan Lu

Authorship First author

AAAI Conference on Artificial Intelligence (AAAI)

A token permutation approach for learning identity-aware representations across visible and infrared spectra.

Why this work matters
Contribution

Uses token permutation to promote both cross-modal interaction and modality-specific representation.

  • Multi-spectral ReID
  • Token Learning
Research figure for Magic Tokens: Select Diverse Tokens for Multi-Modal Object Re-Identification
Published 2024 Conference

Magic Tokens: Select Diverse Tokens for Multi-Modal Object Re-Identification

Pingping Zhang*, Yuhao Wang, Yang Liu, Zhengzheng Tu, Huchuan Lu

Authorship Contributor

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

A token selection strategy that preserves diverse, identity-relevant cues for multimodal object re-identification.

Why this work matters
Contribution

Selects compact, diverse identity tokens before multimodal interaction to reduce background redundancy.

  • Multimodal ReID
  • Token Learning

Publication policy. Published work, work under review, and ongoing research are deliberately separated. Missing links are omitted rather than replaced with placeholders. Latest detailed metadata is verified against public records.