RESEARCH / ONE CONNECTED PROGRAM
Perceive broadly.
Compute deliberately.
Act reliably.
My research begins with a simple systems question: how can multimodal intelligence preserve the right evidence, spend computation where it matters, and remain useful under real-world constraints?
INTERACTIVE RESEARCH MAP
A research system, not a keyword list.
Select a node to trace its questions, methods, papers, and projects.
Intelligence Perceive · Compress · Act
Established RESEARCH THEME
Multimodal Perception
How can heterogeneous sensors describe the same identity under changing viewpoints, spectra, and environments?
I align RGB, near-infrared, thermal, text, and aerial-ground observations into identity-aware representations that remain dependable in complex scenes.
- Cross-modal alignment
- Modality-specific representation
- Text-guided fusion
Growing RESEARCH THEME
Multimodal Foundation Models
How can vision-language priors improve multimodal understanding without erasing task-specific structure?
I study CLIP, VLMs, and multimodal large language models as semantic priors for perception, retrieval, and interface understanding.
- Vision-language adaptation
- Prompt learning
- Semantic knowledge guidance
Growing RESEARCH THEME
Efficient Intelligence
Which visual tokens, states, and memories are truly necessary for accurate multimodal inference?
I reduce redundant computation through token selection, state-space modeling, mixture-of-experts, compression, and cache-aware inference.
- Token pruning and compression
- Mamba / state-space models
- KV and visual cache
Emerging RESEARCH THEME
On-device GUI Agents
How can multimodal agents perceive, remember, and act on interfaces within strict latency, memory, and package-size budgets?
My recent work connects GUI understanding with training-free model pruning and deployment-oriented optimization for resource-constrained devices.
- Training-free pruning
- Visual memory efficiency
- Device-aware optimization
Future RESEARCH THEME
Cloud–Edge Agent Systems
How should perception, reasoning, memory, and action be distributed between device and cloud?
A forward-looking direction toward reliable long-horizon agents that adapt computation to device constraints, connectivity, privacy, and task difficulty.
- Cloud–edge collaboration
- Adaptive computation
- Long-horizon decision-making
RESEARCH PHILOSOPHY
Four principles behind the work.
Perceive beyond pixels.
Multimodal systems should preserve the evidence unique to each sensor while discovering the semantics they share.
Spend compute where it matters.
Efficiency is an information-design problem: select useful tokens, route useful features, and reuse useful memory.
Design for the device.
A method is more valuable when latency, memory, and deployment constraints influence the research question from the beginning.
Connect perception to action.
The next step is not only to understand multimodal environments, but to help agents make reliable decisions within them.
WHERE I’M GOING
From efficient models to reliable agent systems.
The next research questions build directly on the same through-line: multimodal evidence, efficient computation, and deployment-aware decisions.
Current · Now
Efficient multimodal models
Reduce token, feature, and memory redundancy while preserving cross-modal understanding.
Next · Near term
Long-horizon GUI agents
Build agents that perceive interfaces, retain useful visual history, and act under real device budgets.
Future · Research vision
Cloud–edge multimodal agent systems
Coordinate perception, reasoning, memory, and action across device and cloud for reliable autonomous decision-making.
COLLABORATE