Thao Nguyen (Shibe)

Thao at Googleplex  

Hi, I'm Thao! 👋
I'm a CS PhD student @ UW-Madison, researching personal AI agents, personal visual memory, and long-horizon multimodal reasoning with Prof. Yong Jae Lee, and fortunate to have support from senior students: Dr. Prof. Utkarsh Ojha, Dr. Yuheng Li, Dr. Haotian Liu, and Dr. Mu Cai.
I'm also lucky enough to have maintained a long-term collaboration with Adobe Research: Dr. Krishna Kumar Singh, Dr. Eli Shechtman, Dr. Yilin Wang, and many other inspiring folks (that space here is not enough to list all... sorry 😅).

> I work on building personal AI agents—systems that understand your personal context, empower your creativity + self-efficacy, and remain reliable, safe, & trustworthy.
To do this, such an agent must:
    . adapt seamlessly to user's inputs/ intentions [NeurIPS'23, CVPR'24, CVPR'26];
    . understand users' knowledge/ environments [NeurIPS'24, CVPR'25];
    . and reason over user's data to extract relevant, helpful memory — so it can respond with personal context, in a more personalized and context-aware way [camroll (arXiv'26), visualmem (arXiv'26), and... tbd ;)].
Today's models are powerful but general-purpose (fig, left); my research steps toward adapting them into personal (and agentic) ones (fig, right) (.❛ ᴗ ❛.).

Email / GitHub / X / WeChat / Google Scholar / Resume/CV

I’ll be on the job market for Research Scientist/Engineer positions starting in 2027. thao.nguyen@wisc.edu 🤗

About Me

4th year PhD Student – University of Wisconsin–Madison.
Contact: thao.nguyen at wisc dot edu | thao.ntp0414 at gmail dot com.
Research interests: Personalized AI; MLLMs; (+ AI for Education); Recent: Personal AI Agents, Visual Memory.

Thao Nguyen's research vision for personal AI agents and visual memory

Publications [my favorites | still my favorites, but all ]

CamRoll teaser  

Personal AI Agent for Camera Roll VQA
a conversational AI agent with hierarchical memory and tools for reasoning over years of personal photos.
Thao Nguyen, Krishna Kumar Singh, Donghyun Kim, Yong Jae Lee†, Yuheng Li†
arXiv preprint, 2026
(†: equal advising)
[ProjectPage, Demo, Code, Paper]

VisualMem teaser  

VisualMem: Personal Visual Memory from Explicit and Implicit Evidence
building personal visual memory from explicit and implicit evidence in users' photos.
Viet Nguyen, Thao Nguyen, Vishal M. Patel†, Yuheng Li†
arXiv preprint, 2026
(†: equal advising)
[ProjectPage, Code, Paper]

relsim teaser  

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
enabling diffusion samples to collaborate through cross-sample attention.
Sicheng Mo,Thao Nguyen, Richard Zhang, Nicholas Kolkin, Siddharth Srinivasan Iyer, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
[ProjectPage, Code, Paper]

relsim teaser  

Relational Visual Similarity
measuring similarity through relations, not just appearance.
Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee†, Yuheng Li†
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Best Paper Award @ 🏆 CVPR 2026 Workshop: "CogVL: Cognitive Foundations for Multimodal Models"
(†: equal advising)
[ProjectPage, Poster, Code, Paper]

X-Fusion teaser  

X-Fusion: Introducing New Modality to Frozen Large Language Models
adding visual understanding and generation to a frozen LLM.
Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li
International Conference on Computer Vision (ICCV), 2025
Best Paper Award @ 🏆 CVPR 2025 Workshop: "Transformers for Vision (T4V)"
[ProjectPage, Code, Paper]

Yo'Chameleon teaser  

Yo'Chameleon: Personalized Vision and Language Generation
personalizing an MLLM to understand and generate a user-specific concept.
Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui, Yong Jae Lee†, Yuheng Li†
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
(†: equal advising; also accepted at 🧷 CVPR 2025 Workshop: "What is Next in Multimodal Foundation Models?")
[ProjectPage, Poster, Code, Paper]

Yo'LLaVA teaser  

Yo'LLaVA: Your Personalized Language and Vision Assistant
teaching a VLM to converse about a user-specific subject from a few images.
Thao Nguyen, Haotian Liu, Yuheng Li, Mu Cai, Utkarsh Ojha, Yong Jae Lee
Neural Information Processing Systems (NeurIPS), 2024
[ProjectPage, Poster, Code, Paper]

Edit One for All teaser  

Edit One for All: Interactive Batch Image Editing
make one edit on a single image, then propagate it across a whole batch interactively.
Thao Nguyen, Utkarsh Ojha, Yuheng Li, Haotian Liu, Yong Jae Lee
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
[ProjectPage, Poster, Code, Paper]

VISII teaser  

Visual Instruction Inversion: Image Editing via Visual Prompting
editing an image from a visual before-and-after example, instead of a text instruction.
Thao Nguyen, Yuheng Li, Utkarsh Ojha, Yong Jae Lee
Neural Information Processing Systems (NeurIPS), 2023
[ProjectPage, Poster, Code, Paper]

CPM teaser  

Lipstick ain't enough: Beyond Color Matching for In-the-Wild Makeup Transfer
in-the-wild makeup transfer that captures patterns and textures, not just color.
Thao Nguyen, Anh Tran, Minh Hoai
Conference on Computer Vision and Pattern Recognition (CVPR), 2021
[ProjectPage, Code, Video, Paper]

Not CS publication :D

ICQE poster  

Are Dining Expectations Culturally Conditioned? An analysis of Asian vs. American Restaurants (yes, it's not a CS conference :D)
Thao Nguyen (+Yuheng Li)
International Conference on Quantitative Ethnography (ICQE), 2025
Best poster nomination 🏆 | work done during my PhD minor course with Prof. David Williamson Shaffer 😊
[Poster, Paper]

Outreach & Side Project(s)

When I make no progress on my research 🐢, I create and maintain these websites:


and these GitHub's awesome lists:

If you're a Vietnamese student wanting to apply for a CS/EE PhD in the US, you might find viet-wics's blog posts helpful.
Or, if you're curious, I talked about my "expected" unexpected PhD experience in this YouTube video. Good luck 🐢

Misc

Mam the Cat  
--- Meet my pets: Bo (2021-2026) and Mam~ 🐶🐱 (.❛ ᴗ ❛.)

"🎉 Bo and Mam have appeared on multiple figures of my papers. Can you spot them? 😉"

neurips2023 cvpr2024 neurips2024 cvpr2025 iccv2025 cvpr2026 camroll 2026
Bo the Shiba