Thao Nguyen (Shibe)

Thao at Googleplex  

Hi, I'm Thao! 👋
I'm a CS PhD student @ UW-Madison, working with Prof. Yong Jae Lee, and fortunate to have support from senior students: Dr. Prof. Utkarsh Ojha, Dr. Yuheng Li, Dr. Haotian Liu, and Dr. Mu Cai.
I'm also lucky enough to have maintained a long-term collaboration with Adobe Research: Dr. Krishna Kumar Singh, Dr. Eli Shechtman, Dr. Yilin Wang, and many other inspiring folks (that space here is not enough to list all... sorry 😅).

> I work on building Personal AI Agent—one that understands your personal context, empowers your creativity + self-efficacy, and remains reliable, safe, & trustworthy.
To do this, such an agent must:
    . adapt seamlessly to user's inputs/ intentions [NeurIPS'23, CVPR'24, CVPR'26];
    . understand users' knowledge/ environments [NeurIPS'24, CVPR'25];
    . and reason over user's data to extract relevant, helpful memory — so it can respond with personal context, in a more personalized and context-aware way [camroll (arXiv'26), visualmem (arXiv'26), and... tbd ;)].
Today's models are powerful but general-purpose (fig, left); my research steps toward adapting them into personal (and agentic) ones (fig, right) (.❛ ᴗ ❛.).

Email / GitHub / X / WeChat / Google Scholar / Resume/CV

I’ll be on the job market for Research Scientist/Engineer positions starting in 2027. thao.nguyen@wisc.edu 🤗

About Me

4th year PhD Student – Univeristy of Wisconsin - Madison.
Contact: thao.nguyen at wisc dot edu | thao.ntp0414 at gmail dot com.
Research interests: Personalized AI; MLLMs; (+ AI for Education); Recent: Personal AI Agent, Visual Memory.

research interests

Publications [my favorites | still my favorites, but all ]

CamRoll teaser  

Personal AI Agent for Camera Roll VQA
if an AI could see your whole camera roll, what would you ask?
Thao Nguyen, Krishna Kumar Singh, Donghyun Kim, Yong Jae Lee†, Yuheng Li†
arXiv preprint, 2026
(†: equal advising)
[ProjectPage, Demo, Code, Paper]

VisualMem teaser  

VisualMem: Personal Visual Memory from Explicit and Implicit Evidence
building a personal visual memory of a user from explicit and implicit evidence in their photos.
Viet Nguyen, Thao Nguyen, Vishal M. Patel†, Yuheng Li†
arXiv preprint, 2026
(†: equal advising)
[ProjectPage, Code, Paper]

relsim teaser  

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
letting diffusion samples collaborate across a batch to improve image generation.
Sicheng Mo,Thao Nguyen, Richard Zhang, Nicholas Kolkin, Siddharth Srinivasan Iyer, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
[ProjectPage, Code, Paper]

relsim teaser  

Relational Visual Similarity
measuring visual similarity through relations — closer to how people intuitively compare images.
Thao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang, Jing Shi, Nicholas Kolkin, Eli Shechtman, Yong Jae Lee†, Yuheng Li†
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Best Paper Award @ 🏆 CVPR 2026 Workshop: "CogVL: Cognitive Foundations for Multimodal Models"
(†: equal advising)
[ProjectPage, Poster, Code, Paper]

X-Fusion teaser  

X-Fusion: Introducing New Modality to Frozen Large Language Models
adding a new (vision) modality to a frozen large language model, without retraining the LLM.
Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li
International Conference on Computer Vision (ICCV), 2025
Best Paper Award @ 🏆 CVPR 2025 Workshop: "Transformers for Vision (T4V)"
[ProjectPage, Code, Paper]

Yo'Chameleon teaser  

Yo'Chameleon: Personalized Vision and Language Generation
personalizing a unified multimodal model to understand and generate a specific user's concept.
Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui, Yong Jae Lee†, Yuheng Li†
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
(†: equal advising; also accepted at 🧷 CVPR 2025 Workshop: "What is Next in Multimodal Foundation Models?")
[ProjectPage, Poster, Code, Paper]

Yo'LLaVA teaser  

Yo'LLaVA: Your Personalized Language and Vision Assistant
teaching a multimodal LLM your personal concepts (e.g., your pet) from just a few images.
Thao Nguyen, Haotian Liu, Yuheng Li, Mu Cai, Utkarsh Ojha, Yong Jae Lee
Neural Information Processing Systems (NeurIPS), 2024
[ProjectPage, Poster, Code, Paper]

Edit One for All teaser  

Edit One for All: Interactive Batch Image Editing
make one edit on a single image, then propagate it across a whole batch interactively.
Thao Nguyen, Utkarsh Ojha, Yuheng Li, Haotian Liu, Yong Jae Lee
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
[ProjectPage, Poster, Code, Paper]

VISII teaser  

Visual Instruction Inversion: Image Editing via Visual Prompting
editing an image from a visual before-and-after example, instead of a text instruction.
Thao Nguyen, Yuheng Li, Utkarsh Ojha, Yong Jae Lee
Neural Information Processing Systems (NeurIPS), 2023
[ProjectPage, Poster, Code, Paper]

CPM teaser  

Lipstick ain't enough: Beyond Color Matching for In-the-Wild Makeup Transfer
in-the-wild makeup transfer that captures patterns and textures, not just color.
Thao Nguyen, Anh Tran, Minh Hoai
Conference on Computer Vision and Pattern Recognition (CVPR), 2021
[ProjectPage, Code, Video, Paper]

Not CS publication :D

ICQE poster  

Are Dining Expectations Culturally Conditioned? An analysis of Asian vs. American Restaurants (yes, it's not a CS conference :D)
Thao Nguyen (+Yuheng Li)
International Conference on Quantitative Ethnography (ICQE), 2025
Best poster nomination 🏆 | work done during my PhD minor course with Prof. David Williamson Shaffer 😊
[Poster, Paper]

Outreach & Side Project(s)

When I make no progress on my research 🐢, I create and maintain these websites:


and these GitHub's awesome lists:

If you're a Vietnamese student wanting to apply for a CS/EE PhD in the US, you might find viet-wics's blog posts helpful.
Or, if you're curious, I talked about my "expected" unexpected PhD experience in this YouTube video. Good luck 🐢

Misc

Mam the Cat  
--- Meet my pets: Bo (2021-2026) and Mam~ 🐶🐱 (.❛ ᴗ ❛.)

"🎉 Bo and Mam have appeared on multiple figures of my papers. Can you spot them? 😉"

neurips2023 cvpr2024 neurips2024 cvpr2025 iccv2025 iccv2025
Bo the Shiba