MONKEY: Masking ON KEY-Value Activation Adapter for Personalization

dc.contributor.authorBaker, James
dc.date.accessioned2025-11-21T00:29:38Z
dc.date.issued2025-10-09
dc.description.abstractPersonalizing diffusion models allows users to generate new images that incorporate a given subject, allowing more control than a text prompt. These models often suffer somewhat when they end up just recreating the subject image, and ignoring the text prompt. We observe that one popular method for personalization, the IP-Adapter automatically generates masks that we definitively segment the subject from the background during inference. We propose to use this automatically generated mask on a second pass to mask the image tokens, thus restricting them to the subject, not the background, allowing the text prompt to attend to the rest of the image. For text prompts describing locations and places, this produces images that accurately depict the subject while definitively matching the prompt. We compare our method to a few other test time personalization methods, and find our method displays high prompt and source image alignment.
dc.description.urihttp://arxiv.org/abs/2510.07656
dc.format.extent13 pages
dc.genrejournal articles
dc.genrepreprints
dc.identifierdoi:10.13016/m2s2nz-x6lt
dc.identifier.urihttps://doi.org/10.48550/arXiv.2510.07656
dc.identifier.urihttp://hdl.handle.net/11603/40771
dc.language.isoen
dc.relation.isAvailableAtThe University of Maryland, Baltimore County (UMBC)
dc.relation.ispartofUMBC Computer Science and Electrical Engineering Department
dc.relation.ispartofUMBC Student Collection
dc.rightsAttribution 4.0 International
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subjectComputer Science - Computer Vision and Pattern Recognition
dc.titleMONKEY: Masking ON KEY-Value Activation Adapter for Personalization
dc.typeText

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
251007656v1.pdf
Size:
29.1 MB
Format:
Adobe Portable Document Format