Meta tags:
Headings (most frequently used words):
prompt, to, image, attention, examples, editing, cross, control, with, abstract, text, style, transfer, followup, works, word, swap, refinement, re, weighting, bibtex, acknowledgements,
Text of the page (most frequently used words):
the (70), #prompt (29), image (21), editing (19), and (18), #attention (18), with (13), text (13), cross (9), control (9), our (8), that (7), this (7), models (7), diffusion (7), maps (7), images (6), tokens (6), for (5), original (5), which (5), word (5), using (4), while (4), new (4), words (4), only (4), synthesis (4), their (3), from (3), can (3), preserve (3), specific (3), examples (3), global (3), swap (3), process (3), controlling (3), approach (3), semantic (3), cat (3), enables (3), generated (3), method (3), spatial (3), are (3), hat (3), prompts (3), imagen (2), hertz (2), amir (2), mokady (2), ron (2), tenenbaum (2), jay (2), aberman (2), kfir (2), pritch (2), yael (2), cohen (2), daniel (2), arxiv (2), real (2), first (2), input (2), null (2), follow (2), adding (2), style (2), injecting (2), various (2), structure (2), extent (2), weighting (2), perform (2), refinement (2), basket (2), therefore (2), main (2), during (2), pixels (2), attend (2), applications (2), show (2), methods (2), modify (2), results (2), key (2), observation (2), layout (2), below (2), bunny (2), doll (2), example (2), modifying (2), over (2), lying (2), beach (2), chair (2), textual (2), model (2), driven (2), capabilities (2), diverse (2), these (2), provide (2), content (2), based (2), even (2), different (2), paper (2), edited (2), thank, noa, glaser, adi, zicher, yaron, brodsky, shlomi, fruchter, david, salesin, valuable, inputs, helped, improve, work, mohammad, norouzi, chitwan, saharia, william, chan, providing, support, pretrained, website, template, borrowed, dreambooth, acknowledgements, article, hertz2022prompt, title, author, booktitle, preprint, 2208, 01626, year, 2022, bibtex, applying, inverting, pivotal, optimization, training, instruction, network, synthetic, data, obtained, combining, gpt3, stable, instructpix2pix, learning, instructions, inversion, guided, followup, works, description, source, create, desired, styles, transfer, reducing, increasing, marked, arrow, influences, generation, extending, initial, local, case, others, apples, oranges, idea, inject, steps, apply, creative, several, through, simple, interface, token, dog, fixing, scene, composition, second, add, freeze, previous, allowing, flow, object, third, increase, decrease, weights, specified, amplification, attenuation, effect, behind, geometry, depend, more, describe, them, fluffy, another, influence, amplify, attenuate, fluffiness, unicorn, clown, musketeer, pirates, swimming, straw, police, cylinder, floral, koala, giraffe, elephant, turtle, horse, lion, leopard, here, generate, then, easily, replace, character, recent, large, scale, have, attracted, much, thanks, remarkable, generating, highly, given, natural, build, upon, however, challenging, generative, since, innate, property, technique, some, small, modification, often, leads, completely, outcome, state, art, mitigate, requiring, users, mask, localize, edit, hence, ignoring, within, masked, region, pursue, intuitive, framework, where, edits, controlled, analyze, conditioned, depth, observe, layers, relation, between, each, propose, along, monitor, paving, way, myriad, caption, such, localized, replacing, specification, reflected, present, demonstrating, high, quality, fidelity, abstract, code, google, research, tel, aviv, university,
Text of the page (random words):
prompt to prompt prompt to prompt image editing with cross attention control amir hertz 1 2 ron mokady 1 2 jay tenenbaum 1 kfir aberman 1 yael pritch 1 daniel cohen or 1 2 1 google research 2 tel aviv university paper code abstract recent large scale text driven synthesis diffusion models have attracted much attention thanks to their remarkable capabilities of generating highly diverse images that follow given text prompts therefore it is only natural to build upon these synthesis models to provide text driven image editing capabilities however editing is challenging for these generative models since an innate property of an editing technique is to preserve some content from the original image while in the text based models even a small modification of the text prompt often leads to a completely different outcome state of the art methods mitigate this by requiring the users to provide a spatial mask to localize the edit hence ignoring the original structure and content within the masked region in this paper we pursue an intuitive prompt to prompt editing framework where the edits are controlled by text only we analyze a text conditioned model in depth and observe that the cross attention layers are the key to controlling the relation between the spatial layout of the image to each word in the prompt with this observation we propose to control the attention maps of the edited image by injecting the attention maps of the original image along the diffusion process our approach enables us to monitor the synthesis process by editing the textual prompt only paving the way to a myriad of caption based editing applications such as localized editing by replacing a word global editing by adding a specification and even controlling the extent to which a word is reflected in the image we present our results over diverse images and prompts with different text to image models demonstrating high quality synthesis and fidelity to the edited prompts prompt to prompt image editing our method enables editing generated images by only modifying the textual prompt for example here we first generate an image from the input prompt a cat with a hat is lying on a beach chair using the imagen text to image diffusion model then with our approach we can easily replace the hat or the main character a cat leopard lion horse turtle elephant giraffe koala with a floral cylinder police straw swimming pirates musketeer clown unicorn hat is lying on a beach chair another prompt editing example is modifying the semantic influence of specific words in the prompt over the generated image using our method we can amplify or attenuate the fluffiness of the bunny doll in the image below my fluffy bunny doll cross attention control the key observation behind our method is that the spatial layout and geometry of an image depend on the cross attention maps below we show that pixels are attend more to the words that describe them therefore our main idea is to inject the cross attention maps during the diffusion process controlling which pixels attend to which tokens of the prompt text during which diffusion steps to apply our approach to various creative editing applications we show several methods to control the cross attention maps through a simple and semantic interface in the word swap control we modify a token in the prompt e g dog to cat while fixing the cross attention maps to preserve the scene composition in the second prompt refinement control we add new words to the prompt and freeze the attention to previous tokens while allowing new attention to flow to the new tokens this enables us to perform global editing or modify a specific object in the third attention re weighting control we increase or decrease the attention weights of specified tokens this results with amplification or attenuation of the semantic effect of the tokens on the generated image word swap examples in this case we swap tokens of the original prompt with others e g a basket with apples to a basket with oranges prompt refinement examples by extending the initial prompt we perform local or global editing attention re weighting examples by reducing or increasing the cross attention of specific words marked with an arrow we control the extent to which it influences the generation text to image style transfer by adding a style description to the prompt while injecting the source attention maps we can create various images in the new desired styles that preserve the structure of the original image followup works null text inversion for editing real images using guided diffusion models applying prompt to prompt editing on real images by first inverting an input image using pivotal null text optimization instructpix2pix learning to follow image editing instructions training an instruction to image network on synthetic data obtained by combining gpt3 and prompt to prompt on stable diffusion bibtex article hertz2022prompt title prompt to prompt image editing with cross attention control author hertz amir and mokady ron and tenenbaum jay and aberman kfir and pritch yael and cohen or daniel booktitle arxiv preprint arxiv 2208 01626 year 2022 acknowledgements we thank noa glaser adi zicher yaron brodsky shlomi fruchter and david salesin for their valuable inputs that helped improve this work and to mohammad norouzi chitwan saharia and william chan for providing us with their support and the pretrained models of imagen the website template is borrowed from dreambooth
|