Meta tags:
description= GRIT: Teaching MLLMs to Think with Images;
keywords= GRIT, MLLMs, visual reasoning, grounded reasoning;
Headings (most frequently used words):
reasoning, with, to, the, grounded, model, mllms, natural, language, bounding, boxes, in, think, images, paradigm, chains, that, as, and, tokens, subsequent, grit, teaching, grpo, gr, reinforcement, learning, for, interleave, generate, of, needed, trained, an, process, generated, directly, on, without, any, image, then, abstract, overview, main, results, examples, bibtex, by, generating, key, innovations, models, interleaving, explicit, box, coordinates, algorithm, which, employs, novel, rewards, enable, ability, efficiently, no, manual, annotations, we, propose, light, efficient, way, pre, is, generates, one, continuous, pass, special, facilitate, extensible, during, autoregressive, generation, influencing, external, decoding, retrieval, flexibly, can, be, placed, anywhere, within, dynamic, number, from, zero, multiple, provides, answer, regions, reflects, it, first, grounds, critical, region, its, analyzes, correctly, handles, queries, about, non, existent, entities, grounding, action,
Text of the page (most frequently used words):
the (43), reasoning (38), and (27), with (25), answer (15), grit (14), think (14), bounding (14), image (12), for (12), #grounded (12), that (11), mllms (10), model (10), #coordinates (9), box (8), rethink (8), language (8), images (7), boxes (7), grpo (7), natural (7), chains (7), this (6), ground (6), truth (6), are (6), reinforcement (6), learning (6), there (5), knife (5), question (5), truck (5), zebras (5), models (5), explicit (5), generate (5), from (4), teaching (4), zheng (4), output (4), grounding (4), cat (4), regions (4), generated (4), gpt (4), trained (4), reward (4), tokens (4), paradigm (4), examples (3), subsequent (3), example (3), counting (3), effectively (3), visual (3), tasks (3), using (3), accuracy (3), directly (3), existing (3), datasets (3), approach (3), annotations (3), data (3), format (3), employs (3), process (3), final (3), interleave (3), which (3), website (2), licensed (2), under (2), creative (2), commons (2), attribution (2), sharealike (2), international (2), license (2), yue (2), fan (2), xuehai (2), diji (2), yang (2), kaizhi (2), ching (2), chen (2), kuo (2), yuting (2), sravana (2), jyothi (2), narayanaraju (2), xinze (2), guan (2), xin (2), eric (2), wang (2), 2505 (2), 15879 (2), arxiv (2), qwen2 (2), present (2), pot (2), soup (2), carrots (2), other (2), ingredients (2), would (2), outside (2), area (2), without (2), any (2), yes (2), top (2), beneath (2), its (2), then (2), six (2), 159 (2), 186 (2), results (2), efficiently (2), achieving (2), strong (2), performance (2), across (2), measures (2), correctness (2), judge (2), method (2), pre (2), comprehensive (2), evaluations (2), demonstrate (2), triplets (2), eliminates (2), need (2), labels (2), efficiency (2), few (2), training (2), number (2), count (2), special (2), key (2), needed (2), during (2), propose (2), algorithm (2), novel (2), rewards (2), ability (2), produce (2), visually (2), adapted, gligen, nerfies, misc, fan2025gritteachingmllmsthink, title, author, year, 2025, eprint, archiveprefix, primaryclass, url, https, org, abs, bibtex, inference, focus, shows, but, correctly, handles, queries, about, non, existent, entities, action, approximately, 209, 488, 364, positioned, first, grounds, critical, region, analyzes, picture, follows, 200, 168, 248, 202, 169, 214, 167, 108, 192, 173, 197, 163, 191, 413, 441, 189, 463, 171, 483, provided, accurate, cover, all, visible, overlapping, missing, how, many, pictured, here, provides, reflects, absence, handling, spatial, confirm, previously, separated, diverse, capabilities, unifies, alignment, between, iou, giou, score, acc, metrics, train, two, internvl, only, vsr, tallyqa, overall, outperform, baselines, testing, sets, main, remarkable, while, maintaining, combines, evaluation, bleu, similarity, aided, ans, verifies, matches, target, encourages, proper, use, structure, related, valid, syntax, optimizes, policy, sequences, three, components, generates, flexibly, dynamic, zero, multiple, can, placed, anywhere, within, influencing, external, decoding, retrieval, autoregressive, generation, facilitate, extensible, one, continuous, pass, response, continues, further, analysis, initial, light, efficient, way, generating, innovations, interleaving, enable, manual, overview, recent, studies, have, demonstrated, efficacy, building, articulate, thoughts, prior, producing, answers, however, despite, ongoing, advances, aim, enabling, vision, open, source, typically, content, pure, lacking, integration, information, limits, their, clearly, articulated, end, introduces, these, point, input, consults, additionally, equipped, built, upon, robust, focused, chain, result, achieves, exceptional, requiring, trains, coherent, showing, successful, unification, abilities, texts, abstract, live, demo, code, paper, ebay, santa, cruz,
Text of the page (random words):
grit teaching mllms to think with images grit teaching mllms to think with images yue fan 1 xuehai he 1 diji yang 1 kaizhi zheng 1 ching chen kuo 2 yuting zheng 2 sravana jyothi narayanaraju 2 xinze guan 2 xin eric wang 1 1 uc santa cruz 2 ebay paper code live demo abstract recent studies have demonstrated the efficacy of using reinforcement learning rl in building reasoning models that articulate chains of thoughts prior to producing final answers however despite ongoing advances that aim at enabling reasoning for vision language tasks existing open source visual reasoning models typically generate reasoning content with pure natural language lacking explicit integration of visual information this limits their ability to produce clearly articulated and visually grounded reasoning chains to this end we propose reasoning with images and texts grit a novel method for training mllms to think with images grit introduces a grounded reasoning paradigm in which models generate reasoning chains that interleave natural language and explicit bounding box coordinates these coordinates point to regions of the input image that the model consults during its reasoning process additionally grit is equipped with a reinforcement learning approach grpo gr built upon the grpo algorithm grpo gr employs robust rewards focused on the final answer accuracy and format of the grounded reasoning output which eliminates the need for data with reasoning chain annotations or explicit bounding box labels as a result grit achieves exceptional data efficiency requiring as few as 20 image question answer triplets from existing datasets comprehensive evaluations demonstrate that grit effectively trains mllms to produce coherent and visually grounded reasoning chains showing a successful unification of reasoning and grounding abilities overview grit teaching mllms to think with images by generating reasoning chains that interleave natural language with bounding boxes key innovations 1 grounded reasoning paradigm models generate reasoning chains interleaving natural language with explicit bounding box coordinates 2 grpo gr a reinforcement learning algorithm which employs novel rewards that enable the grounded reasoning ability of mllms efficiently no manual reasoning annotations needed grounded reasoning paradigm we propose grounded reasoning paradigm as a light and efficient way for pre trained mllms to think with images the model is trained to generate reasoning chains that interleave natural language with bounding boxes think initial natural language reasoning with bounding boxes generated x1 y1 x2 y2 think rethink reasoning continues with further analysis on image regions rethink answer final response the model generates reasoning in one continuous pass with special tokens to facilitate an extensible reasoning process bounding boxes tokens generated directly during autoregressive generation influencing on subsequent reasoning directly as tokens without any external decoding image retrieval process flexibly generated can be placed anywhere within the natural language reasoning with dynamic number of boxes from zero to multiple as needed grpo gr reinforcement learning for grounded reasoning grit employs grpo gr a reinforcement learning approach that optimizes a policy π θ to generate grounded reasoning sequences with three key reward components 1 grounded reasoning format reward r format encourages proper use of reasoning structure related special tokens and valid bounding box syntax 2 grounded target counting reward r count for counting tasks verifies the number of bounding boxes matches the ground truth count 3 gpt aided answer accuracy reward r ans combines gpt 4o evaluation of answer correctness with bleu similarity to ground truth this approach eliminates the need for explicit reasoning annotations or bounding box labels achieving remarkable data efficiency with as few as 20 training examples while maintaining strong performance main results directly using the grit method with grpo gr reinforcement learning we train two pre trained mllms qwen2 5 vl 3b and internvl 3 2b with only 20 image question answer triplets from existing datasets vsr and tallyqa comprehensive evaluations demonstrate that grit trained models overall outperform baselines across six testing sets datasets metrics gpt as judge answer accuracy acc measures answer correctness using gpt 4o as judge 0 1 score grounding iou giou measures alignment between model generated bounding boxes and ground truth regions results confirm that grit effectively unifies previously separated grounding and reasoning capabilities in mllms efficiently and effectively achieving strong performance across diverse visual reasoning tasks examples example 1 counting example 2 spatial reasoning example 3 absence handling model provides an answer with regions and then reflects on it in the subsequent reasoning question how many zebras are pictured here model output think there are six zebras in the picture the coordinates for the zebras are as follows 1 200 168 248 202 2 169 159 214 186 3 76 167 108 192 4 24 173 50 197 5 51 163 70 191 6 413 159 441 189 7 463 171 483 186 think rethink the coordinates provided for the zebras are accurate and cover all the zebras visible in the image there are no overlapping or missing coordinates rethink answer 7 ground truth answer 7 model first grounds critical image region in its reasoning and then analyzes in the subsequent reasoning question is the truck beneath the cat model output think the bounding box coordinates for the truck are approximately 0 209 488 364 the cat is positioned on top of this truck think rethink the cat is on top of the truck so the truck is beneath the cat rethink answer yes ground truth answer yes model correctly handles queries about non existent entities without any grounding action question is there a knife in the image model output think there is no knife present in the image the focus is on a pot of soup with carrots and other ingredients the bounding box coordinates for the knife would be outside the image area think rethink the image shows a pot of soup with carrots and other ingredients but there is no knife present the bounding box coordinates for the knife would be outside the image area rethink answer no ground truth answer no inference examples of qwen2 5 vl grit bibtex misc fan2025gritteachingmllmsthink title grit teaching mllms to think with images author yue fan and xuehai he and diji yang and kaizhi zheng and ching chen kuo and yuting zheng and sravana jyothi narayanaraju and xinze guan and xin eric wang year 2025 eprint 2505 15879 archiveprefix arxiv primaryclass cs cv url https arxiv org abs 2505 15879 this website is adapted from nerfies and gligen licensed under a creative commons attribution sharealike 4 0 international license this website is licensed under a creative commons attribution sharealike 4 0 international license
|