If you are not sure if the website you would like to visit is secure, you can verify it here. Enter the website address of the page and see parts of its content and the thumbnail images on this site. None (if any) dangerous scripts on the referenced page will be executed. Additionally, if the selected site contains subpages, you can verify it (review) in batches containing 5 pages.
favicon.ico: grounded-reasoning.github.io - GRIT: Teaching MLLMs to Think .

site address: grounded-reasoning.github.io redirected to: grounded-reasoning.github.io

site title: GRIT: Teaching MLLMs to Think with Images

Our opinion (on Wednesday 22 July 2026 4:02:07 UTC):

GREEN status (no comments) - no comments
After content analysis of this website we propose the following hashtags:



Meta tags:
description=GRIT: Teaching MLLMs to Think with Images;
keywords=GRIT, MLLMs, visual reasoning, grounded reasoning;

Headings (most frequently used words):

reasoning, with, to, the, grounded, model, mllms, natural, language, bounding, boxes, in, think, images, paradigm, chains, that, as, and, tokens, subsequent, grit, teaching, grpo, gr, reinforcement, learning, for, interleave, generate, of, needed, trained, an, process, generated, directly, on, without, any, image, then, abstract, overview, main, results, examples, bibtex, by, generating, key, innovations, models, interleaving, explicit, box, coordinates, algorithm, which, employs, novel, rewards, enable, ability, efficiently, no, manual, annotations, we, propose, light, efficient, way, pre, is, generates, one, continuous, pass, special, facilitate, extensible, during, autoregressive, generation, influencing, external, decoding, retrieval, flexibly, can, be, placed, anywhere, within, dynamic, number, from, zero, multiple, provides, answer, regions, reflects, it, first, grounds, critical, region, its, analyzes, correctly, handles, queries, about, non, existent, entities, grounding, action,

Text of the page (most frequently used words):
the (43), reasoning (38), and (27), with (25), answer (15), grit (14), think (14), bounding (14), image (12), for (12), #grounded (12), that (11), mllms (10), model (10), #coordinates (9), box (8), rethink (8), language (8), images (7), boxes (7), grpo (7), natural (7), chains (7), this (6), ground (6), truth (6), are (6), reinforcement (6), learning (6), there (5), knife (5), question (5), truck (5), zebras (5), models (5), explicit (5), generate (5), from (4), teaching (4), zheng (4), output (4), grounding (4), cat (4), regions (4), generated (4), gpt (4), trained (4), reward (4), tokens (4), paradigm (4), examples (3), subsequent (3), example (3), counting (3), effectively (3), visual (3), tasks (3), using (3), accuracy (3), directly (3), existing (3), datasets (3), approach (3), annotations (3), data (3), format (3), employs (3), process (3), final (3), interleave (3), which (3), website (2), licensed (2), under (2), creative (2), commons (2), attribution (2), sharealike (2), international (2), license (2), yue (2), fan (2), xuehai (2), diji (2), yang (2), kaizhi (2), ching (2), chen (2), kuo (2), yuting (2), sravana (2), jyothi (2), narayanaraju (2), xinze (2), guan (2), xin (2), eric (2), wang (2), 2505 (2), 15879 (2), arxiv (2), qwen2 (2), present (2), pot (2), soup (2), carrots (2), other (2), ingredients (2), would (2), outside (2), area (2), without (2), any (2), yes (2), top (2), beneath (2), its (2), then (2), six (2), 159 (2), 186 (2), results (2), efficiently (2), achieving (2), strong (2), performance (2), across (2), measures (2), correctness (2), judge (2), method (2), pre (2), comprehensive (2), evaluations (2), demonstrate (2), triplets (2), eliminates (2), need (2), labels (2), efficiency (2), few (2), training (2), number (2), count (2), special (2), key (2), needed (2), during (2), propose (2), algorithm (2), novel (2), rewards (2), ability (2), produce (2), visually (2), adapted, gligen, nerfies, misc, fan2025gritteachingmllmsthink, title, author, year, 2025, eprint, archiveprefix, primaryclass, url, https, org, abs, bibtex, inference, focus, shows, but, correctly, handles, queries, about, non, existent, entities, action, approximately, 209, 488, 364, positioned, first, grounds, critical, region, analyzes, picture, follows, 200, 168, 248, 202, 169, 214, 167, 108, 192, 173, 197, 163, 191, 413, 441, 189, 463, 171, 483, provided, accurate, cover, all, visible, overlapping, missing, how, many, pictured, here, provides, reflects, absence, handling, spatial, confirm, previously, separated, diverse, capabilities, unifies, alignment, between, iou, giou, score, acc, metrics, train, two, internvl, only, vsr, tallyqa, overall, outperform, baselines, testing, sets, main, remarkable, while, maintaining, combines, evaluation, bleu, similarity, aided, ans, verifies, matches, target, encourages, proper, use, structure, related, valid, syntax, optimizes, policy, sequences, three, components, generates, flexibly, dynamic, zero, multiple, can, placed, anywhere, within, influencing, external, decoding, retrieval, autoregressive, generation, facilitate, extensible, one, continuous, pass, response, continues, further, analysis, initial, light, efficient, way, generating, innovations, interleaving, enable, manual, overview, recent, studies, have, demonstrated, efficacy, building, articulate, thoughts, prior, producing, answers, however, despite, ongoing, advances, aim, enabling, vision, open, source, typically, content, pure, lacking, integration, information, limits, their, clearly, articulated, end, introduces, these, point, input, consults, additionally, equipped, built, upon, robust, focused, chain, result, achieves, exceptional, requiring, trains, coherent, showing, successful, unification, abilities, texts, abstract, live, demo, code, paper, ebay, santa, cruz,


Text of the page (random words):
grit teaching mllms to think with images grit teaching mllms to think with images yue fan 1 xuehai he 1 diji yang 1 kaizhi zheng 1 ching chen kuo 2 yuting zheng 2 sravana jyothi narayanaraju 2 xinze guan 2 xin eric wang 1 1 uc santa cruz 2 ebay paper code live demo abstract recent studies have demonstrated the efficacy of using reinforcement learning rl in building reasoning models that articulate chains of thoughts prior to producing final answers however despite ongoing advances that aim at enabling reasoning for vision language tasks existing open source visual reasoning models typically generate reasoning content with pure natural language lacking explicit integration of visual information this limits their ability to produce clearly articulated and visually grounded reasoning chains to this end we propose reasoning with images and texts grit a novel method for training mllms to think with images grit introduces a grounded reasoning paradigm in which models generate reasoning chains that interleave natural language and explicit bounding box coordinates these coordinates point to regions of the input image that the model consults during its reasoning process additionally grit is equipped with a reinforcement learning approach grpo gr built upon the grpo algorithm grpo gr employs robust rewards focused on the final answer accuracy and format of the grounded reasoning output which eliminates the need for data with reasoning chain annotations or explicit bounding box labels as a result grit achieves exceptional data efficiency requiring as few as 20 image question answer triplets from existing datasets comprehensive evaluations demonstrate that grit effectively trains mllms to produce coherent and visually grounded reasoning chains showing a successful unification of reasoning and grounding abilities overview grit teaching mllms to think with images by generating reasoning chains that interleave natural language with bounding boxes key innovations 1 grounded reasoning paradigm models generate reasoning chains interleaving natural language with explicit bounding box coordinates 2 grpo gr a reinforcement learning algorithm which employs novel rewards that enable the grounded reasoning ability of mllms efficiently no manual reasoning annotations needed grounded reasoning paradigm we propose grounded reasoning paradigm as a light and efficient way for pre trained mllms to think with images the model is trained to generate reasoning chains that interleave natural language with bounding boxes think initial natural language reasoning with bounding boxes generated x1 y1 x2 y2 think rethink reasoning continues with further analysis on image regions rethink answer final response the model generates reasoning in one continuous pass with special tokens to facilitate an extensible reasoning process bounding boxes tokens generated directly during autoregressive generation influencing on subsequent reasoning directly as tokens without any external decoding image retrieval process flexibly generated can be placed anywhere within the natural language reasoning with dynamic number of boxes from zero to multiple as needed grpo gr reinforcement learning for grounded reasoning grit employs grpo gr a reinforcement learning approach that optimizes a policy π θ to generate grounded reasoning sequences with three key reward components 1 grounded reasoning format reward r format encourages proper use of reasoning structure related special tokens and valid bounding box syntax 2 grounded target counting reward r count for counting tasks verifies the number of bounding boxes matches the ground truth count 3 gpt aided answer accuracy reward r ans combines gpt 4o evaluation of answer correctness with bleu similarity to ground truth this approach eliminates the need for explicit reasoning annotations or bounding box labels achieving remarkable data efficiency with as few as 20 training examples while maintaining strong performance main results directly using the grit method with grpo gr reinforcement learning we train two pre trained mllms qwen2 5 vl 3b and internvl 3 2b with only 20 image question answer triplets from existing datasets vsr and tallyqa comprehensive evaluations demonstrate that grit trained models overall outperform baselines across six testing sets datasets metrics gpt as judge answer accuracy acc measures answer correctness using gpt 4o as judge 0 1 score grounding iou giou measures alignment between model generated bounding boxes and ground truth regions results confirm that grit effectively unifies previously separated grounding and reasoning capabilities in mllms efficiently and effectively achieving strong performance across diverse visual reasoning tasks examples example 1 counting example 2 spatial reasoning example 3 absence handling model provides an answer with regions and then reflects on it in the subsequent reasoning question how many zebras are pictured here model output think there are six zebras in the picture the coordinates for the zebras are as follows 1 200 168 248 202 2 169 159 214 186 3 76 167 108 192 4 24 173 50 197 5 51 163 70 191 6 413 159 441 189 7 463 171 483 186 think rethink the coordinates provided for the zebras are accurate and cover all the zebras visible in the image there are no overlapping or missing coordinates rethink answer 7 ground truth answer 7 model first grounds critical image region in its reasoning and then analyzes in the subsequent reasoning question is the truck beneath the cat model output think the bounding box coordinates for the truck are approximately 0 209 488 364 the cat is positioned on top of this truck think rethink the cat is on top of the truck so the truck is beneath the cat rethink answer yes ground truth answer yes model correctly handles queries about non existent entities without any grounding action question is there a knife in the image model output think there is no knife present in the image the focus is on a pot of soup with carrots and other ingredients the bounding box coordinates for the knife would be outside the image area think rethink the image shows a pot of soup with carrots and other ingredients but there is no knife present the bounding box coordinates for the knife would be outside the image area rethink answer no ground truth answer no inference examples of qwen2 5 vl grit bibtex misc fan2025gritteachingmllmsthink title grit teaching mllms to think with images author yue fan and xuehai he and diji yang and kaizhi zheng and ching chen kuo and yuting zheng and sravana jyothi narayanaraju and xinze guan and xin eric wang year 2025 eprint 2505 15879 archiveprefix arxiv primaryclass cs cv url https arxiv org abs 2505 15879 this website is adapted from nerfies and gligen licensed under a creative commons attribution sharealike 4 0 international license this website is licensed under a creative commons attribution sharealike 4 0 international license
Thumbnail images (randomly selected): * Images may be subject to copyright.GREEN status (no comments)
  • Example 3

The site also has references to the 2 subdomain(s)

  eric-xw.github.io  Verify   gligen.github.io  Verify


Top 50 hastags from of all verified websites.

Supplementary Information (add-on for SEO geeks)*- See more on header.verify-www.com

Header

HTTP/1.1 301 Moved Permanently
Connection close
Content-Length 162
Server GitHub.com
Content-Type text/html
Location htt????/grounded-reasoning.github.io/
X-GitHub-Request-Id 48FA:123CB3:4EB1C:53BF2:6A6040BB
Accept-Ranges bytes
Age 0
Date Wed, 22 Jul 2026 04:02:06 GMT
Via 1.1 varnish
X-Served-By cache-lcy-egml8630044-LCY
X-Cache MISS
X-Cache-Hits 0
X-Timer S1784692927.843458,VS0,VE83
Vary Accept-Encoding
X-Fastly-Request-ID 56358c09a72f1e49efaaf252e426424371b788f5
HTTP/2 200
server GitHub.com
content-type text/html; charset=utf-8
last-modified Mon, 01 Dec 2025 15:55:08 GMT
access-control-allow-origin *
strict-transport-security max-age=31556952
etag W/ 692dba5c-70b7
expires Wed, 22 Jul 2026 04:12:07 GMT
cache-control max-age=600
content-encoding gzip
x-proxy-cache MISS
x-github-request-id 5A46:292E42:183516:18C00E:6A6040BE
accept-ranges bytes
age 0
date Wed, 22 Jul 2026 04:02:07 GMT
via 1.1 varnish
x-served-by cache-rtm-ehrd2290038-RTM
x-cache MISS
x-cache-hits 0
x-timer S1784692927.952165,VS0,VE131
vary Accept-Encoding
x-fastly-request-id a36ad975918960983100c5228b62606483d91101
content-length 6607

Meta Tags

title="GRIT: Teaching MLLMs to Think with Images"
charset="utf-8"
name="description" content="GRIT: Teaching MLLMs to Think with Images"
name="keywords" content="GRIT, MLLMs, visual reasoning, grounded reasoning"
name="viewport" content="width=device-width, initial-scale=1"

Load Info

page size6607
load time (s)0.283863
redirect count1
speed download23346
server IP 185.199.111.153
* all occurrences of the string "http://" have been changed to "htt???/"