Meta tags:
description= VLM4D: Towards Spatiotemporal Awareness in Vision Language Models;
keywords= VLM, spatiotemporal reasoning;
Headings (most frequently used words):
vlm4d, towards, spatiotemporal, awareness, in, vision, language, models, abstract, distribution, of, dataset, sources, and, annotations, interactive, demo, model, performance, leaderboard, bibtex, which, direction, is, he, running, toward,
Text of the page (most frequently used words):
and (29), the (18), vlms (14), spatiotemporal (13), 2024 (13), 2025 (12), #vision (10), models (8), benchmark (7), reasoning (7), vlm4d (6), language (6), centric (6), real (6), awareness (5), across (5), this (4), from (4), towards (4), video (4), qwen2 (4), open (4), source (4), synthetic (4), our (4), videos (4), visual (4), dynamic (4), international (3), llava (3), gpt (3), performance (3), exo (3), ego (3), question (3), dataset (3), first (3), capabilities (3), understanding (3), humans (3), perspective (3), website (2), licensed (2), under (2), creative (2), commons (2), attribution (2), sharealike (2), license (2), zhou (2), shijie (2), vilesov (2), alexander (2), xuehai (2), wan (2), ziyu (2), zhang (2), shuwang (2), nagachandra (2), aditya (2), chang (2), chen (2), dongdong (2), wang (2), eric (2), xin (2), kadambi (2), achuta (2), next (2), internvideo2 (2), shanghai (2), lab (2), videollama3 (2), 72b (2), internvl2 (2), deepseek (2), phi (2), microsoft (2), llama (2), 17b (2), gemini (2), pro (2), proprietary (2), random (2), human (2), average (2), model (2), cot (2), leaderboard (2), example (2), leading (2), comparison (2), accuracy (2), types (2), questions (2), right (2), your (2), with (2), person (2), further (2), translational (2), rotational (2), integrating (2), effortlessly (2), reason (2), for (2), world (2), current (2), paper (2), designed (2), evaluate (2), motion (2), struggle (2), temporal (2), deeper (2), spatial (2), time (2), moving (2), adapted, gligen, nerfies, inproceedings, zhou2025vlm4d, title, author, booktitle, proceedings, ieee, cvf, conference, computer, pages, 8600, 8612, year, bibtex, 34b, one, damo, alibaba, aria, rhymes, pixtral, 12b, mistral, 38b, vl2, multimodal, scout, maverick, meta, image, grok, xai, claude, sonnet, anthropic, google, openai, latest, selection, 100, user, study, directional, overall, release, organization, report, evaluation, using, various, responses, two, statistics, entire, between, chain, thought, direct, output, prompting, left, score, which, direction, running, toward, browser, does, not, support, tag, previous, test, skills, interactive, demo, overview, composition, illustrating, proportions, third, data, categorized, annotation, including, action, counting, false, positive, queries, targeting, nonexistent, events, assess, critical, cosmos, ego4d, youtube, vos, davis, distribution, sources, annotations, have, shown, remarkable, linguistic, but, remain, fundamentally, limited, interactions, track, about, object, movements, rotations, shifts, abilities, essential, robust, yet, notably, lacking, introduce, specifically, comprises, diverse, accompanied, carefully, curated, answer, pairs, emphasizing, motions, continuity, through, comprehensive, evaluations, state, art, closed, identify, significant, gaps, compared, baselines, highlighting, fundamental, deficiencies, existing, extensive, analysis, reveals, that, particularly, multiple, cues, maintaining, coherence, explore, promising, directions, such, leveraging, feature, field, reconstruction, targeted, supervised, fine, tuning, demonstrating, their, effectiveness, enhancing, comprehension, work, aims, encourage, exploration, into, improving, grounding, paving, way, more, capable, reliable, intelligence, environments, abstract, intuitively, space, reconstructing, trajectory, objects, any, contrast, typically, rely, aggregating, features, incorrect, predictions, when, interpretation, requires, correctly, perceive, car, while, vlm, inaccurately, predicts, leftward, movement, suggesting, perform, explicitly, code, iccv, honolulu, equal, contribution, usc, ucsc, ucla,
Text of the page (random words):
vlm4d towards spatiotemporal awareness in vision language models vlm4d towards spatiotemporal awareness in vision language models shijie zhou 1 alexander vilesov 1 xuehai he 2 3 ziyu wan 2 shuwang zhang 1 aditya nagachandra 1 di chang 4 dongdong chen 2 xin eric wang 3 achuta kadambi 1 1 ucla 2 microsoft 3 ucsc 4 usc equal contribution iccv 2025 honolulu paper dataset leaderboard code tl dr the first benchmark explicitly designed to evaluate the spatiotemporal 4d reasoning capabilities of vision language models vlms spatiotemporal 4d awareness humans intuitively reason in 4d 3d space time effortlessly reconstructing the dynamic spatial trajectory of moving objects from any perspective in contrast current vision language models vlms typically rely on aggregating 2d visual features across time leading to incorrect predictions when motion understanding and interpretation requires deeper spatiotemporal reasoning in this example humans correctly perceive the car moving to the right while the vlm gpt 4o inaccurately predicts leftward movement suggesting vlms struggle to perform spatiotemporal reasoning abstract vision language models vlms have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions humans effortlessly track and reason about object movements rotations and perspective shifts abilities essential for robust dynamic real world understanding yet notably lacking in current vlms in this paper we introduce vlm4d the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of vlms our benchmark comprises diverse real world and synthetic videos accompanied by carefully curated question answer pairs emphasizing translational and rotational motions perspective awareness and motion continuity through comprehensive evaluations of state of the art open and closed source vlms we identify significant performance gaps compared to human baselines highlighting fundamental deficiencies in existing models extensive analysis reveals that vlms struggle particularly with integrating multiple visual cues and maintaining temporal coherence we further explore promising directions such as leveraging 4d feature field reconstruction and targeted spatiotemporal supervised fine tuning demonstrating their effectiveness in enhancing spatiotemporal comprehension our work aims to encourage deeper exploration into improving vlms spatial and temporal grounding paving the way towards more capable and reliable visual intelligence for dynamic environments distribution of dataset sources and annotations overview of the dataset composition illustrating the proportions of real third person exo centric videos davis youtube vos real first person ego centric videos ego4d and synthetic videos cosmos the real video data is further categorized by annotation types including translational rotational action counting and false positive queries targeting nonexistent events to assess critical reasoning interactive demo test your spatiotemporal reasoning skills with our vlm4d benchmark questions exo centric ego centric synthetic question 1 of 3 previous next your browser does not support the video tag which direction is he running toward score 0 0 0 model performance an example question from our benchmark and the responses from two leading models gpt 4o and gemini 2 5 pro statistics across the entire benchmark left comparison of accuracy across types of spatiotemporal questions right accuracy comparison between chain of thought cot and direct output do prompting across vlms leaderboard we report the evaluation using cot on vlm4d benchmark across various proprietary and open source vlms organization model release real synthetic overall ego centric exo centric average directional fp average user study human performance 99 6 99 7 99 7 95 8 100 96 2 98 8 random random selection 24 4 23 2 23 6 25 5 24 7 25 4 24 1 latest proprietary vlms openai gpt 4o 2024 11 55 5 62 2 60 0 49 5 53 3 49 9 57 5 google gemini 2 5 pro 2025 6 64 6 62 9 63 5 54 8 80 0 57 3 62 0 anthropic claude sonnet 4 2025 5 52 6 52 1 52 2 44 0 86 7 48 3 51 3 xai grok 2 vision 2024 12 48 8 49 7 49 4 49 3 66 7 51 0 49 8 open source image vlms meta llama 4 maverick 17b 2025 4 52 6 54 3 53 8 53 3 51 1 53 0 53 6 llama 4 scout 17b 2025 4 48 6 56 2 53 7 53 3 75 6 55 5 54 1 microsoft phi 4 multimodal 2025 3 41 0 35 4 37 2 37 5 11 1 34 8 36 6 phi 3 5 vision 2024 7 33 4 38 8 37 1 23 3 37 8 24 7 34 0 deepseek deepseek vl2 2024 12 33 6 32 9 33 1 31 8 46 7 33 3 33 2 shanghai ai lab internvl2 5 38b 2024 11 46 6 50 1 48 9 43 3 57 8 44 7 47 9 internvl2 5 8b 2024 11 39 0 44 0 42 4 40 8 42 2 40 9 42 0 mistral ai pixtral 12b 2024 9 32 3 25 8 27 9 24 3 22 2 24 0 27 0 rhymes aria 2024 11 47 2 44 0 45 1 38 5 71 1 41 8 44 3 open source video vlms alibaba qwen2 5 vl 7b 2025 1 42 3 43 7 43 3 43 5 64 4 45 6 43 8 qwen2 5 vl 72b 2025 1 54 3 52 5 53 1 49 5 80 0 52 6 53 0 qwen2 vl 7b 2024 8 36 1 34 7 35 2 40 5 35 6 40 0 36 3 qwen2 vl 72b 2024 9 48 1 43 0 44 6 40 8 73 3 44 0 44 5 damo videollama3 2b 2025 1 53 2 42 5 46 0 34 3 55 6 36 4 43 7 videollama3 7b 2025 1 49 4 45 1 46 5 42 8 53 3 43 8 45 9 shanghai ai lab internvideo2 5 8b 2025 1 57 2 50 5 52 7 44 3 46 7 44 5 50 7 internvideo2 8b 2024 8 35 6 39 3 38 1 43 0 0 0 38 7 38 2 llava llava one vision 7b 2024 9 36 8 35 6 36 0 37 8 35 6 37 5 36 3 llava next video 34b 2024 6 29 6 31 6 30 9 24 5 55 6 27 6 30 1 bibtex inproceedings zhou2025vlm4d title vlm4d towards spatiotemporal awareness in vision language models author zhou shijie and vilesov alexander and he xuehai and wan ziyu and zhang shuwang and nagachandra aditya and chang di and chen dongdong and wang eric xin and kadambi achuta booktitle proceedings of the ieee cvf international conference on computer vision pages 8600 8612 year 2025 this website is adapted from nerfies and gligen licensed under a creative commons attribution sharealike 4 0 international license this website is licensed under a creative commons attribution sharealike 4 0 international license
|