If you are not sure if the website you would like to visit is secure, you can verify it here. Enter the website address of the page and see parts of its content and the thumbnail images on this site. None (if any) dangerous scripts on the referenced page will be executed. Additionally, if the selected site contains subpages, you can verify it (review) in batches containing 5 pages.
favicon.ico: scanents3d.github.io - ScanEnts3D: Exploiting Phrase-.

site address: scanents3d.github.io redirected to: scanents3d.github.io

site title: ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes

Our opinion (on Wednesday 22 July 2026 14:42:36 UTC):

GREEN status (no comments) - no comments
After content analysis of this website we propose the following hashtags:



Meta tags:
description=ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes;
keywords=ScanEnts3D, Dataset, 3D Visual Grounding, 3D Dense Captioning, Referit3D, ScanRefer, ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes;

Headings (most frequently used words):

3d, in, dataset, scanents3d, grounded, language, exploiting, phrase, to, object, correspondences, for, improved, visio, linguistic, models, scenes, abstract, motivation, scan, entities, method, comprehension, production, qualitative, results, download, browser, citation,

Text of the page (most frequently used words):
the (42), and (29), scanents3d (11), #object (11), #dataset (11), objects (9), for (8), model (8), our (8), can (7), language (7), architectures (7), all (7), correspondences (6), with (6), their (6), neural (6), that (6), scenes (5), task (5), modifications (5), losses (5), target (5), listening (5), this (4), visio (4), linguistic (4), grounded (4), propose (4), existing (4), show (4), proposed (4), scene (4), are (4), annotations (4), exploiting (3), phrase (3), improved (3), models (3), here (3), two (3), trans2cap (3), learning (3), loss (3), mentioned (3), between (3), both (3), nr3d (3), scanrefer (3), referential (3), scan (3), website (2), which (2), abdelreheem (2), ahmed (2), olszewski (2), kyle (2), lee (2), hsin (2), ying (2), wonka (2), peter (2), achlioptas (2), panos (2), research (2), you (2), download (2), production (2), above (2), figure (2), scanents (2), caption (2), green (2), transfer (2), use (2), trained (2), same (2), time (2), modular (2), predict (2), instances (2), comprehension (2), three (2), functions (2), flexible (2), listeners (2), mvt (2), centric (2), crucially (2), anchor (2), several (2), additional (2), training (2), speaking (2), grounding (2), beyond (2), provide (2), large (2), scale (2), extending (2), underlying (2), entities (2), more (2), from (2), they (2), its (2), them (2), incorporating (2), natural (2), real (2), world (2), including (2), improving (2), sota (2), based, licensed, under, creative, commons, attribution, sharealike, international, license, nerfies, template, article, abdelreheem2022scanents, author, title, journal, computing, repository, corr, volume, abs, 2212, 06250, year, 2022, citation, browse, few, sampled, examples, browser, qualitative, results, corresponding, appropriate, attend, tell, m2cap, adapting, operate, given, set, outputs, table, box, exploits, cross, modal, knowledge, inputs, together, counterpart, images, adopts, student, teacher, paradigm, boxes, yellow, approach, finetuning, pre, encoder, promote, discriminative, feature, representations, guides, network, ground, truth, new, generic, serve, auxiliary, add, ons, demonstrates, adjusted, applied, independently, top, context, aware, features, extended, shown, purple, class, distractor, red, default, only, predicts, state, art, utilize, provided, during, explore, tasks, multiple, per, main, goal, demonstrate, inherent, value, curated, simple, implement, lead, substantial, improvements, therefore, conjecture, similar, will, possible, extant, future, making, method, share, community, each, explicitly, any, mentions, introduce, utterances, includes, 369, 039, than, times, number, original, works, itie, ent, when, humans, describe, typically, enumerating, ego, properties, texture, geometry, instead, refer, direct, relations, other, dubbed, anchors, work, investigate, rigorous, exploitation, such, annotating, modern, via, motivation, popular, datasets, connect, data, paper, curate, complementary, aforementioned, ones, associating, sentence, inside, specifically, provides, explicit, 369k, across, 84k, sentences, covering, 705, intuitive, enable, novel, significantly, improve, performance, recently, introduced, benchmarks, respectively, moreover, experiment, competitive, baselines, recent, methods, generation, speakers, also, noticeably, benefit, cider, points, benchmark, overall, carefully, conducted, experimental, studies, strongly, support, conclusion, commonly, used, become, efficient, interpretable, generalization, without, needing, these, newly, collected, test, referit3d, abstract, released, arxiv, wacv, 2024, snap, inc, kaust,


Text of the page (random words):
scanents3d exploiting phrase to 3d object correspondences for improved visio linguistic models in 3d scenes scanents3d exploiting phrase to 3d object correspondences for improved visio linguistic models in 3d scenes ahmed abdelreheem 1 2 kyle olszewski 2 hsin ying lee 2 peter wonka 1 panos achlioptas 2 1 kaust 2 snap inc wacv 2024 arxiv dataset released abstract the two popular datasets scanrefer and referit3d connect natural language to real world 3d data in this paper we curate a large scale and complementary dataset extending both the aforementioned ones by associating all objects mentioned in a referential sentence to their underlying instances inside a 3d scene specifically our scan entities in 3d scanents3d dataset provides explicit correspondences between 369k objects across 84k natural referential sentences covering 705 real world scenes crucially we show that by incorporating intuitive losses that enable learning from this novel dataset we can significantly improve the performance of several recently introduced neural listening architectures including improving the sota in both the nr3d and scanrefer benchmarks by 4 3 and 5 0 respectively moreover we experiment with competitive baselines and recent methods for the task of language generation and show that as with neural listeners 3d neural speakers can also noticeably benefit by training with scanents3d including improving the sota by 13 2 cider points on the nr3d benchmark overall our carefully conducted experimental studies strongly support the conclusion that by learning on scanents3d commonly used visio linguistic 3d architectures can become more efficient and interpretable in their generalization without needing to provide these newly collected annotations at test time motivation when humans describe an object in a 3d scene they typically go beyond enumerating its ego centric properties e g its texture or geometry instead they refer to direct relations between the target and other co existing objects in the scene dubbed as anchors in this work we investigate the rigorous exploitation of such anchor objects by annotating them and incorporating them in modern neural listening and speaking architectures via modular and flexible loss functions scanents3d scan ent itie s in 3d dataset we share with the research community grounding annotations that go beyond each target object and explicitly provide all the correspondences between all 3d objects and any of their mentions we introduce a large scale dataset extending both nr3d and scanrefer by grounding all objects mentioned in their referential utterances to their underlying 3d scenes our scanents3d dataset scan entities in 3d includes an additional 369 039 language to object correspondences more than three times the number from the original works method we propose modifications to several existing state of the art architectures to utilize the additional annotations provided by scanents3d during training we explore two tasks neural listening and speaking and multiple architectures per task our main goal is to demonstrate the inherent value of the curated annotations all proposed modifications are simple to implement and lead to substantial improvements we therefore conjecture that similar modifications are or will be possible to extant and future architectures making use of scanents3d 3d grounded language comprehension for the 3d grounded language comprehension task we propose three new loss functions which are flexible generic and can serve as auxiliary add ons to existing neural listeners the above figure demonstrates our proposed listening losses adjusted for the mvt model the proposed losses are applied independently on top of object centric and context aware features crucially the extended mvt scanents model can predict all anchor objects shown in purple same class distractor objects red and the target green the default model only predicts the target grounded language production in 3d for the grounded language production in 3d task we propose corresponding modifications and appropriate losses to two existing architectures show attend tell model and x trans2cap in the above figure we propose the m2cap scanents model adapting x trans2cap model to operate with our proposed losses the model is given a set of 3d objects in a 3d scene and outputs a caption for the target object e g the table in the green box the x trans2cap model exploits cross modal knowledge transfer 3d inputs together with their counterpart 2d images and adopts a student teacher paradigm boxes in yellow show our modifications here we use a transfer learning approach by finetuning a pre trained object encoder trained on the listening task to promote discriminative object feature representations at the same time our modular loss guides the network to predict all object instances mentioned in the ground truth caption qualitative results dataset download you can download the dataset here dataset browser you can browse a few sampled examples of scanents3d dataset here citation article abdelreheem2022scanents author abdelreheem ahmed and olszewski kyle and lee hsin ying and wonka peter and achlioptas panos title scanents3d exploiting phrase to 3d object correspondences for improved visio linguistic models in 3d scenes journal computing research repository corr volume abs 2212 06250 year 2022 this website is based on the nerfies website template which is licensed under a creative commons attribution sharealike 4 0 international license
Thumbnail images (randomly selected): * Images may be subject to copyright.GREEN status (no comments)

Verified site has: 1 subpage(s). Do you want to verify them? Verify pages:

1-1


Top 50 hastags from of all verified websites.

Supplementary Information (add-on for SEO geeks)*- See more on header.verify-www.com

Header

HTTP/1.1 301 Moved Permanently
Connection close
Content-Length 162
Server GitHub.com
Content-Type text/html
Location htt????/scanents3d.github.io/
X-GitHub-Request-Id C5F6:154439:1E0B05:1F8F49:6A60CC61
Accept-Ranges bytes
Age 2682
Date Wed, 22 Jul 2026 14:42:35 GMT
Via 1.1 varnish
X-Served-By cache-lcy-egml8630036-LCY
X-Cache HIT
X-Cache-Hits 0
X-Timer S1784731355.343123,VS0,VE1
Vary Accept-Encoding
X-Fastly-Request-ID cebfd4b43112753be52b62c8755e871df7e79a11
HTTP/2 200
server GitHub.com
content-type text/html; charset=utf-8
last-modified Wed, 21 Aug 2024 21:26:02 GMT
access-control-allow-origin *
strict-transport-security max-age=31556952
etag W/ 66c65b6a-44e4
expires Wed, 22 Jul 2026 14:52:35 GMT
cache-control max-age=600
content-encoding gzip
x-proxy-cache MISS
x-github-request-id AE3A:5070D:B06811:B31572:6A60D6DB
accept-ranges bytes
age 0
date Wed, 22 Jul 2026 14:42:35 GMT
via 1.1 varnish
x-served-by cache-rtm-ehrd2290052-RTM
x-cache MISS
x-cache-hits 0
x-timer S1784731355.372123,VS0,VE114
vary Accept-Encoding
x-fastly-request-id 70a5316f5c788d55293d1e7b2774fadcc92c586f
content-length 5285

Meta Tags

title="ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes"
charset="utf-8"
name="description" content="ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes"
name="keywords" content="ScanEnts3D, Dataset, 3D Visual Grounding, 3D Dense Captioning, Referit3D, ScanRefer, ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D Scenes"
name="viewport" content="width=device-width, initial-scale=1"

Load Info

page size5285
load time (s)0.180291
redirect count1
speed download29361
server IP 185.199.109.153
* all occurrences of the string "http://" have been changed to "htt???/"