If you are not sure if the website you would like to visit is secure, you can verify it here. Enter the website address of the page and see parts of its content and the thumbnail images on this site. None (if any) dangerous scripts on the referenced page will be executed. Additionally, if the selected site contains subpages, you can verify it (review) in batches containing 5 pages.
favicon.ico: vggsounder.github.io - VGGSounder: Audio-Visual Evalu.

site address: vggsounder.github.io redirected to: vggsounder.github.io

site title: VGGSounder: Audio-Visual Evaluations for Foundation Models

Our opinion (on Wednesday 22 July 2026 19:05:30 UTC):

GREEN status (no comments) - no comments
After content analysis of this website we propose the following hashtags:



Meta tags:
description=VGGSounder, a multi-label audio-visual classification dataset with modality annotations.;
keywords=VGGSounder, Audio-Visual Benchmark, Foundation Models;

Headings (most frequently used words):

models, vggsounder, modality, audio, visual, evaluations, for, foundation, does, your, model, actually, listen, or, just, look, vs, vggsound, dataset, preview, adding, can, make, worse, confusion, across, video, abstract, bibtex, what, we, found, many, are, vision, centric, some, overfit, to, speech, see, better, when, things, move,

Text of the page (most frequently used words):
the (20), audible (19), and (17), #modality (16), models (14), vggsound (14), video (13), visible (13), people (11), only (11), your (10), playing (10), audio (9), over (9), meta (9), does (9), visual (8), label (8), speech (8), browser (8), not (8), support (8), tag (8), for (7), can (7), original (7), dog (7), what (7), foundation (6), model (6), water (6), voice (6), speaking (6), every (5), this (5), both (5), evaluations (4), test (4), confusion (4), see (4), many (4), they (4), male (4), man (4), vggsounder (3), vision (3), dataset (3), our (3), limitations (3), that (3), annotations (3), when (3), are (3), per (3), sees (3), splashing (3), barking (3), background (3), music (3), sniggering (3), clip (3), you (3), labels (3), occurring (3), tübingen (3), daniil (2), zverev (2), thaddäus (2), wiedemer (2), ameya (2), prabhu (2), matthias (2), bethge (2), wieland (2), brendel (2), sophia (2), koepke (2), several (2), classes (2), misaligned (2), these (2), multi (2), set (2), performance (2), adding (2), input (2), answer (2), then (2), once (2), worse (2), motorboat (2), speedboat (2), acceleration (2), screaming (2), bow (2), wow (2), drum (2), kit (2), sloshing (2), crowd (2), cheering (2), belly (2), laughing (2), sea (2), guitar (2), explore (2), samples (2), task (2), whether (2), check (2), ground (2), truth (2), try (2), image (2), class (2), one (2), lean (2), just (2), heard (2), seen (2), hears (2), university (2), inproceedings, zverevwiedemer2025vggsounder, author, title, booktitle, proceedings, ieee, cvf, international, conference, computer, iccv, year, 2025, bibtex, emergence, underscores, importance, reliably, assessing, their, multimodal, understanding, commonly, used, benchmark, evaluating, classification, however, analysis, identifies, including, incomplete, labelling, partially, overlapping, modalities, lead, distorted, auditory, capabilities, address, introduce, comprehensively, annotated, extends, specifically, designed, evaluate, features, detailed, enabling, precise, analyses, specific, furthermore, reveal, analysing, degradation, another, with, new, metric, abstract, better, things, move, some, overfit, balance, centric, let, profile, even, surface, biases, data, were, trained, across, correctly, from, alone, also, call, toggle, below, happen, lose, make, engine, accelerating, revving, vroom, fire, crackling, wind, noise, owl, hooting, female, woman, cricket, chirping, whimpering, growling, baying, violin, fiddle, trumpet, trombone, flute, orchestra, swimming, children, shouting, child, kid, lacrosse, hockey, air, horn, lion, goat, bleating, footsteps, snow, fireworks, banging, spraying, waves, electric, bass, scroll, more, preview, mechanical, turk, annotators, faced, watch, decide, each, proposed, neither, answers, against, annotation, guess, controlled, confounder, free, evaluation, profiling, makes, measurable, split, annotates, synonym, superclass, merging, static, aligned, subsets, filter, has, train, splits, counted, errors, confounders, information, single, same, videos, richer, tagged, evidence, spoken, narration, shortcut, stumble, gone, bias, noticeably, clips, still, than, real, motion, stills, forget, hear, added, found, exactly, quantify, which, prefers, good, should, understand, equally, well, often, doesn, quietly, favour, sense, other, standard, benchmarks, like, catch, because, never, record, actually, listen, look, code, arxiv, paper, equal, contributors, mpi, intelligent, systems, ellis, institute, center, technical, munich, mcml,


Text of the page (random words):
vggsounder audio visual evaluations for foundation models vggsounder audio visual evaluations for foundation models daniil zverev 1 thaddäus wiedemer 2 3 4 ameya prabhu 2 3 matthias bethge 2 3 wieland brendel 3 4 a sophia koepke 1 2 3 1 technical university of munich mcml 2 university of tübingen 3 tübingen ai center 4 mpi for intelligent systems ellis institute tübingen equal contributors paper arxiv video code does your model actually listen or just look a good audio visual model should understand what it hears and what it sees equally well often it doesn t many models quietly favour one sense over the other and standard benchmarks like vggsound 1 can t catch this because they never record whether a label can be heard or seen vggsound er labels every class as heard seen or both so you can check exactly what a model sees and hears and quantify which modality it prefers what we found ️ vision over audio many foundation models lean on what they see and forget what they hear once video is added ️ motion over stills many models do noticeably worse on clips that are just a still image than on real video ️ speech bias several models lean on spoken narration as a shortcut and stumble when it s gone explore dataset see the evidence vggsound er vs vggsound same videos richer ground truth every co occurring label tagged as audible visible or both vggsound one single label per clip no modality information 48 of samples are modality misaligned no meta labels for confounders co occurring classes counted as errors has both train and test splits vggsound er multi label every co occurring class per label audible visible both modality aligned subsets you can filter to meta labels background music voice over static image synonym superclass merging test split only re annotates vggsound s test set what this makes measurable modality profiling modality confusion controlled confounder free evaluation can you guess every label in this clip try the annotation task try this is the task our mechanical turk annotators faced watch the clip and decide for each proposed label whether it is audible visible both or neither then check your answers against vggsound er s ground truth modality annotations dataset preview scroll to explore more samples your browser does not support the video tag audible only male speech man speaking playing bass guitar playing drum kit playing electric guitar background music meta voice over meta visible only sea waves audible visible sloshing water splashing water spraying water motorboat speedboat acceleration original your browser does not support the video tag audible only male speech man speaking people sniggering audible visible fireworks banging original footsteps on snow your browser does not support the video tag audible only people belly laughing people sniggering visible only splashing water audible visible dog barking original dog bow wow goat bleating sea lion barking your browser does not support the video tag audible only air horn male speech man speaking voice over meta audible visible people cheering people crowd playing hockey playing lacrosse your browser does not support the video tag audible only people belly laughing voice over meta audible visible child speech kid speaking children shouting male speech man speaking people cheering people crowd people screaming sloshing water original swimming people sniggering your browser does not support the video tag audible only orchestra playing drum kit playing flute playing trombone playing trumpet playing violin fiddle background music meta audible visible dog barking dog bow wow original dog baying dog growling dog whimpering your browser does not support the video tag audible only cricket chirping female speech woman speaking owl hooting original wind noise voice over meta audible visible fire crackling your browser does not support the video tag audible only people screaming voice over meta audible visible engine accelerating revving vroom motorboat speedboat acceleration original splashing water adding a modality can make models worse a model can answer correctly from audio alone then lose that answer once it also sees the video we call this modality confusion toggle the input below to see it happen modality confusion across models vggsound er s per modality annotations let us profile every model and even surface biases in the data they were trained on many models are vision centric modality balance some models overfit to speech models see better when things move video abstract the emergence of audio visual foundation models underscores the importance of reliably assessing their multimodal understanding the vggsound dataset is commonly used as a benchmark for evaluating audio visual classification however our analysis identifies several limitations of vggsound including incomplete labelling partially overlapping classes and misaligned modalities these lead to distorted evaluations of auditory and visual capabilities to address these limitations we introduce vggsound er a comprehensively re annotated multi label test set that extends vggsound and is specifically designed to evaluate audio visual foundation models vggsound er features detailed modality annotations enabling precise analyses of modality specific performance furthermore we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric bibtex inproceedings zverevwiedemer2025vggsounder author daniil zverev and thaddäus wiedemer and ameya prabhu and matthias bethge and wieland brendel and a sophia koepke title vggsounder audio visual evaluations for foundation models booktitle proceedings of the ieee cvf international conference on computer vision iccv year 2025
Thumbnail images (randomly selected): * Images may be subject to copyright.GREEN status (no comments)

    No Images


    Verified site has: 1 subpage(s). Do you want to verify them? Verify pages:

    1-1


    The site also has references to the 2 subdomain(s)

      drimpossible.github.io  Verify   akoepke.github.io  Verify


    The site also has 1 references to other resources (not html/xhtml )

     vggsounder.github.io/./static/vggsounder.pdf  Verify


    Top 50 hastags from of all verified websites.

    Supplementary Information (add-on for SEO geeks)*- See more on header.verify-www.com

    Header

    HTTP/1.1 301 Moved Permanently
    Connection close
    Content-Length 162
    Server GitHub.com
    Content-Type text/html
    Location htt????/vggsounder.github.io/
    X-GitHub-Request-Id 67DC:3E201F:4C225:5013B:6A611479
    Accept-Ranges bytes
    Age 0
    Date Wed, 22 Jul 2026 19:05:29 GMT
    Via 1.1 varnish
    X-Served-By cache-lcy-egml8630051-LCY
    X-Cache MISS
    X-Cache-Hits 0
    X-Timer S1784747130.533218,VS0,VE88
    Vary Accept-Encoding
    X-Fastly-Request-ID 039f291f99b481939f25aacab8781e08c665635d
    HTTP/2 200
    server GitHub.com
    content-type text/html; charset=utf-8
    last-modified Tue, 07 Jul 2026 18:01:24 GMT
    access-control-allow-origin *
    strict-transport-security max-age=31556952
    etag W/ 6a4d3ef4-7d35
    expires Wed, 22 Jul 2026 19:15:29 GMT
    cache-control max-age=600
    content-encoding gzip
    x-proxy-cache MISS
    x-github-request-id 2A38:1110AF:2607A:26752:6A611479
    accept-ranges bytes
    age 0
    date Wed, 22 Jul 2026 19:05:29 GMT
    via 1.1 varnish
    x-served-by cache-rtm-ehrd2290034-RTM
    x-cache MISS
    x-cache-hits 0
    x-timer S1784747130.647484,VS0,VE108
    vary Accept-Encoding
    x-fastly-request-id 8b2256f11e6f52a11c74b7c08672f2e7672c595d
    content-length 6403

    Meta Tags

    title="VGGSounder: Audio-Visual Evaluations for Foundation Models"
    charset="utf-8"
    name="description" content="VGGSounder, a multi-label audio-visual classification dataset with modality annotations."
    name="keywords" content="VGGSounder, Audio-Visual Benchmark, Foundation Models"
    name="viewport" content="width=device-width, initial-scale=1"

    Load Info

    page size6403
    load time (s)0.268333
    redirect count1
    speed download23891
    server IP 185.199.110.153
    * all occurrences of the string "http://" have been changed to "htt???/"