Запись архива

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
Journalists Protest against rising violence during march in Mexi | by Knight Foundation | openverse | by-sa

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

Что произошло

Источник arXiv cs.CL зафиксировал сигнал: arXiv:2607.28640v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. Specifically, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap, a method that constructs such inputs by replacing textual concepts with semantically aligned images, resulting in sequences where visual tokens are interleaved with text tokens. Based on TokenSwap, we transform existing text-based benchmarks such as MMLU into image-interleaved counterparts, resulting in TokenSwap-Bench. Across 42 MLLMs, we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% +/- 3.3% across models. Notably, we observe that reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models. In contrast, neither prompting strategies nor scaling training compute alone reliably reduces the modality gap. Finally, we demonstrate that incorporating TokenSwap during training effectively mitigates this gap while preserving strong text-only and vision-language performance.

Почему это обсуждают

Для аудитории COMRAD404 это повод проверить, касается ли тема моделей, агентов, промптов, инструментов или разработки с ИИ. Социальный источник сам по себе не является доказательством, поэтому выводы нужно держать осторожными.

Что подтверждено

Punkt Detail
Платформа arxiv
Источник arXiv cs.CL
Проверка https://arxiv.org/

Что проверить дальше

Нужно открыть первичный источник, документацию продукта, GitHub, блог лаборатории или публикацию автора и отделить факт релиза от реакции сообщества.

Источник: arXiv cs.CL – https://arxiv.org/abs/2607.28640; проверка: https://arxiv.org/