Post
975
Two rival models, one memory.
Four multimodal models from four different companies (two at a time). One of them indexes a thousand photographs. Then a different company's model searches that index, having never been trained to work with it.
The same pictures come back.
Cross-company retrieval is statistically indistinguishable from a model searching its own index: retention 0.988, 95% CI [0.955, 1.023], across 12 cross directions and 1,000 held-out items. The interval contains 1.0, so we report indistinguishability and keep the interval attached to the number. Four of the twelve cross directions score above the within-vendor baseline of the gallery they search.
A single linear tower also places image, audio and video beside text in one searchable space, beating one tower per modality on every modality. Video gains most, 0.076.
Two strips of results for the same request, one per company:
RiverRider/srt-omni-demo
The four hosts:
Qwen/Qwen3-Omni-30B-A3B-Instruct
google/gemma-4-31B-it
mistralai/Mistral-Small-3.1-24B-Instruct-2503
rhymes-ai/Aria
States, fitted towers and the manifest are published, so you can add a host we have not tried:
RiverRider/srt-omni-xvendor-towers
Scope. The four-vendor result covers images, since two of the hosts handle images only. Raw cosine between unrelated items on these states is +0.869 and raw retrieval sits at chance, so centering is not optional. All four hosts train on overlapping web-scale corpora, so this shows the readable structure survives a change of vendor, architecture and training run. The stronger reading, that independently trained systems would converge on that structure, remains untested and we do not assert it.
Four multimodal models from four different companies (two at a time). One of them indexes a thousand photographs. Then a different company's model searches that index, having never been trained to work with it.
The same pictures come back.
Cross-company retrieval is statistically indistinguishable from a model searching its own index: retention 0.988, 95% CI [0.955, 1.023], across 12 cross directions and 1,000 held-out items. The interval contains 1.0, so we report indistinguishability and keep the interval attached to the number. Four of the twelve cross directions score above the within-vendor baseline of the gallery they search.
A single linear tower also places image, audio and video beside text in one searchable space, beating one tower per modality on every modality. Video gains most, 0.076.
Two strips of results for the same request, one per company:
RiverRider/srt-omni-demo
The four hosts:
Qwen/Qwen3-Omni-30B-A3B-Instruct
google/gemma-4-31B-it
mistralai/Mistral-Small-3.1-24B-Instruct-2503
rhymes-ai/Aria
States, fitted towers and the manifest are published, so you can add a host we have not tried:
RiverRider/srt-omni-xvendor-towers
Scope. The four-vendor result covers images, since two of the hosts handle images only. Raw cosine between unrelated items on these states is +0.869 and raw retrieval sits at chance, so centering is not optional. All four hosts train on overlapping web-scale corpora, so this shows the readable structure survives a change of vendor, architecture and training run. The stronger reading, that independently trained systems would converge on that structure, remains untested and we do not assert it.