
0.35% trained, 100% competitive: the frozen-tower architecture behind jina-embeddings-v5-omni
The latest jina embeddings model generates multimodal embeddings for text, images, video and audio, competing with models nearly 6x its size on vector search while training just 0.35% of the weights.