Repository navigation
EmbeddingGemma 2 images: ViT on Core ML, 79 images/s - #32
Open
Alex-Wengg wants to merge 4 commits into
Open
Alex-Wengg wants to merge 4 commits into
Alex-Wengg wants to merge 4 commits into
Conversation
…ext space EmbeddingGemma2Vision runs the model's vision encoder (Gemma 4 ViT, 16 layers, 2-D axial RoPE, 3x3 pooling, embed_vision projection) from EmbeddingGemma2Vision.mlpackage on FluidInference/embeddinggemma-2-coreml (revision f9d567d5, its own pinned bundle). One fixed-shape function per soft-token budget, shared weights: vision_70 / vision_140 / vision_280 (630 / 1,260 / 2,520 patches). The pooled tokens go through the text model as <bos> <|image> tokens <image|> <eos>, so images share the text space. - Host side matches HF Gemma4ImageProcessor: aspect-preserving resize to multiples of 48 px within the patch budget (target sizes unit-tested against get_aspect_ratio_preserving_size), patches row by row [16][16][RGB], (x, y) positions, 3x3 pooling matrix, and the patch embedder's position tables (position_embeddings.f16, rows 0-1023) added on the host: the in-graph gather put the model's first ops on the CPU. - GPU (default): 70 tokens 14.6 ms, 140 tokens 34 ms per image. Every op also runs on the Neural Engine, but its time grows with the square of the patch count (37 / 102 / 341 ms), so the GPU wins at every budget. - Accuracy: fp32 wrapper == sentence-transformers (cos 1.000000); Swift (CoreGraphics resize + fp16) vs PyTorch fp32 cos mean 0.998, min 0.991. Zero-shot Oxford Pets (370 photos, 37 breeds): 87.6% in Swift at 70 tokens (PyTorch 87.0%; 140 tokens 88.4%, 280 tokens 89.5%). - Swift end to end at 70 tokens: 79 images/s (vision on the GPU, text on the Neural Engine, four images in flight). ImageSearchCheck scores zero-shot Pets (photos cached by ImageSortDemo) and timing, with --budget and --dump.
… 74 photos/s Generic picture-grid challenges built from a Caltech-256 subset (21 everyday categories: traffic light, fire hydrant, bus, bicycle, motorcycle, boat, bridge, palm tree, zebra, …), solved zero-shot with EmbeddingGemma 2: each photo is embedded (vision on the GPU, text on the Neural Engine) and ticked when the asked-for category beats the other 20. Nothing is taken from or sent to a real CAPTCHA service; the look is generic. Hands-free: three mixed grids and three lookalike grids (bus vs fire truck, horse vs zebra) at a watchable pace, then turbo (--grids=200): the next batch of four grids is decoded and embedded while the current one is drawn. M5 Pro: 206 grids / 1,854 photos in 25 s, 74 photos/s, 99% of photos right, 194/206 grids perfect. demo.sh adds macmon above the live log. EmbeddingGemma2Vision.embed(images:) takes maxInFlight.
…ccuracy to 0.1% M5 Pro: 1,006 grids / 9,054 photos in 160.6 s, 99.5% of photos right, 965/1,006 grids perfect (56 photos/s with other demos sharing the GPU; 74 photos/s alone).
Alex-Wengg
force-pushed
the
feat/embeddinggemma2-vision
branch
from
October 9, 2026 05:32
22f7668 to
a6f6df9
Compare
… the stats One square per grid drawn in a Canvas, sized so the whole run fits (1,006 grids at ~9 pt); nothing scrolls away.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #31 (audio). Adds EmbeddingGemma 2's vision encoder (FluidInference/embeddinggemma-2-coreml @ f9d567d5): images embed into the same space as text and audio.
🤖 Generated with Claude Code