Skip to content

EmbeddingGemma 2 images: ViT on Core ML, 79 images/s - #32

Open
Alex-Wengg wants to merge 4 commits into
mainfrom
feat/embeddinggemma2-vision
Open

Alex-Wengg wants to merge 4 commits into
mainfrom
feat/embeddinggemma2-vision

Conversation

@Alex-Wengg

Copy link
Copy Markdown
Member

Stacked on #31 (audio). Adds EmbeddingGemma 2's vision encoder (FluidInference/embeddinggemma-2-coreml @ f9d567d5): images embed into the same space as text and audio.

  • EmbeddingGemma2Vision: HF-matching preprocessing (resize, patches, positions, 3×3 pooling) + Core ML vision_70/140/280 (shared weights) on the GPU + text model on the ANE. 79 images/s end to end at 70 tokens.
  • Parity: fp32 == sentence-transformers (cos 1.0); Swift vs PyTorch cos ≥ 0.991 (CoreGraphics resize + fp16). Zero-shot Oxford Pets 87.6% (PyTorch 87.0%; 140 / 280 tokens: 88.4 / 89.5%).
  • GPU over ANE: all ops do run on the ANE, but full attention over 630–2,520 patches is 2.5–5× slower there.
  • ImageSearchCheck CLI; unit test for the resize targets vs HF.

🤖 Generated with Claude Code

…ext space

EmbeddingGemma2Vision runs the model's vision encoder (Gemma 4 ViT, 16
layers, 2-D axial RoPE, 3x3 pooling, embed_vision projection) from
EmbeddingGemma2Vision.mlpackage on FluidInference/embeddinggemma-2-coreml
(revision f9d567d5, its own pinned bundle). One fixed-shape function per
soft-token budget, shared weights: vision_70 / vision_140 / vision_280
(630 / 1,260 / 2,520 patches). The pooled tokens go through the text model
as <bos> <|image> tokens <image|> <eos>, so images share the text space.

- Host side matches HF Gemma4ImageProcessor: aspect-preserving resize to
  multiples of 48 px within the patch budget (target sizes unit-tested
  against get_aspect_ratio_preserving_size), patches row by row
  [16][16][RGB], (x, y) positions, 3x3 pooling matrix, and the patch
  embedder's position tables (position_embeddings.f16, rows 0-1023) added
  on the host: the in-graph gather put the model's first ops on the CPU.
- GPU (default): 70 tokens 14.6 ms, 140 tokens 34 ms per image. Every op
  also runs on the Neural Engine, but its time grows with the square of the
  patch count (37 / 102 / 341 ms), so the GPU wins at every budget.
- Accuracy: fp32 wrapper == sentence-transformers (cos 1.000000); Swift
  (CoreGraphics resize + fp16) vs PyTorch fp32 cos mean 0.998, min 0.991.
  Zero-shot Oxford Pets (370 photos, 37 breeds): 87.6% in Swift at 70
  tokens (PyTorch 87.0%; 140 tokens 88.4%, 280 tokens 89.5%).
- Swift end to end at 70 tokens: 79 images/s (vision on the GPU, text on
  the Neural Engine, four images in flight).

ImageSearchCheck scores zero-shot Pets (photos cached by ImageSortDemo)
and timing, with --budget and --dump.
… 74 photos/s

Generic picture-grid challenges built from a Caltech-256 subset (21 everyday
categories: traffic light, fire hydrant, bus, bicycle, motorcycle, boat,
bridge, palm tree, zebra, …), solved zero-shot with EmbeddingGemma 2: each
photo is embedded (vision on the GPU, text on the Neural Engine) and ticked
when the asked-for category beats the other 20. Nothing is taken from or sent
to a real CAPTCHA service; the look is generic.

Hands-free: three mixed grids and three lookalike grids (bus vs fire truck,
horse vs zebra) at a watchable pace, then turbo (--grids=200): the next batch
of four grids is decoded and embedded while the current one is drawn.
M5 Pro: 206 grids / 1,854 photos in 25 s, 74 photos/s, 99% of photos right,
194/206 grids perfect. demo.sh adds macmon above the live log.

EmbeddingGemma2Vision.embed(images:) takes maxInFlight.
…ccuracy to 0.1%

M5 Pro: 1,006 grids / 9,054 photos in 160.6 s, 99.5% of photos right, 965/1,006 grids perfect
(56 photos/s with other demos sharing the GPU; 74 photos/s alone).
@Alex-Wengg
Alex-Wengg force-pushed the feat/embeddinggemma2-vision branch from 22f7668 to a6f6df9 Compare October 9, 2026 05:32
@Alex-Wengg
Alex-Wengg changed the base branch from feat/embeddinggemma2-audio to main October 9, 2026 05:32
… the stats

One square per grid drawn in a Canvas, sized so the whole run fits (1,006 grids at ~9 pt);
nothing scrolls away.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant