mlboydaisuke commited on
Commit
fdb16fc
Β·
verified Β·
1 Parent(s): a27dc73

README: ios/ is the JIT .aimodel, ios-h18p/ the h18p bundle; iPhone 18 Pro load numbers

Browse files
Files changed (1) hide show
  1. README.md +40 -12
README.md CHANGED
@@ -28,7 +28,7 @@ Zoo card, recipe and gate transcript: [coreai-model-zoo/models/granite-embedding
28
 
29
  IBM's 97M-parameter **multilingual text embedder** β€” a ModernBERT encoder, 384-d CLS-pooled
30
  unit vectors, Japanese and English among its languages β€” as a static `.aimodel` for macOS 27
31
- and, ahead-of-time compiled, for the iPhone 17 Pro.
32
  [`ibm-granite/granite-embedding-97m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2)
33
  (Apache-2.0, revision `835ad1408…`) is the **smallest embedder in this catalog** (390 MB fp32,
34
  against 1.2 GB for EmbeddingGemma-300m and 1.1 GB for Qwen3-Embedding-0.6B) and its **first
@@ -104,6 +104,25 @@ The Mac h16c AOT twin also passed 70/70 (same numerics), but its timings were ta
104
  lane's GPU evaluation running and are not reported. The w8 variant on Mac is gated on **CPU
105
  only** (min cosine 0.999410, max |err| 5.67e-3, ranking exact); Mac GPU for w8 was not run.
106
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
107
  The fixed grid computes every position, so pick the smallest grid that covers the text: S=128
108
  for queries and short notes, S=512 for passages. **fp32 is the default.** w8 is a storage
109
  option only β€” 22% smaller, not faster here β€” because the 180,000Γ—384 fp32 vocabulary table is
@@ -123,7 +142,8 @@ matched to the wrong text) must FAIL.
123
  **63 instead of 64** passes the embedding gate (cos 0.99995) and fails only the layer gate β€”
124
  which is why the layer gate exists. Whole-model **fp16 fails** this layer gate on both grids.
125
  - **Export**: the torch-exported, decomposed graph is gated before conversion, on both grids.
126
- - **Runtime**: Mac CPU and GPU (JIT), Mac h16c AOT, iPhone h18p AOT β€” the tables above.
 
127
  - **w8**: the same gate at prepared, finalized and decomposed stages, 48 `lut_to_dense` ops
128
  counted, palettes hashed; the iOS w8 export reuses the Mac palettes byte for byte.
129
 
@@ -142,17 +162,24 @@ platform β†’ folder.
142
  |---|---|---|---|---:|
143
  | `macos/fp32-s512/` **(default)** | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
144
  | `macos/fp32-s128/` | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
145
- | `ios/fp32-s512/` **(default)** | iOS 27, **h18p only** | AOT `.aimodelc` | `granite97m_fp32_s512_bound.h18p.aimodelc` | 390,308,788 |
146
- | `ios/fp32-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_fp32_s128_bound.h18p.aimodelc` | 390,081,410 |
147
  | `macos/w8-fp32table-s512/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
148
  | `macos/w8-fp32table-s128/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
149
- | `ios/w8-fp32table-s512/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s512_r02.h18p.aimodelc` | 305,479,184 |
150
- | `ios/w8-fp32table-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s128_r02.h18p.aimodelc` | 305,251,774 |
151
-
152
- The `ios/` bundles are compiled for one device architecture (`h18p`, the iPhone 17 Pro) with
 
 
 
 
 
 
153
  `xcrun coreai-build compile --platform iOS --min-deployment-version 27.0 --preferred-compute gpu
154
- --architecture h18p` (coreai-build 3600.83.1). **Never load an iOS bundle on a Mac.** Other
155
- phones need their own compile from the recipe; the source IR is reproducible, not shipped.
 
156
 
157
  Convert yourself: [`conversion/granite_embedding/`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/README.md)
158
  β€” five staged scripts; [`recipe.toml`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/recipe.toml) names the commands.
@@ -181,6 +208,7 @@ epsilon because the converter's `F.normalize` decomposition drops it.
181
 
182
  Apache-2.0 at the pinned upstream revision; this repo carries IBM's unmodified card as
183
  `UPSTREAM_README.md` and a `LICENSE-NOTE.md` listing the changes (static graph, in-graph
184
- pooling, optional w8 palettes, h18p compile). Not tested: other phones or OS builds, the Mac GPU
185
- with w8, the Neural Engine, dynamic or batched shapes, S > 512, languages beyond the JA/EN
 
186
  fixtures, retrieval quality on a benchmark, sustained thermals, true cache-cold load.
 
28
 
29
  IBM's 97M-parameter **multilingual text embedder** β€” a ModernBERT encoder, 384-d CLS-pooled
30
  unit vectors, Japanese and English among its languages β€” as a static `.aimodel` for macOS 27
31
+ and iOS 27, with bundles compiled ahead of time for the iPhone 17 Pro beside it.
32
  [`ibm-granite/granite-embedding-97m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2)
33
  (Apache-2.0, revision `835ad1408…`) is the **smallest embedder in this catalog** (390 MB fp32,
34
  against 1.2 GB for EmbeddingGemma-300m and 1.1 GB for Qwen3-Embedding-0.6B) and its **first
 
104
  lane's GPU evaluation running and are not reported. The w8 variant on Mac is gated on **CPU
105
  only** (min cosine 0.999410, max |err| 5.67e-3, ranking exact); Mac GPU for w8 was not run.
106
 
107
+ **iPhone 18 Pro** (iPhone19,2, h19p), iOS 27.0 build **24A437**, the JIT `.aimodel` bundles in `ios/`
108
+ (the same files as `macos/`), loaded 2026-09-26 by the zoo's DecideGate app in its load-only mode,
109
+ GPU-preferred, without the increased-memory entitlement. Each first load was the first after a fresh
110
+ install of the app. The call is one run on all-zero inputs. Each first load wrote a specialization of
111
+ about the bundle's size into the app container, and the load after a relaunch reused it. One measurement
112
+ per graph
113
+ ([knowledge/jit-distribution.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/jit-distribution.md)).
114
+
115
+ | JIT bundle in `ios/` | MB | first load | first call | load after relaunch |
116
+ |---|---:|---:|---:|---:|
117
+ | `fp32-s128/granite97m_fp32_s128_bound.aimodel` | 390 | 1.01 s | 619 ms | 0.26 s |
118
+ | `fp32-s512/granite97m_fp32_s512_bound.aimodel` | 390 | 0.40 s | 88 ms | 0.26 s |
119
+ | `w8-fp32table-s128/granite97m_w8_fp32table_s128.aimodel` | 305 | 0.63 s | 132 ms | 0.21 s |
120
+ | `w8-fp32table-s512/granite97m_w8_fp32table_s512.aimodel` | 306 | 0.33 s | 93 ms | 0.21 s |
121
+
122
+ The iPhone gate above (the iPhone 17 Pro rows) ran the `h18p` export, now in `ios-h18p/`. The JIT IR in
123
+ `ios/` was only loaded and called once on the iPhone 18 Pro. That call returned a 384-value float32
124
+ embedding with no non-finite values; its parity with the reference was not re-measured.
125
+
126
  The fixed grid computes every position, so pick the smallest grid that covers the text: S=128
127
  for queries and short notes, S=512 for passages. **fp32 is the default.** w8 is a storage
128
  option only β€” 22% smaller, not faster here β€” because the 180,000Γ—384 fp32 vocabulary table is
 
142
  **63 instead of 64** passes the embedding gate (cos 0.99995) and fails only the layer gate β€”
143
  which is why the layer gate exists. Whole-model **fp16 fails** this layer gate on both grids.
144
  - **Export**: the torch-exported, decomposed graph is gated before conversion, on both grids.
145
+ - **Runtime**: Mac CPU and GPU (JIT), Mac h16c AOT, iPhone h18p AOT β€” the tables above; iPhone 18 Pro
146
+ JIT: load and one zero-input call only.
147
  - **w8**: the same gate at prepared, finalized and decomposed stages, 48 `lut_to_dense` ops
148
  counted, palettes hashed; the iOS w8 export reuses the Mac palettes byte for byte.
149
 
 
162
  |---|---|---|---|---:|
163
  | `macos/fp32-s512/` **(default)** | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
164
  | `macos/fp32-s128/` | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
165
+ | `ios/fp32-s512/` **(default)** | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
166
+ | `ios/fp32-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
167
  | `macos/w8-fp32table-s512/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
168
  | `macos/w8-fp32table-s128/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
169
+ | `ios/w8-fp32table-s512/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
170
+ | `ios/w8-fp32table-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
171
+ | `ios-h18p/fp32-s512/` | iOS 27, **h18p only** | AOT `.aimodelc` | `granite97m_fp32_s512_bound.h18p.aimodelc` | 390,308,788 |
172
+ | `ios-h18p/fp32-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_fp32_s128_bound.h18p.aimodelc` | 390,081,410 |
173
+ | `ios-h18p/w8-fp32table-s512/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s512_r02.h18p.aimodelc` | 305,479,184 |
174
+ | `ios-h18p/w8-fp32table-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s128_r02.h18p.aimodelc` | 305,251,774 |
175
+
176
+ `ios/` holds the same JIT bundles as `macos/`; every iPhone generation specializes them on its first
177
+ load. The `ios-h18p/` bundles moved there from `ios/` in revision `a27dc73e` (2026-09-26). They are
178
+ compiled for one device architecture (`h18p`, the iPhone 17 Pro) with
179
  `xcrun coreai-build compile --platform iOS --min-deployment-version 27.0 --preferred-compute gpu
180
+ --architecture h18p` (coreai-build 3600.83.1), from a separate iOS export whose IR is reproducible
181
+ from the recipe but not shipped. The iPhone 18 Pro refuses an h18p bundle with
182
+ `incompatibleCompiledAssetArchitecture`. **Never load an `ios-h18p/` bundle on a Mac.**
183
 
184
  Convert yourself: [`conversion/granite_embedding/`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/README.md)
185
  β€” five staged scripts; [`recipe.toml`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/recipe.toml) names the commands.
 
208
 
209
  Apache-2.0 at the pinned upstream revision; this repo carries IBM's unmodified card as
210
  `UPSTREAM_README.md` and a `LICENSE-NOTE.md` listing the changes (static graph, in-graph
211
+ pooling, optional w8 palettes, h18p compile). Not tested: phones other than the iPhone 17 Pro (the
212
+ h18p gate) and the iPhone 18 Pro (JIT load and one call), other OS builds, the JIT bundles' embedding
213
+ parity on an iPhone, the Mac GPU with w8, the Neural Engine, dynamic or batched shapes, S > 512, languages beyond the JA/EN
214
  fixtures, retrieval quality on a benchmark, sustained thermals, true cache-cold load.