projections · keep separate
DeepSeek-V4-Flash × pd-bridge × 2 Spark + Studio × 10 GbE × 105K
Axonometric of discrete cells — not a metric rack scan. Measured boards are solid. Named rack ghosts are outlines. No line means no evidence of a link.
SPACE · boards as places
Two Sparks and one Studio are the measured join. The other boards in the photo are named, not joined.
On a narrow screen, scroll the volume sideways to the Studios.
- measured
- documented (other lab / other mechanism)
- recipe, unmeasured
- posted, not RESULTS.md
- goal
- blocked
- unknown (blank)
model × harness · one slice at T
Every pairing we will admit
Columns are models this bench could even name. Rows are harnesses that have touched this silicon family in public. Most intersections stay empty on purpose. Click a blank: the inspector will say unknown.
mind communication · capability at this T
What can cross the ISA wall
A mind here is tokens leaving metal after a prefix has been scheduled. Communication is a file the consumer can open — not a CUDA object, not a shared address space. The lists below are the Capability projection written as sentences.
instance · 2026-09-06 · RESULTS.md
Measured windows
DeepSeek-V4-Flash 284B / 13B active. Prefill: 2× DGX Spark, vLLM TP2, FP8, 149 GB. Decode: Mac Studio M3 Ultra, oMLX 0.6.4, MXFP4, 156 GB. Link: ordinary 10 GbE. Decode rate unchanged both ways (23–25 tok/s). Quality eval 5/5, same as native — five questions, not “equivalent.”
| cold prompt | Studio alone | Sparks → Studio | ratio | status |
|---|---|---|---|---|
| ~25K | 42.6 s | 28.2 s | 1.5× | measured |
| ~82K | 205.8 s | 72.9 s | 2.8× | measured (old writer ceiling; fixed) |
| ~105K | 245.6 s | 75.5 s | 3.3× | measured · validated envelope top |
| ~241K | 732.3 s | 200.3 s | 3.7× | measured at 262K prefill ceiling |
| 241K warm | — | 4.9 s / 19.3 s | bypass | measured · file already scheduled |
Payload ~10 KB/token (0.80 GB for 81K, pulled in 1.08 s). Network is not the bottleneck; prefill is ~65% of wall. Warm turns never take the bridge. A silent native fallback can no longer enter a results table: every reply carries X-PD-Bridge (complete / partial B/T / declined).
Posted, not RESULTS.md
A 2026-09-07 lab post added longer rows and a rack photo (7 Sparks + 5 Studios). Those numbers are not the validated envelope. They sit in this map as posted, so they cannot be mistaken for RESULTS.md.
| load | Studio alone | 2 Sparks → Studio | ratio | status |
|---|---|---|---|---|
| 100k | 5.4 min | 0.9 min | 5.8× | posted |
| 241k cold | 12 min | 3 min | ~4× | posted (aligns ~732 s / ~200 s) |
| 500k | 37.9 min | 6.3 min | 6.1× | posted |
| 1.25M | 2.43 h | 23.1 min | 6.3× | posted |
| 2.09M | 5.64 h | 52.3 min | 6.5× | posted |
Rack photo = goal inventory (896 GB CUDA + ~1.5 TB Metal). Bench = 2 Sparks + 1 Studio. Goal window 262K → 524K → 2M and a 1–2T 4-bit model on five Studios are named, not measured.
nine strata · the same volume, as matter
Substrate
The 4D slice sits on geology. Click a stratum if you want the mineral clock instead of a configuration cell.
- 01
Mineral
SiO2, Cu, Au, Li, water, rare earths in fans. Geology before software.
- 02
Power
Watts in. Decode is cheap. Prefill is the expensive clock.
- 03
Silicon
GB10 CUDA vs M3 Ultra Metal. Same element, two ISAs, two caches.
- 04
Cooling
Every watt that is not a token is heat. Thermal ceiling delays the first word.
- 05
Metal
The board you unplug. Hostname is a path. Die and serial are the place.
- 06
File
Weights, prefix, memory files. Do not ship an unreadable KV.
- 07
Schedule
Intent etched, not yet speaking. Existence of the prefix is not execute.
- 08
Prompt
Unlock. First word is the cost. Then 23–25 tok/s, all day.
- 09
Mind
Tokens leaving metal. Loop: mineral → die → file → schedule → prompt → token → new file.
CUDA prefill → Metal prefix store
Two halves. No shared cache format.