|
Download GUIDE.md from DaisyChainAI/DaisyChain-Infer: direct link, hf CLI and curl.
- Browser
- Download file 9.78 kB
-
https://huggingface.co/DaisyChainAI/DaisyChain-Infer/resolve/main/GUIDE.md
- Command line
-
hf download hf://DaisyChainAI/DaisyChain-Infer/GUIDE.md
-
curl -L -o GUIDE.md https://huggingface.co/DaisyChainAI/DaisyChain-Infer/resolve/main/GUIDE.md
9.78 kB
| # β DaisyChain-Infer β run one model across your devices | |
| Part of **DaisyChain** β https://huggingface.co/DaisyChainAI | |
| Point it at a **Hugging Face model**, open the page on two or more devices, and | |
| each one takes a slice of the model's layers. A token is produced by passing | |
| the hidden state around the ring, peer-to-peer over WebRTC. Every multiply runs | |
| through the same **verified INT8 units** as the rest of DaisyChain β WebGPU | |
| where available, the identical math on CPU where not. | |
| **This runs locally.** It is a project you clone and start on your own | |
| machines; there is no hosted deployment, and it is not a Space. | |
| ```bash | |
| npm install | |
| npm start # http://localhost:8788 | |
| ``` | |
| Built from [DaisyChain-Train](https://huggingface.co/DaisyChainAI/DaisyChain-Train) | |
| and [DaisyChain-Web](https://huggingface.co/spaces/Quazim0t0/DaisyChain-Web) β | |
| the verified units, the WGSL kernels and their exact gates, and the WebRTC mesh | |
| are carried over unchanged. | |
| --- | |
| ## What is actually new here | |
| DaisyChain-Train is explicit about its limit: it **pools compute, not memory**. | |
| Every node holds a full replica, so a model bigger than one machine cannot be | |
| trained, and chaining five laptops does not give you one big machine. | |
| Inference is where that limit can be lifted, because a forward pass is a chain: | |
| layer `l` needs layer `l-1`'s **output**, never its **weights**. | |
| And because safetensors gives every tensor an exact byte range, **each device | |
| downloads only its own layers, straight from the Hub**. No device β not even | |
| the one driving the run β ever holds the whole model. | |
| | | DaisyChain-Train | DaisyChain-Infer | | |
| |---|---|---| | |
| | What is split | the **batch** | the **model** | | |
| | What crosses the wire | gradients (whole-model sized, every step) | hidden states (`TΓhidden` floats, per hop) | | |
| | Every device holds | the entire model | its own layers only | | |
| | Pools | compute | **memory** | | |
| | More devices means | more throughput | a **bigger model fits** | | |
| Running SmolLM-135M across three devices, each one downloads and holds | |
| 135β243 MB of a 513 MB model. The activation moving between them is tens of KB. | |
| That asymmetry is why this works. | |
| ## The ring | |
| ``` | |
| βββββββββββββββββ hidden state (TΓhidden f32) ββββββββββββββ | |
| βΌ β | |
| βββββββββββ βββββββββββ βββββββββββ β | |
| β stage 0 βββββββββΆβ stage 1 βββββββββΆβ stage 2 βββββββββββββββββ | |
| β HEAD β β layers β β layers β | |
| β embed β β 10β19 β β 20β29 β | |
| β layers β βββββββββββ βββββββββββ | |
| β 0β9 β | |
| β lm_head ββββ the returning state becomes the next token | |
| βββββββββββ | |
| ``` | |
| It is a **ring, not a line**, and weight tying forces that. When `lm_head` is | |
| tied to the embedding table β as it is in most small models β the largest | |
| tensor is needed at *both* ends: to embed the prompt and to produce the logits. | |
| Copying it to the last stage would hand back most of the memory just pooled, so | |
| the last stage returns its hidden state to the head, which owns the embedding | |
| **once** and does both ends. | |
| A consequence worth stating plainly: a middle stage receives layer weights | |
| only. It never gets the embedding table, so **it never sees the vocabulary** β | |
| it passes floats it cannot interpret. | |
| ## Models it can run | |
| Any Hugging Face repo with `.safetensors` weights whose architecture is: | |
| - **Llama-style** β Llama, Mistral, Qwen2/2.5, SmolLM, TinyLlama | |
| (RMSNorm, RoPE, grouped-query attention, SwiGLU) | |
| - **GPT-2-style** β LayerNorm with bias, learned positions, fused QKV, GELU | |
| Verified working end to end: `HuggingFaceTB/SmolLM-135M`, | |
| `openai-community/gpt2`, `Qwen/Qwen2.5-0.5B`. | |
| Anything else is **refused by name**, not approximated. Treating an unknown | |
| architecture as a known one produces fluent, confident, wrong output, which is | |
| the failure this project spends its verification budget making impossible. The | |
| same applies to tokenizers: byte-level BPE is implemented, and a | |
| SentencePiece/Unigram repo is rejected rather than tokenized approximately. | |
| Practical ceiling: weights are held as f32, so budget ~4 bytes per parameter | |
| across the group, and remember the head also carries the embedding table. | |
| ## The token | |
| Gated or private models need a Hugging Face token. You are asked for one | |
| **once**, when a request actually fails for want of it β not up front. | |
| It is held in memory for that tab and nowhere else: **not** localStorage, not | |
| sessionStorage, not a cookie, not the URL, never written to the log, and | |
| **never sent to another device**. Each device is asked for its own, because | |
| each device fetches its own layers. Reloading the tab forgets it, and there is | |
| a *Forget token* button. | |
| ## Verification | |
| The trainer's stack carries over unchanged: exact init gates on every kernel on | |
| every boot, the continuous random-cell audit at live shapes, and the | |
| cross-device kernel probe. | |
| But a pipeline moves the risk somewhere those instruments cannot reach: | |
| > In data-parallel **training**, every peer computes the same thing, so a | |
| > device with broken arithmetic shows up as a diverging replica. In a | |
| > **pipeline**, each stage computes something *different* and nobody else | |
| > repeats it. There is no replica to compare against. A wrong middle stage | |
| > produces a fluent, confident, wrong answer, and no consistency check | |
| > anywhere in the system would notice. | |
| Four things close that: | |
| 1. **The kernel probe**, which matters more here than in the trainer β the same | |
| seeded GEMM on every device, so it stays comparable even when the real work | |
| is not. A stage whose probe disagrees is flagged before it is given layers. | |
| 2. **Activation integrity hashes** on every hop, plus the model fingerprint, so | |
| a stage still holding a slice of a *different* model refuses the hop instead | |
| of silently mixing two models. | |
| 3. **Structural validation** of every assignment β a plan that does not cover | |
| each layer exactly once is refused, because it would still generate fluent | |
| text with a layer missing. | |
| 4. **The differential check** (tick *Verify*). The driving device re-runs the | |
| identical prompt with every layer locally and compares token ids. It is the | |
| only instrument that can show a distributed answer is *right* rather than | |
| merely self-consistent. | |
| ```bash | |
| npm test # pipeline equivalence + wire protocol + loader | |
| ``` | |
| `test_pipeline.js` asserts that splitting changes **no bit**, for *both* | |
| architecture families, across several uneven splits including a head that keeps | |
| no layers. Two cases are mutation checks β a stage that silently drops a layer, | |
| and stages applied out of order β both of which *must* fail the comparison. | |
| Without those, a test where both sides call the same code proves nothing. | |
| This was also confirmed against a real model: SmolLM-135M's 30 layers split | |
| `[10,10,10]`, `[1,14,15]`, `[0,15,15]` and `[5,5,5,5,5,5]` all produced token | |
| sequences identical to the single-device run. | |
| `test_loader.js` builds safetensors files and reads them back, checks BF16/F16 | |
| widening is exact (including subnormals and signed zero), and round-trips the | |
| BPE tokenizer. `test_wire.js` round-trips the protocol and refuses the | |
| malformed messages that would otherwise produce plausible wrong answers β a | |
| plan with a gap, a crafted repo id, a one-ulp flip deep in an activation, a | |
| `-0` flipped to `+0`. | |
| ### One bug, and what it cost to find | |
| The first live two-device run stalled. The head is both stage 0 *and* the | |
| ring's terminus, so an activation addressed to index 0 meant "start the lap" | |
| outbound and "the lap is finished" inbound. The head read the return leg as its | |
| own turn, re-ran its own layers, forwarded again, and the lap never closed. | |
| Every number was correct. Every message round-tripped. **Neither test suite | |
| could see it** β the pipeline test calls the stages in order itself, and the | |
| codec test only checks bytes. The defect lived in the *route*, precisely the | |
| category DaisyChain-Web's own self-corpus writeup identified as needing | |
| different instruments rather than better oracles. The fix gave the return leg | |
| its own address, and the routing decision moved out of the event handler into a | |
| pure function so `walkLap` in `test_wire.js` can walk a lap and assert it | |
| visits each stage once and terminates. That test fails against the old | |
| behaviour; it was checked. | |
| ## Honest limits | |
| - **Latency, not bandwidth, is the cost.** Every token pays one round trip per | |
| stage. More stages buy capacity, not speed β expect tokens/sec to *fall* as | |
| you add devices. | |
| - **No KV cache.** Every token re-runs the whole window, which is also what | |
| keeps the split bit-comparable against an unsplit run. Context length is the | |
| dominant per-token cost. | |
| - **f32 in memory.** Weights are widened from F16/BF16 on load, so a stage | |
| costs 4 bytes per parameter it holds. | |
| - **The head is a single point of failure**, and a stage that drops stalls the | |
| ring; press Generate again to re-plan around whoever is still connected. | |
| - **No authentication of activations.** A malicious stage that runs correct | |
| math but returns a crafted activation is not caught by any of this β the | |
| verification proves the *computation* is right on every honest device. Peers | |
| see each other's IPs, and the head sees your prompt. Run rings with people | |
| you trust. | |
| --- | |
| **License:** MIT Β· **Author:** Dean Byrne (Quazim0t0) Β· **Org:** DaisyChainAI | |