Low-Zi-Hong/ESP32s3-LLM-Cluster: 7-node ESP32-S3 cluster operating a 0.4B LLM by way of 1.58-bit (BitNet) ternary quantization over SPI daisy-chain · GitHub


A distributed pipeline inference engine on a number of ESP32S3 operating 1.58-bit (BitNet) Language mannequin.

This undertaking runs a sliced 0.5B LLM throughout a cluster of seven ESP32s3. One act as grasp and others are node. The grasp node runs the tokenizer and embeding and the opposite consideration layer and MLP ran on the nodes. The grasp and node talk via excessive velocity SPI Daisy-Chain.

┌─────────────────────────────────────────────────────────┐
│                     MASTER NODE                         │
│                                                         │
│  [ Prompt ] ---> BPE Tokenizer                          │
│                       │                                 │
│                 Token Embedding                         │
│             (INT4, ~14MB in Flash)                      │
│                       │                                 │
│             (SPI CH A - TX to Node 1)                   │
└───────────────────────┬─────────────────────────────────┘
                        │ Hidden State Vector (FP32)
                        ▼
┌─────────────────────────────────────────────────────────┐
│                    COMPUTE NODE 1                       │
│             (SPI CH B - RX from Master)                 │
│                                                         │
│  ► Layer 0 to three (4x Transformer Blocks)                 │
│    • RMSNorm (FP16 scaled to FP32)                      │
│    • 1.58-bit Attention (Q, Ok, V, O proj) + RoPE        │
│    • KV Cache (PSRAM)                                   │
│    • 1.58-bit MLP (Gate, Up, Down proj)                 │
│                                                         │
│             (SPI CH A - TX to Node 2)                   │
└───────────────────────┬─────────────────────────────────┘
                        │
                       ... (Nodes 2 to five)
                        │
                        ▼
┌─────────────────────────────────────────────────────────┐
│                    COMPUTE NODE 6                       │
│             (SPI CH B - RX from Node 5)                 │
│                                                         │
│  ► Layer 20 to 23 (4x Transformer Blocks)               │
│    • Same 1.58-bit Architecture                         │
│                                                         │
│             (SPI CH A - TX again to Master)              │
└───────────────────────┬─────────────────────────────────┘
                        │
                        ▼
┌─────────────────────────────────────────────────────────┐
│                     MASTER NODE                         │
│             (SPI CH B - RX from Node 6)                 │
│                                                         │
│                 Final RMS Norm                          │
│             (FP16, 64KB in 'fnorm' partition)           │
│                       │                                 │
│         LM Head (Tied to INT4 Embeddings)               │
│                       │                                 │
│               Greedy Sampling                           │
│                       │                                 │
│  [ Output ] <--- Next Token ID                          │
└─────────────────────────────────────────────────────────┘

pls refer workflow guide to start out with the undertaking.

.
├── README.md                   # Project documentation
├── workflow.md                 # Step-by-step flashing, mannequin prep & wiring information
├── .gitignore                  # Git ignore guidelines for construct information & binaries
│
├── docs/                       
│   └── photographs/                 # Architecture diagrams and {hardware} photographs
│
├── master_board/               # Firmware for the Master Node (ESP-IDF)
│   ├── most important/
│   │   ├── most important.cpp            # Master orchestrator, person I/O & BPE tokenizer
│   │   ├── embedding.cpp       # INT4 embedding lookup logic
│   │   ├── lm_head.cpp         # LM Head mapping and grasping sampling
│   │   └── spi_bus.cpp         # Master dual-channel SPI driver
│   ├── partitions.csv          # Custom partition desk (token, mannequin, fnorm)
│   └── CMakeLists.txt
│
├── node_firmware/              # Firmware for the Compute Nodes (ESP-IDF)
│   ├── most important/
│   │   ├── most important.cpp            # Node employee entry level & inference loop
│   │   ├── bitlinear.cpp       # 1.58-bit ternary linear layer implementation
│   │   ├── bitlinear_forward.S # Assembly optimized MAC ops for 1.58-bit
│   │   ├── qwen_attention.cpp  # Qwen Attention, RoPE & KV-Cache runtime
│   │   ├── lut_table.cpp       # Look-up tables for excessive optimization
│   │   └── spi_bus.cpp         # Daisy-chain SPI DMA receiver/transmitter
│   ├── partitions.csv          # Layer partition format for Node
│   └── CMakeLists.txt
│
├── python_tools/               # PC-side quantization & preprocessing suite
    ├── crop_token.py           # Vocabulary pruning (scales right down to 32K tokens)
    ├── crop_model_weight.py    # Embedding matrix slicing
    ├── qat_158.py              # BitNet QAT (Quantization-Aware Training) fine-tuning
    ├── bit4_embedding.py       # INT4 weight packer for embeddings
    ├── pack_tokenizer_bin.py   # Serializes tokenizer guidelines into ESP32 .bin
    ├── pack_model_bin.py       # Packs 1.58-bit layer chunks for bodily alignment
    ├── look_model_structure.py # Debug software for inspecting .safetensors
    └── flash_*.bat             # Multi-threaded quick flashing scripts

This undertaking is licensed below the MIT License – see the LICENSE file for particulars.

Inspiration, associated works, and references:



Source link