FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in a single repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card. · GitHub

An open-source AI accelerator, developed by AI.
openTPU brings the teachings of auto-arch-tournament
to AI accelerators. It asks two questions: how far can AI brokers go at {hardware} design, and might
they construct the chip that runs their very own inference?
otpu-chat working LFM2.5-230M on the FPGA card (left), with otpu-smi exhibiting the cardboard’s
utilization and DRAM bandwidth (proper).
openTPU can be a studying undertaking. The complete accelerator lives in a single small monorepo that you simply
can learn finish to finish: the {hardware} design (SystemVerilog), the instruction set, a bit-exact
simulator, a kernel language and its compiler, and the host software program that drives an actual PCIe
card. If you need to perceive how an AI accelerator works, from a matmul in Python all the way down to
the wires, this can be a good place to start out.
The design runs ten trendy fashions with their actual weights on an Inspur YPCB-00338 card
(Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the cardboard produces the identical tokens because the
simulator, bit for bit.
| Model | Weights | Decode, gadget | Decode, wall | Prefill, gadget | DRAM whereas decoding |
|---|---|---|---|---|---|
| LFM2.5-230M | int8 | 59.0 tok/s | 52.3 tok/s | 295.6 tok/s | 14.5 GB/s (85% of peak) |
| LFM2.5-230M | 4-bit, int8 head | 85.8 tok/s | 82.1 tok/s | 335.4 tok/s | 14.1 GB/s (82%) |
| Qwen3-0.6B | int8 | 21.6 tok/s | 21.3 tok/s | 92.1 tok/s | 14.4 GB/s (84%) |
| Qwen3-0.6B | 4-bit, int8 head | 31.3 tok/s | 30.7 tok/s | 103.4 tok/s | 13.9 GB/s (82%) |
| Qwen3.5-0.8B | int8 | 17.6 tok/s | 16.3 tok/s | 61.4 tok/s | 14.5 GB/s (85%) |
| Qwen3.5-0.8B | 4-bit, int8 head | 24.5 tok/s | 23.3 tok/s | 66.7 tok/s | 14.1 GB/s (83%) |
| Gemma 4 E2B | 4-bit, int8 head | 10.57 tok/s | 10.53 tok/s | 32.1 tok/s | 15.6 GB/s (92%) |
| Gemma 4 E2B | 4-bit, 4-bit head | 12.14 tok/s | 12.09 tok/s | 29.9 tok/s | 15.5 GB/s (91%) |
| LFM2-2.6B | int8 | 6.05 tok/s | 6.03 tok/s | 21.4 tok/s | 16.1 GB/s (94%) |
| LFM2-2.6B | 4-bit, int8 head | 10.96 tok/s | 10.93 tok/s | 20.6 tok/s | 15.8 GB/s (93%) |
| SmolLM3-3B | int8 | 5.00 tok/s | 4.99 tok/s | 21.1 tok/s | 16.0 GB/s (94%) |
| SmolLM3-3B | 4-bit, int8 head | 8.74 tok/s | 8.72 tok/s | 22.8 tok/s | 15.7 GB/s (92%) |
| Phi-4-mini (3.8B) | int8 | 3.99 tok/s | 3.98 tok/s | 13.8 tok/s | 16.0 GB/s (94%) |
| Phi-4-mini (3.8B) | 4-bit, int8 head | 6.56 tok/s | 6.55 tok/s | 15.0 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-2B | int8 | 8.02 tok/s | 8.00 tok/s | 38.2 tok/s | 16.0 GB/s (94%) |
| Qwen3.5-2B | 4-bit, int8 head | 12.09 tok/s | 12.03 tok/s | 41.7 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-4B | 4-bit, int8 head | 5.88 tok/s | 5.87 tok/s | 12.9 tok/s | 15.7 GB/s (92%) |
| Gemma 4 E4B | int8, 4-bit head and down 0-23 | 3.78 tok/s | 3.75 tok/s | 14.8 tok/s | 16.0 GB/s (94%) |
Measured on the cardboard: the primary three fashions on 2026-09-29 with the manufacturing picture
deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and
4B and Gemma 4 on 2026-10-01, with construct B, deploy_fused133c_79c5707a, manufacturing since then.
Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% sooner than e698dcd7 (Gemma 4 E2B 10%),
at 91-94% of the DRAM peak as an alternative of 82-87%. Qwen3.5-4B’s int8 picture is over 4 GiB.
- The picture: principal e698dcd at 133.33 MHz, one bitstream for all fashions. It has LiteDRAM
controllers calibrated by a small CPU contained in the reminiscence core, a four-column systolic matrix
unit and the stream engine (docs/stream.md). DDR3-1066, with a 17.1 GB/s
peak. - The host: the cardboard sits in opentpu (Intel Core i7-4790).
- Method,
instruments/qual/perf.py: decode is 64 grasping tokens after a 512-token immediate, with the
host’s argmax within the loop (not streamed). “Device” counts solely the cycles the accelerator runs;
“wall” provides the host. Prefill is the 512-token immediate, on the gadget. - DRAM site visitors comes from the cardboard’s personal counters whereas it runs.
- Gemma 4 E2B retains its per-layer embedding tables on the cardboard (3.5-3.6 GiB photographs;
docs/gemma4.md); in int8 it doesn’t match. It matches Hugging Face’s grasping
tokens on three prompts with both head. E4B’s desk (2.95 GB) stays on the host, which
copies one 11 KB row into the cardboard per token; its picture is 3.96 GiB, int8 with the top and
the primary 24 layers’ down projections in 4-bit (docs/gemma4_e4b.md).
In the cardboard’s personal decode loop (the cardboard selecting each token) Gemma 4 decodes sooner: E2B
11.01 / 12.73 tok/s (int8 / 4-bit head), E4B 3.83 tok/s, on the gadget. - Every configuration matches the simulator token for token, per-position and with the
resident decode program. More element in docs/board.md.
With the logits streamed again whereas the cardboard runs (instruments/decode_profile.py, 96 tokens), 4-bit
decode is quicker, in gadget / wall tok/s:
- LFM2: 89.5 / 84.5;
- Qwen3: 33.7 / 33.3;
- Qwen3.5: 24.6 / 24.2;
- LFM2-2.6B: 11.07 / 11.02 (construct B);
- SmolLM3-3B: 8.92 / 8.89 (construct B);
- Phi-4-mini: 6.69 / 6.67 (construct B).
The earlier manufacturing picture, se-cand3, was constructed with the Xilinx MIG, a two-column matrix
unit and a 120.755 MHz clock. Measured the identical manner, the brand new picture:
- decode: inside 2.3% of se-cand3’s in each configuration. Decode is sure by DRAM, and
LiteDRAM reads at 82-85% of the DDR3 peak, because the MIG did. - prefill: 1.3x (Qwen3.5) to 2.0x (LFM2 4-bit) sooner.
- calibration: when the picture begins, the core’s CPU calibrates each DDR3 channels in 12 s,
with no host involvement.
The earlier photographs and their numbers are in docs/board.md, part 5.
Mixture-of-experts fashions larger than the cardboard’s 4 GiB run with their consultants streamed from host
storage (docs/offload.md, part 10). The card routes every token and computes
each professional, and it retains the consultants in per-layer slots in its DRAM. The host solely copies
lacking consultants from a pool file into these slots, on the hyperlink’s charge (part 10.1). Measured
on 2026-10-01 with construct B (79c5707a), the cardboard’s personal decode loop selecting each token, 4-bit
consultants, int8 head:
- LFM2.5-8B-A1B (8.5B parameters, 1.7B lively): 10.6 tok/s over 160 tokens. 98.5% of professional
makes use of hit the slots, and 5.2 MB streamed per token. - Qwen3.5-35B-A3B (34.7B parameters, 3.0B lively): 3.95 tok/s, with Hugging Face’s 16
grasping tokens. 62% of professional makes use of hit, and 153 MB streamed per token at 1.41 GB/s over PCIe
(part 10.3). - Both match the simulator bit for bit.
4-bit weights (docs/quant.md) use FP4 values with two-level block scales, 4.25
bits per weight, and preserve the LM head in int8 for accuracy. They reduce the bytes per token by about
a 3rd and lift decode velocity by 40% (Qwen3.5) to 45% (Qwen3, LFM2), at a measurable value in
perplexity that docs/quant.md experiences per mannequin.
The host is almost out of the way in which. For LFM2 and Qwen3 the cardboard runs one decode program compiled
as soon as, which reads the place from a register and appears up its personal embedding and RoPE rows, and
the logits stream again whereas the cardboard remains to be working: the host provides 0.17 to 0.30 ms per
token on omarchy (0.45 to 1.3 ms on opentpu).
Qwen3.5 runs the identical manner for decode; its prefill nonetheless compiles every chunk’s program on the host,
forward of the cardboard.
Kernels in ol mlp, consideration, full mannequin layers
| @ol.jit
Language + compiler layouts, affine loop addressing, fusion
|
ISA 8 x 32-bit phrases per instruction
|
ISA simulator <======> RTL similar bits, checked by the checks
(Python) (SystemVerilog)
| Vivado bitstream
FPGA card Kintex-7 xc7k480t
| PCIe
Host otpu-chat, otpu-smi, otpu-lens
The machine is intentionally easy. A sequencer points one instruction per cycle to a couple
items: DMA strikes knowledge, the matrix unit multiplies int8 weights streamed from DRAM, the vector
unit does fp32 math, and a quantizer turns outcomes again into int8. There is not any cache and no
hidden scheduling: each knowledge motion is an instruction, so a hint reveals precisely the place the
cycles go. docs/isa.md describes the entire instruction set.
A kernel seems to be like this:
from opentpu import language as ol
@ol.jit
def mlp(h, gamma, w_gate, w_up, w_down, out, eps): # simplified; see kernels/mlp.py
x = ol.load(h)
xs = ol.quantize(rmsnorm(x, ol.load(gamma), eps))
g = ol.dot(xs, w_gate)
u = ol.dot(xs, w_up)
a = ol.all_gather(silu(g) * u)
y = ol.all_gather(ol.dot(a, w_down))
if ol.program_id() == 0:
ol.retailer(out, x + y)
Because each knowledge motion is an instruction, a hint of a run explains its velocity. Lens, the
profiler, data a run from the RTL, the simulator or the cardboard and opens it within the browser,
with a roofline, a timeline and per-instruction tables (docs/lens.md).
Lens replaying a part of a Qwen3 decode step. Colours present what every unit is doing in every
cycle: busy, ready on DRAM, or ready on one other instruction.
Everything besides the cardboard runs on a laptop computer.
pip set up -e .
pip set up pytest torch transformers
python3 -m pytest -q # RTL checks additionally want Verilator 5
hf obtain LiquidAI/LFM2.5-230M --local-dir fashions/LFM2.5-230M
otpu-chat --model lfm2 --backend isa # chat on the simulator
With a card, construct the bitstream (make bit in boards/ypcb-00338), load
it over JTAG, then run sudo otpu-setup and otpu-chat --backend board.
docs/board.md walks by the bring-up.
| Command | What it does |
|---|---|
otpu-chat |
chat with Qwen3-0.6B, LFM2.5-230M (--model lfm2), Qwen3.5-0.8B (--model qwen35), LFM2-2.6B (lfm2-2.6b), SmolLM3-3B (smollm3), Phi-4-mini (phi4-mini) or Qwen3.5-2B / 4B (qwen35-2b, qwen35-4b) |
otpu-smi |
temperature, energy, DRAM bandwidth and per-unit utilization |
otpu-lens |
file a run and open it within the profiler |
otpu-selftest, otpu-diag |
examine that the cardboard works |
- docs/isa.md: the instruction set. Everything else is constructed on it.
opentpu/kernelsand docs/compiler.md: how a kernel
turns into directions.opentpu/isasim.py: the simulator, which is the spec.rtl/: the {hardware}, ranging fromrtl/top/otpu_top.sv.- docs/lfm2.md, docs/qwen35.md, docs/llama.md,
docs/benchmarks.md: complete fashions and the place their cycles go. - docs/board.md: the bodily card, from clocks to PCIe.
- The previous few % of DRAM. Decode is sure by DRAM effectivity: it reads 82 to 85% of the
DDR3-1066 peak. Work on the LiteDRAM path’s effectivity is underneath manner. - Timing margin and space. The design closes 133.33 MHz, the clock at which the 128-byte port
matches the 2 DDR3 channels, however solely simply (WNS +0.032 ns). A event of Vivado runs retains
engaged on its margin and space. Decode is sure by DRAM, so a sooner clock largely helps prefill. - Faster prefill. The four-column systolic matrix unit is within the manufacturing picture; prefill is
nonetheless restricted by the matrix unit’s multiply charge.
Issues and pull requests are welcome, and many of the work wants solely Python and Verilator, not
an FPGA. Changes to the ISA, the simulator or the RTL should preserve python3 -m pytest -q passing,
and efficiency claims ought to say how they had been measured.
Apache License 2.0. See LICENSE.

