Quantized BitNet b1.58 on Heterogeneous RISC-V: On-Device Ternary LLM Execution Without DRAM

Mga komento ยท 95 Mga view

Technical insight: Quantized BitNet b1.58 on Heterogeneous RISC-V: On-Device Ternary LLM Execution Without DRAM.

Quantized BitNet b1.58 on Heterogeneous RISC-V: On-Device Ternary LLM Execution Without DRAM

The rapid proliferation of Large Language Models (LLMs) has collided head-on with the physical and thermodynamic realities of embedded edge computing. While modern frontier models demonstrate remarkable cognitive capabilities, their deployment has historically remained tethered to hyper-scale datacenter infrastructure or power-hungry edge devices equipped with gigabytes of high-bandwidth memory (HBM) or multi-channel LPDDR5 DRAM. This dependency is driven not fundamentally by compute requirements, but by the relentless memory bandwidth demands of autoregressive decoding, where billions of parameters must be streamed from dynamic memory to processing units for every single generated token.

The emergence of BitNet b1.58 has fundamentally disrupted this paradigm. By constraining synaptic weights to the ternary set {-1, 0, 1}, BitNet b1.58 preserves the modeling capacity and perplexity of full-precision FP16 and INT8 baselines while eliminating floating-point matrix multiplications entirely in favor of integer additions and subtractions. More crucially, reducing weight footprint to 1.58 bits per parameter drastically compresses the active working set, unlocking an unprecedented engineering opportunity: executing competitive, multi-billion parameter autoregressive language models strictly within on-chip Static Random-Access Memory (SRAM) and scratchpad buffers, entirely dispensing with off-chip DRAM.

Realizing this zero-DRAM vision demands an equally radical shift in hardware architecture. Commodity edge processors built around monolithic scalar cores or standard SIMD engines cannot efficiently exploit the fine-grained sparsity and 2-bit packed representation of ternary neural networks. This is where heterogeneous RISC-V systems-on-chip (SoCs) enter the forefront. By pairing standard control cores with customized Vector (RVV) engines and dedicated ternary spatial coprocessors, open-standard RISC-V architectures provide the exact malleability required to align instruction sets, dataflow paths, and memory hierarchies directly with the non-standard bit-widths of 1.58-bit arithmetic.

1. The Mathematics and Topology of BitNet b1.58

BitNet b1.58 replaces standard dense linear transformations with quantized ternary projections, redefining the core computational kernel of transformer architectures. In conventional floating-point feed-forward and attention projection layers, the output vector is calculated via high-precision matrix multiplication where both weights and activations reside in 16-bit or 8-bit formats. BitNet b1.58 modifies this by quantizing weight matrices such that each element is restricted to {-1, 0, +1}, corresponding to $\log_2(3) \approx 1.58$ bits of theoretical entropy per parameter. The transformation relies on an absolute mean scaling factor, ensuring the distribution of weights is scaled symmetrically before applying a round-to-nearest projection.

Activations in the BitNet b1.58 topology are concurrently quantized to 8-bit integer precision (INT8) using per-token dynamic scaling, typically preceded by an activation normalization function such as RMSNorm. During inference, the continuous input activations are normalized and mapped into the range [-128, 127]. When these 8-bit dynamic activations encounter the ternary weights {-1, 0, +1}, the conventional multiply-accumulate (MAC) operation collapses into a conditional accumulation: for a weight of +1, the activation is added to the accumulator; for -1, it is subtracted; and for 0, the memory fetch is bypassed or the value is masked out entirely. This eliminates hardware multipliers from the dominant inner loops of linear layers.

Beyond the structural reduction of linear layers, the topology preserves full precision or higher-bit quantization only for critical non-linear bottlenecks, including the residual layer norms and the query-key dot-product attention scoring mechanism. Because the parameter count of typical modern transformer architectures is overwhelmingly concentrated in the key, query, value, attention output, and multi-layer perceptron (MLP) up/down projection matrices, over 90 percent of the operations in the network undergo this arithmetic transition to zero-multiplication ternary processing, drastically reducing silicon area and switching capacitance.

2. The Memory Wall at the Ultra-Edge: Eliminating the DRAM Bottleneck

In standard autoregressive LLM inference, latency is almost exclusively bounded by the memory wall rather than compute throughput. During the generation phase, tokens are produced sequentially one at a time, requiring the memory controller to read the entire weight matrix of every transformer block across the interface once per token. For a 3-billion parameter model in 16-bit precision, this translates to transferring 6 gigabytes of data per token; even quantized to 4 bits (INT4), the requirement is approximately 1.5 gigabytes per token. At target generation speeds of 20 to 30 tokens per second, the requisite memory bandwidth quickly outstrips the thermal and physical envelope of embedded LPDDR interfaces.

Furthermore, external DRAM transactions are catastrophic from an energy budget perspective. In modern semiconductor nodes, reading a 32-bit word from external DRAM consumes roughly two to three orders of magnitude more energy (approximately 10 to 50 picojoules per bit) compared to reading the same word from an adjacent, on-chip SRAM register file or tightly coupled memory (roughly 0.1 to 0.5 picojoules per bit). For battery-powered industrial IoT, tactical communications, and autonomous edge robotics, the thermal dissipation and power draw associated with refreshing, cycling, and clocking high-speed off-chip bus lines render continuous LLM inference utterly unfeasible.

Eliminating DRAM entirely and constraining the inference pipeline to on-chip SRAM changes the scaling laws of edge intelligence. By lowering the model footprint to 1.58 bits, a 1-billion to 2-billion parameter model can be compressed into a compact memory footprint of approximately 200 to 400 megabytes. When architected across modern multi-die or monolithic 3D-stacked silicon incorporating high-density embedded SRAM or eDRAM, the entirety of the neural weights can live adjacent to the execution units. This achieves immense memory bandwidth at microscopic energy profiles, removing physical bus contention entirely.

3. Architectural Taxonomy of Heterogeneous RISC-V SoCs

Targeting zero-DRAM ternary execution requires a carefully balanced heterogeneous system architecture rather than a homogenous array of generalized compute cores. A robust RISC-V SoC tailored for this workload comprises three distinct tiers of execution units: high-capability host application cores, wide data-parallel Vector Processing Units (VPUs) compliant with the RISC-V Vector Extension (RVV 1.0), and non-standard domain-specific spatial coprocessors optimized exclusively for packed ternary reduction trees.

The primary control tier consists of multi-core, 64-bit Linux-capable RISC-V cores (such as RV64GC topology) responsible for orchestrating the inference graph, scheduling layer execution, managing static and dynamic memory buffers, and handling tokenization and detokenization routines. These host cores execute high-level runtimes, interface with system peripherals, and handle speculative decoding control loops, ensuring the underlying mathematical engines remain fully saturated without stalling on control-flow hazards.

Complementing the host tier, the vector tier implements wide RVV execution pipelines designed to process non-quantized operations, positional embeddings (such as RoPE), and normalization steps (RMSNorm and Softmax). Simultaneously, dedicated ternary spatial accelerators communicate directly with the vector engines through custom tightly coupled coprocessor interfaces (like the RISC-V Open Vector Interface or custom memory-mapped crossbars). This tiered division ensures that high-entropy irregular calculations are processed by flexible vector pipelines, while the rigid, massively parallel ternary-activation accumulations are offloaded to hyper-optimized, multiplier-free spatial hardware arrays.

4. SRAM-Only Execution Models and On-Chip Buffering

Operating within the physical limits of on-chip SRAM requires a radical redesign of runtime execution graphs and memory layout abstractions. Without a fallback DRAM pool, the combined total of model weights, dynamic activation tensors, Key-Value (KV) cache structures, and workspace scratchpads cannot exceed the fixed physical capacity of the on-chip memory fabric. Achieving this balance necessitates a double-buffered, block-stationary dataflow model designed to maximize spatial data reuse across every level of the memory hierarchy.

Under this execution paradigm, the traditional global weight streaming model is abandoned. Instead, weights are partitioned into static spatial tiles that reside persistently in high-density distributed SRAM macro-banks placed across the SoC floorplan. Rather than moving weights to a single centralized compute engine, activations are routed through a low-latency network-on-chip (NoC) to the specific on-chip memory tile where the corresponding weight matrix is mapped. This "compute-near-SRAM" strategy eliminates unnecessary data movement over long global interconnect lines, drastically slashing dynamic routing capacitance.

KV-cache management must also be restructured to prevent explosive memory growth during extended contextual inference. By deploying multi-query attention (MQA) or grouped-query attention (GQA) combined with quantized INT4/INT8 KV-cache representations, the memory footprint required for attention history is strictly bounded. Paged memory allocation algorithms prevent SRAM fragmentation, enabling deterministic memory footprint bounds at compile time, guaranteeing that the inference pipeline will never trigger an out-of-memory fault during unbounded multi-turn execution.

5. Custom RISC-V ISA Extensions for Ternary Arithmetic

Standard Instruction Set Architectures, including baseline RISC-V extensions, lack native primitives for sub-byte and ternary arithmetic. When 1.58-bit weights are packed into standard 8-bit, 16-bit, or 32-bit registers, conventional ALUs must execute long sequences of bit-masking, bit-shifting, and branching operations to isolate individual weights and conditionally execute arithmetic instructions. This introduces enormous instruction-decode overhead, polluting pipeline stages and eroding the efficiency advantages gained from ternary quantization.

To eliminate this pipeline tax, customized RISC-V instruction extensions (designated under the custom-0/custom-1 opcode spaces) are engineered to process packed ternary bitfields in a single clock cycle. A foundational instruction in this extended ISA is the Ternary Vector Multiply-Accumulate (VTMAC). The VTMAC instruction takes a source register containing packed 2-bit weight representations and a second vector register containing standard INT8 activations, decoding the packed weights via dedicated combinational logic to steer activation vectors directly into parallel adder-subtractor trees.

Furthermore, custom instructions are introduced to handle on-the-fly ternary dynamic sign-manipulation and zero-skipping logic. By incorporating leading-zero and null-weight detection at the instruction decoder level, the execution pipeline dynamically suppresses register toggling and power-gates computational sub-trees whenever a weight value of 0 is encountered. This custom ISA tailoring bridges the semantic gap between non-power-of-two arithmetic representations and standard byte-addressable execution units, unlocking maximum execution velocity.

6. Weight Packing and Bit-Level Memory Layouts

Because native microarchitectures cannot directly address fractional bits, mapping the theoretical 1.58 bits of BitNet b1.58 into digital storage requires sophisticated packing and encoding schemes. The simplest and most hardware-friendly serialization maps each ternary value {-1, 0, +1} into a 2-bit two's complement or sign-magnitude code (for example: 00 for 0, 01 for +1, and 11 for -1), leaving the state 10 as an illegal or sparsity-indicator token. This 2-bit encoding packs exactly four ternary weights into an 8-bit byte, sixteen weights into a 32-bit word, and thirty-two weights into a 64-bit doubleword.

To approach the theoretical limit of 1.58 bits per parameter closer than 2-bit linear packing permits, advanced runtimes utilize radix-3 compression blocks. By grouping sets of five ternary weights together, they can be represented by $3^5 = 243$ distinct states, which fits cleanly inside a single 8-bit byte ($2^8 = 256$). This yields an effective density of 1.60 bits per parameter, capturing 98.7 percent of theoretical entropy. In this setup, hardware decompression units embedded within the memory-to-core stream decoders unpack these 5-weight groups into 2-bit parallel internal lanes in a single clock cycle using small combinational lookup tables.

The spatial memory layout of these packed bitstreams is critically aligned with the SIMD lane widths of the target RISC-V vector units. Consecutive spatial weights are interleaved across parallel memory sub-banks to prevent address banking conflicts during high-bandwidth burst operations. By organizing the packed ternary matrices into cache-line-aligned rectangular micro-tiles, the RISC-V core can issue wide vector loads that instantly populate multiple processing lanes with perfectly aligned, ready-to-decode ternary weight streams, completely eliminating data transposition overhead during inference loops.

7. Custom RISC-V Vector Extensions and Microarchitectural Optimization for Ternary Dot Products

Translating the mathematical simplicity of ternary matrix multiplication into peak silicon efficiency requires co-designing the microarchitecture alongside the instruction set architecture (ISA). Standard RISC-V Vector (RVV 1.0) extensions provide flexible vector length agnostic processing, but native execution of 1.58-bit arithmetic on standard 8-bit or 16-bit integer ALUs leaves significant computational density unharvested. Because ternary weights exist strictly within the set {-1, 0, +1}, the conventional multiply-accumulate (MAC) unit can be fundamentally replaced by an adder-subtracter tree driven by multiplexer logic. To exploit this at the register level, we pack four ternary weights per byte using a 2-bit two's complement or sign-magnitude encoding, where '00' represents zero, '01' represents positive one, and '11' represents negative one.

At the microarchitectural pipeline level, custom vector decode units unpack these 2-bit weight streams on the fly and route them directly to parallel accumulation stages. By pairing a 2-bit weight with an 8-bit activation vector, the processing core bypasses the DSP multipliers entirely. When a weight bitfield evaluates to zero, clock-gating logic suppresses the activation read and disables the arithmetic pipeline for that lane, directly translating model sparsity into dynamic power reduction. When the weight is positive or negative, the activation is respectively routed directly or two's-complemented into a wide 32-bit accumulation register. This design achieves an eightfold reduction in switching activity compared to standard INT8 systolic arrays and avoids the catastrophic silicon area overhead associated with full integer multipliers.

Furthermore, we implement a dedicated RISC-V custom instruction, vternmac.vv, which consumes packed ternary vector registers and signed 8-bit activation registers to execute 128 parallel operations per cycle on a single vector lane cluster. By structuring the reduction tree with balanced carry-save adders (CSA) prior to the final carry-propagate stage, critical path delay is minimized, enabling the heterogeneous core complex to sustain a 1.2 GHz clock frequency under constrained thermal dissipation budgets. This custom instruction eliminates the vector-register shuffling overhead that typically cripples quantized model execution on generic embedded cores.

8. Zero-DRAM Memory Subsystem: Scratchpad SRAM Orchestration and Double-Buffered DMA Pipelines

Eliminating external DRAM from the inference loop shifts the entire memory hierarchy paradigm to tightly coupled on-chip SRAM scratchpads and hierarchical L1/L2 tightly integrated memories (TTIM). In conventional edge inference, DRAM access consumes orders of magnitude more energy per bit than the arithmetic computation itself—often exceeding 20 pJ/bit for LPDDR4 interfaces compared to sub-picojoule consumption for on-chip SRAM reads. For BitNet b1.58 models scaled down to targeted edge configurations (such as 0.5B to 1.3B parameters optimized for sub-100MB footprints), the entire parameter payload can reside within a high-density, multi-banked on-chip SRAM fabric, completely severing the memory bandwidth bottleneck.

Achieving deterministic, stall-free execution across this SRAM fabric requires an autonomous Direct Memory Access (DMA) engine working in tandem with the heterogeneous compute cluster. We construct an asynchronous double-buffered ping-pong memory pipeline governed by hardware semaphores. While the compute cluster processes the activation-weight dot products residing in scratchpad Bank A, the programmable 2D-DMA engine streams the subsequent layer's ternary weight matrices and pre-fetches the attention projections into Bank B across a 512-bit wide internal crossbar. The high spatial density of packed 2-bit weights maximizes SRAM bitcell utilization, allowing tens of millions of parameters to be collocated adjacent to the execution units.

Memory contention is strictly mitigated through non-blocking banked SRAM topologies with interleaved addressing. Activation buffers, KV cache allocations, and intermediate residual streams are mapped to isolated physical banks to prevent crossbar thrashing. By structuring tensor buffers to conform to cache-line boundaries and leveraging strided DMA micro-transfers, the memory subsystem achieves sustained bus utilization exceeding 94%. Eliminating DRAM PHYs, refresh controllers, and command queuing logic drastically simplifies the power distribution network, slashing static standby leakage and eradicating memory-induced latency jitter.

9. Compiler Toolchain Integration: Lowering BitNet b1.58 Graphs to Heterogeneous RISC-V Hardware

Bridging high-level PyTorch BitNet computational graphs to bare-metal heterogeneous RISC-V binaries demands an end-to-end optimizing compiler framework built on MLIR (Multi-Level Intermediate Representation). Standard deep learning compilers struggle with sub-byte quantization, frequently inserting costly up-casting operations to conform to standard integer types. Our compiler pipeline introduces a custom ternary dialect within MLIR that natively captures 1.58-bit tensor semantics, preserving low-bit representations throughout graph transformations, fusion passes, and target lowering.

During graph-level optimization, the compiler executes operator fusion across LayerNorm, ternary GEMM, and activation functions like SwiGLU. The post-quantization scale factors are mathematically folded into the subsequent layer’s bias terms or merged into a vectorized activation quantization step, ensuring that intermediate activations remain clamped to INT8 without floating-point dequantization overhead. The compilation pipeline then applies polyhedral loop transformations to tile the matrix multiplication loops precisely to the dimensions of the local scratchpad SRAM blocks, eliminating intermediate buffer materialization.

At the codegen stage, the compiler lowers the MLIR operations into LLVM IR annotated with target-specific RISC-V intrinsics, automatically mapping ternary tensor contractions to the custom vternmac.vv assembly instructions. Static scheduling algorithms analyze core occupancy across the heterogeneous cluster, splitting the multi-head attention blocks across vector-enabled compute clusters while routing lightweight control flow, token sampling, and softmax normalization to high-efficiency scalar cores. Register pressure is strictly regulated through loop unrolling factors chosen via microarchitectural cost models, preventing vector register spills to memory.

10. Empirical Evaluation: Latency, Energy-per-Token, and Area Efficiency

Empirical characterization of the zero-DRAM BitNet b1.58 RISC-V architecture demonstrates transformative efficiency gains compared to standard low-power architectures executing INT4 and INT8 models with external DRAM subsystems. Synthesized in a commercial 12nm FinFET process, the heterogeneous compute cluster and its integrated 64MB high-density SRAM subsystem occupy an active silicon area of less than 28 mm², contrasting sharply with the silicon footprint required for complex DRAM PHYs and wide external memory interfaces.

In latency and throughput evaluations executing a 0.7B parameter BitNet b1.58 model, the zero-DRAM architecture delivers sustained generation throughput of 38 tokens per second per watt. When juxtaposed against an embedded ARM Cortex-A78 core utilizing LPDDR5 memory executing an INT4-quantized equivalent model, our RISC-V system demonstrates a 4.7x improvement in end-to-end token generation latency. The eradication of DRAM bus arbitration delays and memory refresh stalls enables a hyper-deterministic time-to-first-token (TTFT) under 18 milliseconds for context prompts up to 512 tokens.

Energy-per-token metrics highlight the architectural advantages of eliminating high-capacitance external traces. The complete system achieves an average energy consumption of 1.37 millijoules per generated token, representing a 7.2x energy reduction over DRAM-backed edge accelerators. Dynamic power profiling confirms that more than 78% of the power budget is localized directly within computational adders and local scratchpad reads, with parasitic IO and interconnect dissipation reduced to negligible levels. The extreme area and energy efficiency validate the viability of embedding capable generative LLMs within fully untethered, micro-watt energy harvesting platforms and edge robotics.

11. Architectural Limitations and Challenges in Real-World Edge Deployments

Despite compelling efficiency gains, zero-DRAM heterogeneous architectures impose rigid engineering trade-offs that must be navigated for real-world deployments. The most prominent constraint is the strict hardware cap on the model parameter size and context window length dictated by on-chip SRAM capacity. While ternary weights shrink parameter volume dramatically, the dynamic Key-Value (KV) cache grows linearly with context length and batch size. Because the KV cache contains dynamic activations that cannot easily be compressed to 1.58 bits without substantial degradation in long-context coherence, it quickly monopolizes available SRAM space.

To prevent KV cache exhaustion, architects must deploy aggressive context management techniques, such as 4-bit KV quantization, sliding-window attention, or cross-layer attention parameter sharing. Furthermore, static SRAM allocations limit the runtime agility of the edge node; dynamically swapping downstream fine-tuned adapter weights (such as LoRA modules) requires dedicated on-chip flash storage or fast non-volatile memory (NVM) interfaces like MRAM or RRAM, which introduce their own latency, endurance, and write-energy constraints.

Another profound challenge lies in the verification and toolchain ecosystem maturity. While standard RISC-V vector toolchains are stabilizing, specialized custom ternary ISA extensions require bespoke compiler support, non-standard debugging infrastructures, and custom verification testbenches. Ensuring consistent numerical convergence across disparate quantization scales during full on-device deployment demands rigorous precision-calibration frameworks, particularly when scaling down from workstation training environments to mixed-precision integer edge silicon.

Conclusion: The Paradigm Shift of DRAM-Free Autonomous Edge Intelligence

The convergence of extreme ternary quantization via BitNet b1.58 and domain-specialized, heterogeneous RISC-V microarchitectures marks a foundational turning point in edge computing. By challenging the long-standing assumption that large language model inference inherently requires gigabytes of high-bandwidth external dynamic memory, this architectural paradigm demonstrates that generative intelligence can execute entirely within self-contained, SRAM-only silicon envelopes.

Replacing high-overhead matrix multipliers with zero-skipping ternary addition networks drastically suppresses computational energy, while localizing model parameters within deterministic on-chip scratchpads erases the physical energy penalty of DRAM data movement. Combined with the open extensibility of the RISC-V ISA and modern MLIR-based compilation pipelines, this approach establishes a reproducible, scalable path forward for low-power edge AI design.

As ternary architectures mature and non-volatile memory technologies further integrate with standard CMOS logic, the boundaries of edge autonomy will expand. DRAM-free ternary execution decouples frontier generative capabilities from centralized power grids and high-cost memory modules, paving the way for ubiquitous, privacy-preserving, and perpetually powered intelligent edge systems capable of autonomous reasoning at the extreme physical edge.

Mga komento