Geniatech adds 20-TOPS edge AI modules

Geniatech adds 20-TOPS edge AI modules

Geniatech has launched two accelerator modules for offline edge inference. Both use M.2 hardware with different memory capacities for 3B- and 7B-class models.


IN Brief:

  • RK1828 and RK1820 cards each provide 20 TOPS INT8 acceleration in an M.2 2280 M-Key format.
  • RK1828 integrates 5GB of stacked DRAM for 7B-class models, while RK1820 provides 2.5GB for 3B-class workloads.
  • Both modules are available for evaluation and volume orders and connect to existing embedded hosts over PCIe.

Geniatech has launched RK1828 and RK1820 M.2 accelerator cards for adding local language-model and vision-language-model inference to embedded systems. Both modules use Rockchip AI co-processors rated at 20 TOPS INT8 and fit an M.2 2280 M-Key form factor, allowing additional inference capacity to be added without replacing the host processor or rebuilding the entire compute platform.

The principal difference is memory. RK1828 integrates 5GB of 3D-stacked in-package DRAM and is positioned for 7-billion-parameter language and multimodal models, while RK1820 provides 2.5GB for workloads around the 3-billion-parameter class. Both devices support mixed numerical formats including INT4, INT8, INT16, FP8, FP16, and BF16, allowing model developers to trade numerical precision against memory footprint and throughput.

The cards operate as dedicated co-processors rather than sharing the host’s main compute resources. Geniatech connects them over PCIe 2.0, separating the AI workload from a host CPU or SoC that may already be running control logic, video processing, networking, or user-interface functions. That separation does not eliminate transfer overhead, but it can reduce contention for host memory and processing resources when generative-AI functions are added to an existing embedded platform.

Memory architecture is particularly relevant for transformer inference because model weights and intermediate data have to be moved repeatedly during processing. Peak TOPS alone cannot describe performance if the arithmetic units are waiting for data, making local memory capacity and bandwidth significant design variables. Integrating DRAM alongside the co-processor gives the accelerator a dedicated working set rather than forcing every model access across the host memory interface.

Geniatech positions RK1828 for 7B-class models and RK1820 for 3B-class models, with both intended to run inference without a cloud connection. The larger device also supports multi-card configurations aimed at model capacities around 27B parameters. Expanding across several cards introduces additional power, cooling, interconnect, and model-partitioning requirements, but it provides a route to higher model capacity without immediately moving to a larger GPU-class system.

The modules are intended to work with Geniatech platforms based on Rockchip RK3588, RK3576, and RK3568 processors, while the software environment includes RKNN, TensorFlow, PyTorch, and ONNX support. For OEMs already using an M.2-capable embedded computer, that can reduce the hardware changes required to add local generative-AI functions, particularly where the existing host board has already passed mechanical, EMC, or thermal qualification.

Local inference can also be useful where operational data should remain inside the equipment. Industrial inspection images, maintenance records, machine status, or access-control information may be unsuitable for transmission to an external service, while some equipment has to continue operating when network connectivity is unavailable. A local accelerator can address those requirements, although model updates, security, and lifecycle maintenance still have to be handled by the OEM.

The M.2 format imposes practical constraints of its own. Designers have to account for card power, local heat dissipation, host PCIe bandwidth, mechanical clearance, supported operators, and the long-term availability of drivers and model-conversion tools. A 20-TOPS rating therefore remains only one part of the system specification; actual language-model performance depends on memory bandwidth, quantisation, context length, software optimisation, and the architecture of the model being deployed.

Geniatech says both cards are available for evaluation and volume orders, moving the announcement beyond an early development-board demonstration. The useful next comparison will be application-level measurements in representative embedded hosts, particularly where generative inference has to coexist with time-sensitive control or vision workloads. The hardware proposition is straightforward: retain the existing host architecture while moving increasingly memory-intensive local AI workloads onto a dedicated accelerator.


Stories for you


  • Geniatech adds 20-TOPS edge AI modules

    Geniatech adds 20-TOPS edge AI modules

    Geniatech has launched two accelerator modules for offline edge inference. Both use M.2 hardware with different memory capacities for 3B- and 7B-class models.


  • NanoBridge adopts Siemens FPGA synthesis flow

    NanoBridge adopts Siemens FPGA synthesis flow

    NanoBridge has adopted Siemens synthesis software for two FPGA families. The OEM agreement provides a device-tuned flow for low-power and harsh-environment programmable logic.