RaiderChip expands NPU compatibility beyond 50 AI models

RaiderChip expands NPU compatibility beyond 50 AI models

RaiderChip reports support for over fifty generative models across architectures. The Spanish semiconductor developer has added dense and mixture-of-experts Qwen models using firmware scheduling and dedicated transformer acceleration hardware without redesigning its NPU.


IN Brief:

  • RaiderChip reports compatibility with more than 50 generative models, including additional Qwen architectures.
  • Its firmware Model Scheduler works with hardware kernels to map transformer operations onto processing resources.
  • The resource utilisation and model-support figures are vendor-reported, not independent cross-model benchmarks.

Spanish semiconductor developer RaiderChip says its generative AI neural processing architecture now supports more than 50 models, adding newer Qwen language model families without changes to the underlying silicon design. The development uses firmware scheduling and dedicated hardware processing kernels to accommodate differences between dense transformer networks and mixture-of-experts architectures.

The latest reported compatible models include Qwen3.8-27B, a dense language model with 27 billion parameters, alongside Qwen3.5 and Qwen3.6 variants based on mixture-of-experts processing. RaiderChip says its architecture can execute additional models without requiring a new semiconductor implementation or the model preprocessing associated with some accelerator deployment workflows. The company is positioning that flexibility as a way to extend dedicated AI hardware as model architectures evolve.

Supporting both dense and mixture-of-experts transformers requires accelerator hardware to execute recurring matrix and attention operations while accommodating different memory access patterns. Transformer models rely heavily on matrix operations, attention mechanisms and movement of intermediate data between processing units and memory. A hardware architecture designed around these operations can reduce reliance on general-purpose instructions, although it must retain enough flexibility to accommodate changes in model dimensions, execution sequences and memory behaviour.

RaiderChip separates the model execution plan from the circuitry performing arithmetic operations. A firmware-based Model Scheduler develops a processing strategy for each supported architecture, while an Operation Scheduler implemented in hardware distributes kernels among the available processing resources. This division allows the allocation and order of transformer operations to change without replacing the compute fabric.

Dense transformers generally access the parameters used across a given inference pass, while a mixture-of-experts model routes each token through selected expert networks. The supported 35B-A3B variants contain approximately 35 billion parameters but activate around three billion per token, producing different scheduling and memory demands from a dense model of comparable total size.

Changing the execution plan allows an accelerator to respond to those patterns, with processing kernels scheduled according to the operations required by the active network structure. RaiderChip reports that its method can maintain hardware resource utilisation above 90% of the available physical limit, although this remains a company-reported architecture claim rather than a published independent benchmark covering every supported model. Practical throughput also depends on the chosen model, numerical precision, memory and system implementation.

RaiderChip’s accelerator IP for systems-on-chip and ASICs combines specialised transformer arithmetic with programmable scheduling. Its Kairós demonstrations have also shown language, speech and other models operating on local computing hardware, although deployment conditions and performance depend on the final system.

When speech, reasoning, visual interpretation and control networks run concurrently, they compete for memory bandwidth and processing capacity even if each model can execute separately. Scheduling must accommodate that contention within the electrical power and cooling limits of the embedded platform.

RaiderChip intends its architecture to support these combined workloads without requiring every model to be converted into a separate fixed hardware implementation. That does not remove the need to establish whether an individual model fits within a device’s memory capacity or meets application latency requirements. A 27 billion parameter model, for example, imposes substantial storage and data transfer demands even when the arithmetic operations themselves can be accelerated efficiently.

The company has previously demonstrated embedded generative AI using its Kairós platform and developed NPU intellectual property for integration into third-party semiconductor designs. The available evidence establishes claimed architectural compatibility, but does not identify final production configurations, commercial shipment volumes or independently measured performance for every new Qwen variant. System integrators must still evaluate performance under representative workloads and software conditions.

Compatibility with additional Qwen models expands the architectures that RaiderChip says can run on its accelerator without redesigning the silicon. The practical choice depends on how each model performs under the memory, power, latency and concurrency constraints of its intended electronic system, together with the resources available for continuing firmware development.


Stories for you


  • Applied Materials and Intel target advanced chip integration

    Applied Materials and Intel target advanced chip integration

    Applied Materials and Intel are expanding joint semiconductor process research. Their collaboration spans transistor structures, interconnect scaling and three-dimensional chip packaging, with development work taking place at facilities in California and Oregon.


  • GlobalFoundries to manufacture Xanadu quantum photonic components

    GlobalFoundries to manufacture Xanadu quantum photonic components

    Xanadu and GlobalFoundries will industrialise photonic quantum computing components together. Their agreement covers superconducting photon detectors and low-loss silicon nitride circuits, with process development and manufacturing planned on a 300 mm semiconductor production line in New York.