Chips&Media targets NPU memory traffic with lossless compression

Chips&Media targets NPU memory traffic with lossless compression

Chips&Media has developed lossless compression technology for neural activation data. The semiconductor IP developer reports memory traffic reductions of between 35.2% and 53.5% across tested numerical formats, with hardware implementation targeted for the first quarter of 2027.


IN Brief:

  • Feature-map compression targets external memory bandwidth during neural network inference.
  • Vendor testing reports traffic reductions of 35.2% for FP16, 45.7% for BF16 and 53.5% for INT8.
  • The compression algorithm and C-model are complete; hardware IP is planned for Q1 2027.

South Korean semiconductor IP developer Chips&Media has introduced a lossless feature-map compression architecture designed to reduce external memory traffic in neural processing systems. The company has completed its compression algorithm and associated C-model, with hardware implementation scheduled for completion during the first quarter of 2027. The development extends its intellectual property portfolio beyond video processing into data movement within AI accelerators.

Neural network layers generate intermediate activation values known as feature maps, which accelerators often write to and read from external DRAM. During inference, an accelerator may need to store these values temporarily and retrieve them for later operations, particularly when the working data set exceeds local memory capacity. Repeated transfers between processing hardware and external DRAM consume bandwidth and electrical power, while introducing delays when computing units cannot obtain data required for their next operation.

Chips&Media proposes inserting lossless encoding between a neural processor and external memory: feature-map values are compressed before the write and reconstructed exactly during the read. The design targets fewer transferred bits without intentionally changing the numerical data supplied to subsequent neural network layers.

Chips&Media reports bandwidth reductions of 35.2% for FP16, 45.7% for BF16 and 53.5% for INT8, based on testing across vision networks including ResNet, MobileNet and YOLO. These figures describe company results for selected workloads and numerical formats; they do not establish a fixed reduction for every model or application. Compression effectiveness depends on activation data characteristics, including how values are distributed and how effectively the algorithm represents them.

FP16 and BF16 both use 16-bit representations but allocate bits differently between numerical precision and exponent range; INT8 represents quantised values in eight bits. The resulting activation patterns can offer different opportunities for lossless encoding, which helps explain why Chips&Media reports different reductions for each format. A target workload must still be tested before estimating its likely memory bandwidth savings.

Chips&Media designed the compression hardware as an inline AXI bridge positioned between a neural processing unit and its external memory interface. AXI is a widely used on-chip interconnect protocol for transferring information between semiconductor functional blocks. Placing compression in this path allows the accelerator to issue memory transactions while additional logic transforms data passing towards or from the memory system.

The proposed inline AXI design shares the interconnect clock and is intended to avoid a separate internal SRAM buffer. Avoiding dedicated buffer memory may reduce silicon area and complexity, although a complete implementation still requires control logic, datapaths and timing validation. An inline compressor must sustain an appropriate processing rate so that its own latency or throughput limitations do not create a replacement bottleneck.

Reducing external traffic can help when a neural accelerator is constrained by available bandwidth between processing units and DRAM. If compute resources frequently stall while waiting for data, transmitting fewer bits can free bandwidth for useful work. Improvement in inference throughput depends on whether memory transfer is the limiting factor. A workload constrained by arithmetic execution or other resources may see a smaller benefit.

Fewer bits travelling between DRAM and an accelerator can reduce memory and interface activity, although compression and decompression consume power of their own. A net energy saving depends on the workload, transfer volume and implementation overhead, and cannot be inferred directly from the percentage reduction in memory traffic.

Chips&Media’s experience with video codecs and frame buffer compression provides a technical foundation, although feature maps have different characteristics from conventional image data. Its existing intellectual property includes hardware video encoders and decoders, frame buffer compression and specialised neural processing functions. The latest development addresses data exchanged during neural computation rather than storage or transmission of finished video frames.

With the algorithm and C-model completed, developers can evaluate compression behaviour in software, but no fully implemented hardware block has yet been confirmed as having completed RTL verification, silicon integration or physical qualification. The company targets hardware implementation completion in the first quarter of 2027, after which customers still need to assess the design in their own architectures.

Beyond feature maps, Chips&Media plans to investigate compression of model weights and key-value caches used in larger language and vision models. Those workloads contain different data structures and access patterns, so they remain prospective additions rather than established capabilities of the current feature-map design. Production silicon results and commercial adoption have not been disclosed; their eventual value will depend on validated hardware performance and customer workload characteristics.


Stories for you


  • Toray tests film capacitors at 150°C in inverter

    Toray tests film capacitors at 150°C in inverter

    Toray has tested heat-resistant film capacitors in a prototype inverter. Joint research with Nagoya University demonstrated operation at 150°C, while thermal modelling indicates that the technology could reduce the size of an inverter’s heat dissipation structure by approximately half.


  • Chips&Media targets NPU memory traffic with lossless compression

    Chips&Media targets NPU memory traffic with lossless compression

    Chips&Media has developed lossless compression technology for neural activation data. The semiconductor IP developer reports memory traffic reductions of between 35.2% and 53.5% across tested numerical formats, with hardware implementation targeted for the first quarter of 2027.