IN Brief:
- Teledyne SP Devices’ ADQ35 processes a 10 GSPS, 12 bit signal stream equivalent to 120 Gbit/s.
- An FPGA based convolutional neural network identifies individual particle arrivals when detector pulses overlap.
- The implementation delivers 300 ns latency while using less than 15% of DSP resources and 6% of LUT resources.
Teledyne SP Devices has demonstrated convolutional neural network inference directly on a 10 GSPS digitiser, using programmable logic to identify individual particle events when detector pulses overlap.
The implementation uses the company’s ADQ35 digitiser, which combines sampling rates up to 10 GSPS with an onboard AMD Kintex UltraScale KU115 FPGA. At 12 bits per sample and 10 GSPS, the incoming stream represents 120 Gbit/s of data, so moving the first stage of analysis onto the acquisition hardware reduces the amount of raw information that has to leave the card before an event can be identified.
The work was carried out through doctoral research at Fulda University of Applied Sciences with GSI Helmholtzzentrum für Schwerionenforschung in Darmstadt. It addresses particle pileup, where several particles pass through a detector so close together in time that their electrical responses merge.
A detector produces a pulse when a particle passes through the sensing element. At lower event rates, individual pulses remain sufficiently separated for threshold or local maximum methods to identify them, but those methods become less reliable as neighbouring pulses overlap and the combined waveform no longer resembles a sequence of isolated events.
The researchers trained a convolutional neural network to recover particle arrival positions from those merged waveforms and then implemented the model in FPGA fabric. Teledyne says the design processes 32 samples in parallel at 312.5 MHz, giving an effective throughput of 10 GSPS without routing the inference task through a host CPU or GPU.
Keeping the model on the acquisition hardware also fixes the inference path. Teledyne reports deterministic latency of 300 ns, allowing downstream processing or control to operate against a defined delay rather than waiting for data transfer, host scheduling and software execution to complete.
The network occupies less than 15% of the KU115’s available DSP resources and less than 6% of its lookup tables. Most of the programmable fabric therefore remains available for acquisition control, filtering, interfaces or additional processing, although any further functions still have to fit within the device’s timing, memory and routing limits.
FPGA inference is constrained by the hardware map as much as by the trained model. Arithmetic precision, network structure and parallelism have to be translated into finite DSP blocks, memories and logic while the design still closes timing at the required clock rate. A model that performs well during software training may need substantial restructuring before it can operate continuously beside a fast converter.
At 120 Gbit/s, continuously exporting every raw sample pushes the burden into host interfaces, memory and storage even when only a small proportion of the waveform contains information needed by the final application. Local inference allows particle timing and event information to be extracted before every sample has to cross that boundary.
Using the detector data and methods evaluated in the project, the CNN performed better under pileup conditions than the threshold and local maximum approaches tested alongside it. That comparison applies to the detector data and processing methods used in the project; it does not establish that neural processing will outperform conventional signal algorithms in every acquisition problem.
Other FPGA systems are also moving latency sensitive functions closer to the interface that produces the data. Orthogone has opened low latency networking IP for FPGA platforms, where protocol handling and data movement are similarly implemented in programmable logic to avoid unnecessary host processing.
Teledyne identifies RF analysis, image processing and other high rate acquisition tasks as potential areas for the same architecture, particularly where useful features have to be extracted before the raw stream can be stored or transported economically. Whether the approach transfers depends on the availability of representative training data and a model that can meet the resource and latency budget of the target FPGA.
The GSI implementation meets those constraints by matching 32 way parallel processing to a 312.5 MHz FPGA clock, producing effective 10 GSPS inference on the continuous detector stream. The result is a measurement architecture in which conversion and classification occur on the same hardware path instead of being separated by a host processing stage.


