XELERA SILVA

Ultra-Low Latency
AI Inference Platform

Best-in-class throughput and latency for Gradient Boosting Trees and Neural Networks, from sub-microsecond CPU-only scoring to inline FPGA silicon. Three execution modes, one API.

decision treee acceleration software picture

Ultra-low latency AI Inference

Xelera Silva delivers ultra-low latency inference for Gradient-Boosted Tree models and Neural Networks across three execution modes under a single API.
CPU-Only mode reaches sub-microsecond p99 for small GBT models with no additional hardware, and runs Neural Networks in the low-microsecond range on Intel AMX.
Accelerator mode runs each model as an independent pipeline on an AMD Alveo FPGA card, holding latency flat as concurrent strategies scale.
Inline Silicon Mode embeds inference as an FPGA IP core directly in the network path for hard latency determinism in the low hundreds of nanoseconds. The trained model file is identical across all three modes. Switching modes is a single configuration-parameter change. The API call itself stays the same.

Technical features

CPU-Only
Mode

GBT Latency (p99):
0.8 to 7.7 µs

 
NN / LSTM latency (p99):
2.3 to 15.6 µs (Intel AMX)


Best for:
Models that fit in CPU cache, pinned one-per-core

Hardware:

None
(NN requires Intel AMX)


Data path:
Host CPU

Deployment:
Any Linux box: on-prem / colo / cloud

Operating System:
Linux Ubuntu, Rocky, CentOS

Accelerator FPGA Card

GBT Latency (p99):
1.1 to 1.3 µs


NN / LSTM latency (p99):
1.8 to 2.5 µs


Best for:
Larger models and multi-strategy deployments running in parallel

Hardware:
AMD Alveo U50 / U55C / V80; Napatech NT200A02


Data path:
PCIe to FPGA card

Deployment:

On-premise

Operating System:
Linux Ubuntu, Rocky, CentOS

Inline Silicon
(IP Core)

GBT Latency (p99):
~250 ns (LightGBM,
AMD UL3524)
 
NN / LSTM latency (p99):
~1.56 µs (LSTM, AMD Alveo V80)

Best for:
Strategies requiring hard latency determinism


Hardware:

AMD UL3524 or Alveo V80; custom FPGA integration

Data path:
On-chip, no PCIe transfer

Deployment:
On-premise

Operating System:
Linux Ubuntu, Rocky, CentOS

Algorithms by model class. GBT: XGBoost, LightGBM, CatBoost (numerical float32, categorical features). Neural networks: LSTM, linear layers, activations (sigmoid, tanh,relu), MLP, Conv1D, Single-Head Attention in float16 / bfloat16. Batch size 1 across all modes.

Platform note: GBT models run in CPU-Only mode on both AMD and Intel platforms. Neural-network models in CPU-Only mode require Intel AMX. Accelerator and Inline Silicon modes run both model classes on AMD Alveo hardware.
Latency and concurrency: p99 figures in the table are from the May 2026 benchmark kit (AMD Alveo V80, Ubuntu 22.04, isolated CPU cores, batchsize 1, p99 over 100k iterations). Under concurrent load, Accelerator p99 stays essentially flat: less than 2% increase from 1 to 8 concurrent models, with the kit citing flat performance up to 15 concurrent strategies. LSTM state is preserved across calls.

Your benefits

Unmatched Speed

Inference atency from sub-microsecond p99 (CPU-Only, small GBT, 0.8 µs measured) toroughly 250 ns (inline silicon), depending on mode and AI algorithm.

Seamless Integration

Neural networks and boosted tree algorithms under a unified high-performance software API in C/C++ and Python

Bring Your Own Model

Train your own model with the standard frameworks on your data and dynamically deploy on the accelerator card

Model Hot-Swap

Concurrent execution of multiple models on a single accelerator with instantaneous model hot-swapping

No Hardware Needed

Evaluate and deploy in CPU-Only mode with zero procurement lead time, then move to the FPGA card or inline silicon with a single configuration change when models or strategy count grow.

Use Cases

High-Frequency Trading
Accelerator (FPGA Card)

The picture shows the Use Case for High-Frequency Trading: Software Tick-to-Trade
The picture shows the Use Case for High-Frequency Trading: Hardware Tick-to-Trade

High-Frequency Trading
Inline Silicon (IP Core)

High-Frequency Trading
CPU-Only Mode

The picture shows the Use Case for High-Frequency Trading: Software Tick-to-Trade
Get OUR DatasheetS now!

Xelera Silva Datasheets

Software api integrationIP COreProduct brief

Deliverables

Xelera Silva is a turnkey full-stack solution designed to jumpstart best-in-class AI Inference acceleration.

Software packages

DEB / RPM packages for CPU-Only mode (no additional hardware required) and FPGA bitstreams for Accelerator and Inline Silicon modes on AMD AlveoU50, U55C, V80 and Napatech NT200A02.

API Support

Unified API in C, C++, C# and Python, identical across all three modes. Host library to load the model and run inference; runtime is thread-safe and process-safe for concurrent strategy deployments.

Example design

Jumpstart the AI inference acceleration with the provided example design

Support and User Guide

Integration and full lifecycle maintenance support
Periodic software updates

Getting Started

Pricing and Support

We understand that integrating solutions does not only require exceptional functionality but also transparent pricing models and reliable support. As technology evolves, so do we. We are committed to continuous innovation, ensuring that our software remains at the forefront of machine learning acceleration. With regular updates and feature enhancements, you can trust that you're always leveraging the latest advancements in the field. Our commitment to innovation means that you can stay ahead of the competition and unlock new possibilities for your projects. Contact us today to learn more about our pricing plans and support services. Unlock the full potential of your projects.

Latest Product news

OUR Technology partner
AMD Logo
STAC Member Logo
AMD Logo
STAC Member Logo