Adventures in benchmarking ONNX Runtime performance
I’ve recently had the opportunity to revisit ONNX file-format and ONNXRuntime for a project that I recently worked on. The idea of ONNX has always been very alluring: train a deep model in Python using PyTorch or TensorFlow, then convert and run it on a light-weight, fast inference stack. Often times, after training a deep model, we mostly care about deploying it for pure inference workloads. Using full-fat frameworks brings a lot of unnecessary overhead, especially if we are not interested in doing something fancy like online training. In my case particularly, where I am trying to do real-time control, I always enjoy figuring out how to squeeze every last bit of performance available on the device. I was even excited to see that it had first-class C# API, so in theory, if I wanted to do something like an ahead-of-time compiled C# cli tool that took care of ingesting data and running inference, I could do it.
I was excited about ONNX since I found out about it many years ago while I was doing my PhD, but I had never managed to get it working for the models I was exploring back in the day. Now to be fair, I was trying to experiment with very new ops introduced in TF/PT, including attention and friends, and those hadn’t yet been implemented in ONNX. The ones I was able to convert, including some basic CNNs and dense networks, I never saw significant gains on the system I was running them on. Benchmarking using ONNX runtime on Python usually put it within striking distance of just using the base framework itself, so it didn’t seem like there was any significant benefit. During the same period, I had also experimented with TFLite, now LiteRT, which was significantly faster (something on the order of 10-100x if my memory serves me right), so I shelved ONNX entirely. Since then, however, I had experienced many rough edges with TF, including dropped CUDA support on Windows and many bugs I encountered during daily use, that I decided to move to PyTorch fully.
For the Summer 2026 term, I did a side project with a summer undergrad student, which we eventually called QGrip, also on Hackster. It was a fun little project where we tried to deploy a PyTorch model on an Arduino UNO Q to control a robotic arm, the Handi Hand. I made the ML stack while the undergrad student worked on integrating the EMG sensor and the Handi Hand. I deployed my current go-to pipeline: a Transformer time-series classifier with an STFT preprocessing stage. I’ve tested this pipeline for years and feel it’s the best baseline for the task. As bonuses, I also implemented other classifiers, including CNNs and standard dense networks.
I vibe-coded the frontend with Codex and tested the stack on my desktop. Then we did a test on the UNO Q. The latencies were higher than I had hoped using vanilla PyTorch. Of course, I expected some performance drops from a desktop class x86 to the ARM processor on the board, but it was in the order of something like 20-30ms per decision, which was not acceptable for real-time inference. So I spent some time trying to find ways to cut down the latencies. Due to some misunderstanding on the availability of Qualcomm SDK on the platform, we ended up deciding to try ONNX. We thought that there was support for inference acceleration on the platform, but in the end, it turns out it was not available. This happy accident led us to explore ONNX a bit more.
So I modified the code to add ONNX conversion and backend support to my code. I took the opportunity to also do ONNX runtime benchmarking on my x86 desktop to get a sense of performance improvements. I found some interesting results, which I will detail more in a different blog post. For this one, the only key information required is that there are some cases where ONNX was slightly faster and some where it was slightly slower. Nothing that exceeds 2x, but not necessarily trivial either. So reasonably promising.
We went ahead and deployed it on the UNO Q and I was shocked to see that the performance was substantially better than running pure PyTorch on CPU. We are talking 5-10x. We then confirmed that QNN wasn’t the cause – it wasn’t available on the platform we were using, but just the pure CPU path seems to work quite well.
I was pleasantly surprised and happy with those findings - extra performance for free? Sign me up! I am also very curious to figure out why the discrepancy exists between an x86 desktop and the UNO Q. Is it purely just that desktop class processors have enough compute to deal with the extra overhead using vanilla PyTorch? Or is it something else? Perhaps the x86 build is linking to optimised BLAS or MKL libraries, while the ARM build is not? I will investigate it in more depth a bit later when I have more time. I also tried to convert the model to use it with the LiteRT runtime, but I had encountered many issues with even installing the required packages. When I finally got them installed, I had issues converting the models. I decided to set it aside for now.
We also came across an interesting performance bottleneck in ONNX runtime. While we were testing the different model architectures (CNNs, Transformers, etc.), we did some profiling using the runtime to figure out the distribution of ops that are taking up the most compute. Some usual suspects popped up - things like the Error Function from the GELU activation being a bottleneck, conservative pooling/downsampling parameters, etc. So we modified the code to solve some of those issues. Something really surprising popped up after everything was done – the largest contributor to the final latencies was surprisingly the STFT operator. This was unusual - DFT/STFT operators have historically been optimised very well. And what’s weird was that I wasn’t able to reproduce it on basic PyTorch, only on ONNX.
Intriguingly, the student set Claude loose on the problem, and it suggested we replace the operator with ones based on primitive convolution and multiplication ops. I was sceptical, but I figured there’s no harm in trying it. This version was orders of magnitude faster than the STFT operator while producing the same outputs. This had to be a bug of some sort in ONNX runtime.
I finally set some time aside to see if this was a proper bug worthy of filing a bug report. I could indeed confirm that it was reproducible, and the same issue crops up even on x86 desktop! I eventually opened a GH bug report here. My benchmarks showed that the basic STFT operator was anywhere from 200x-750x slower than equivalent operations expressed as convolution operators. There is currently an open PR for it from another contributor to address the issue, hopefully it gets merged soon.
The experience left me with two interesting takeaways: 1) ONNX Runtime can provide substantial inference gains on constrained edge hardware, and 2) the magnitude of those gains depends heavily on the operators used by the model and how efficiently they are implemented by the runtime. In this case, profiling uncovered an STFT implementation that was hundreds of times slower than an equivalent formulation using more primitive convolution operations. It was a good reminder that switching runtimes doesn’t automatically give us free performance; it is always worth profiling to check what the data tells us.