Model zoo performance
Akida 1.0 models for models targeting the Akida Neuromorphic Processor IP 1.0 and the AKD1000 reference SoC,
Akida 2.0 models for models targeting the Akida Neuromorphic Processor IP 2.0,
Akida Pico models for models targeting the Akida Pico Neuromorphic Processor IP,
Upgrading to Akida 2.0 tutorial to understand the architectural differences between 1.0 and 2.0 models and their respective workflows.
Note
The download links provided point towards standard TensorFlow Keras models that must be converted to an Akida model using cnn2snn.convert.
Akida 1.0 models
For 1.0 models, 4-bit accuracy is provided and is always obtained through a QAT phase.
Note
The “8/4/4” quantization scheme stands for 8-bit weights in the input layer, 4-bit weights in other layers and 4-bit activations.
The NPs column provides the minimal number of neural processors required for the model execution on the Akida IP. The numbers given are the result of the map operation using the Minimal MapMode targeting AKD1000/AKD1500 SoC.
Energy per inference is an average, measured on an AKD1500 device for the most efficient mapping.
Image domain
Classification
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
Top-1 accuracy |
Size (KB) |
NPs |
Energy (mJ) |
Download |
|---|---|---|---|---|---|---|---|---|---|
AkidaNet 0.25 |
160 |
ImageNet |
480K |
8/4/4 |
42.58% |
403.3 |
20 |
1.56 |
|
AkidaNet 0.5 |
160 |
ImageNet |
1.4M |
8/4/4 |
57.80% |
1089.1 |
24 |
4.32 |
|
AkidaNet |
160 |
ImageNet |
4.4M |
8/4/4 |
66.94% |
4061.1 |
68 |
13.66 |
|
AkidaNet 0.25 |
224 |
ImageNet |
480K |
8/4/4 |
46.71% |
409.1 |
22 |
2.50 |
|
AkidaNet 0.5 |
224 |
ImageNet |
1.4M |
8/4/4 |
61.30% |
1202.2 |
32 |
6.885 |
|
AkidaNet |
224 |
ImageNet |
4.4M |
8/4/4 |
69.65% |
6294.0 |
116 |
26.34 |
|
AkidaNet 0.5 edge |
160 |
ImageNet |
4.0M |
8/4/4 |
51.66% |
2017.4 |
38 |
6.63 |
|
AkidaNet 0.5 edge |
224 |
ImageNet |
4.0M |
8/4/4 |
54.03% |
2130.5 |
46 |
9.99 |
|
AkidaNet 0.5 |
224 |
PlantVillage |
1.1M |
8/4/4 |
97.92% |
1019.1 |
33 |
7.17 |
|
AkidaNet 0.25 |
96 |
Visual Wake Words |
229K |
8/4/4 |
84.77% |
179.6 |
16 |
0.49 |
|
MobileNetV1 0.25 |
160 |
ImageNet |
467K |
8/4/4 |
36.05% |
376.4 |
20 |
1.53 |
|
MobileNetV1 0.5 |
160 |
ImageNet |
1.3M |
8/4/4 |
54.59% |
1007.0 |
24 |
4.21 |
|
MobileNetV1 |
160 |
ImageNet |
4.2M |
8/4/4 |
65.47% |
3525.8 |
65 |
13.44 |
|
MobileNetV1 0.25 |
224 |
ImageNet |
467K |
8/4/4 |
39.73% |
377.9 |
22 |
2.46 |
|
MobileNetV1 0.5 |
224 |
ImageNet |
1.3M |
8/4/4 |
58.50% |
1065.3 |
32 |
6.68 |
|
MobileNetV1 |
224 |
ImageNet |
4.2M |
8/4/4 |
68.76% |
5223.3 |
110 |
26.28 |
|
GXNOR |
28 |
MNIST |
1.6M |
2/2/1 |
98.03% |
412.8 |
3 |
0.34 |
Object detection
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
mAP |
Size (KB) |
NPs |
Energy (mJ) |
Download |
|---|---|---|---|---|---|---|---|---|---|
YOLOv2 |
224 |
PASCAL-VOC 2007 - person and car classes |
3.6M |
8/4/4 |
41.51% |
3061.4 |
71 |
14.61 |
|
YOLOv2 |
224 |
WIDER FACE |
3.5M |
8/4/4 |
77.63% |
3053.1 |
71 |
14.22 |
Regression
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
MAE |
Size (KB) |
NPs |
Energy (mJ) |
Download |
|---|---|---|---|---|---|---|---|---|---|
VGG-like |
32 |
UTKFace (age estimation) |
458K |
8/2/2 |
6.1791 |
138.6 |
6 |
0.14 |
Face recognition
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
Accuracy |
Size (KB) |
NPs |
Energy (mJ) |
Download |
|---|---|---|---|---|---|---|---|---|---|
AkidaNet 0.5 |
112×96 |
CASIA Webface face identification |
2.3M |
8/4/4 |
70.18% |
1930.1 |
21 |
3.73 |
|
AkidaNet 0.5 edge |
112×96 |
CASIA Webface face identification |
23.6M |
8/4/4 |
71.13% |
6980.2 |
34 |
8.41 |
Audio domain
Keyword spotting
Architecture |
Dataset |
#Params |
Quantization |
Top-1 accuracy |
Size (KB) |
NPs |
Energy (mJ) |
Download |
|---|---|---|---|---|---|---|---|---|
DS-CNN |
Google Speech Commands |
22.7K |
8/4/4 |
91.72% |
23.1 |
5 |
0.07 |
Point cloud
Classification
Architecture |
Dataset |
#Params |
Quantization |
Accuracy |
Size (KB) |
NPs |
Download |
|---|---|---|---|---|---|---|---|
PointNet++ |
ModelNet40 3D Point Cloud |
602K |
8/4/4 |
79.78% |
490.9 |
12 |
Akida 2.0 models
For 2.0 models, both 8-bit PTQ and 4-bit QAT numbers are given. When not explicitly stated, 8-bit PTQ accuracy is given as is (i.e. no further tuning/training, only quantization and calibration). The 4-bit QAT is the same as for 1.0.
Note
The digit in the quantization scheme stands for both the weights and activations bitwidth. Weights in the first layer are always quantized to 8-bit.
The NPs column provides the minimal number of neural processors required for the model execution on the Akida IP. The numbers given are the result of the map operation using the Minimal MapMode targeting a 12-node Akida 2.0 device.
Image domain
Classification
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
Accuracy |
NPs |
Download |
|---|---|---|---|---|---|---|---|
AkidaNet 0.25 |
160 |
ImageNet |
483K |
8 4 |
48.61% 40.69% |
27 26 |
|
AkidaNet 0.5 |
160 |
ImageNet |
1.4M |
8 4 |
61.92% 57.42% |
42 29 |
|
AkidaNet |
160 |
ImageNet |
4.4M |
8 4 |
69.96% 66.80% |
124 60 |
|
AkidaNet 0.25 |
224 |
ImageNet |
483K |
8 4 |
52.38% 44.48% |
31 26 |
|
AkidaNet 0.5 |
224 |
ImageNet |
1.4M |
8 4 |
64.85% 60.53% |
55 34 |
|
AkidaNet |
224 |
ImageNet |
4.4M |
8 4 |
72.23% 69.21% |
206 92 |
|
AkidaNet 0.5 |
224 |
PlantVillage |
1.2M |
8 4 |
99.61% 99.30% |
56 35 |
|
AkidaNet 0.25 |
96 |
Visual Wake Words |
227K |
8 4 |
87.05% 85.70% |
25 24 |
|
AkidaNet18 |
160 |
ImageNet |
2.4M |
8 |
64.72% |
61 |
|
AkidaNet18 |
224 |
ImageNet |
2.4M |
8 |
67.32% |
86 |
|
MobileNetV1 0.25 |
160 |
ImageNet |
469K |
8 4 |
45.72% 36.96% |
31 30 |
|
MobileNetV1 0.5 |
160 |
ImageNet |
1.3M |
8 4 |
60.16% 54.09% |
47 34 |
|
MobileNetV1 |
160 |
ImageNet |
4.2M |
8 4 |
69.04% 64.92% |
114 68 |
|
MobileNetV1 0.25 |
224 |
ImageNet |
469K |
8 4 |
49.58% 40.80% |
36 31 |
|
MobileNetV1 0.5 |
224 |
ImageNet |
1.3M |
8 4 |
63.67% 57.87% |
65 44 |
|
MobileNetV1 |
224 |
ImageNet |
4.2M |
8 4 |
71.31% 67.72% |
184 106 |
|
GXNOR |
28 |
MNIST |
1.6M |
4 |
98.57% |
4 |
Object detection
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
mAP 50 |
NPs |
Download |
|---|---|---|---|---|---|---|---|
YOLOv2 (AkidaNet 0.5 backbone) |
224 |
PASCAL-VOC 2007 |
3.6M |
8 4 |
51.41% 46.74% |
119 70 |
|
CenterNet (AkidaNet18 backbone) |
384 |
PASCAL-VOC 2007 |
2.4M |
8 |
72.77% [1] |
336 |
|
CenterNet (AkidaNet18 backbone) |
224 |
PASCAL-VOC 2007 |
2.4M |
8 |
66.08% [1] |
125 |
|
YOLOv2 (AkidaNet 0.5 backbone) |
224 |
WIDER FACE |
3.6M |
8 4 |
80.51% 78.69% |
117 69 |
Regression
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
MAE |
NPs |
Download |
|---|---|---|---|---|---|---|---|
VGG-like |
32 |
UTKFace (age estimation) |
458K |
8 4 |
6.0299 6.1421 |
7 6 |
Face recognition
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
Accuracy |
NPs |
Download |
|---|---|---|---|---|---|---|---|
AkidaNet 0.5 |
112×96 |
CASIA Webface face identification |
2.3M |
8 4 |
73.02% 68.60% |
40 29 |
Segmentation
Architecture |
Resolution |
Dataset |
#Params |
Quantization |
Binary IOU |
NPs |
Download |
|---|---|---|---|---|---|---|---|
AkidaUNet 0.5 |
128 |
Portrait128 |
1.1M |
8 |
0.9076 [2] |
66 |
PTQ accuracy boosted with 1 epoch QAT.
Audio domain
Keyword spotting
Architecture |
Dataset |
#Params |
Quantization |
Top-1 accuracy |
NPs |
Download |
|---|---|---|---|---|---|---|
DS-CNN |
Google Speech Commands |
23.8K |
8 4 |
92.83% 92.58% |
9 9 |
Point cloud
Classification
Architecture |
Dataset |
#Params |
Quantization |
Accuracy |
NPs |
Download |
|---|---|---|---|---|---|---|
PointNet++ |
ModelNet40 3D Point Cloud |
277K |
8 4 |
79.62% [3] 79.50% |
13 11 |
TENNs
Gesture recognition
Dataset |
#Params |
Quantization |
Accuracy |
NPs |
Download |
|---|---|---|---|---|---|
DVS128 |
165K |
8 |
97.12% |
25 |
|
Jester |
1.3M |
8 |
95.04% |
43 |
Eye tracking
Dataset |
#Params |
Quantization |
Accuracy |
NPs |
Download |
|---|---|---|---|---|---|
Eye tracking CVPR 2024 |
219K |
8 |
p10: 98.58% mean_distance: 2.17 |
22 |
PTQ accuracy boosted with 5 epochs QAT.
Akida Pico models
Pico models are recurrent TENNs targeting the Akida Pico Neuromorphic Processor IP. Please refer to the Recurrent TENNs API for model descriptions and to the Akida Pico layers hardware constraints for mapping limits.
Audio domain
Keyword spotting
Architecture |
Dataset |
#Params |
Quantization |
Accuracy |
Download |
|---|---|---|---|---|---|
TENN recurrent |
Google Speech Commands |
46.6K |
8 |
93.80% |
Vibration domain
Fault classification
Architecture |
Dataset |
#Params |
Quantization |
Avg AUROC |
Download |
|---|---|---|---|---|---|
TENN recurrent |
UORED-VAFCLS |
16.6K |
8 |
0.9420 |