Published evidence · keyword spotting

Recognition quality and model size, measured together.

The public BC-ResNet-8 reference achieves higher command accuracy. The reqmote compact model has a much smaller file. Both are evaluated on the same audio.

Maintained by reqmote / Qurosphere · Updated

The measured tradeoff

Offline development/calibration comparison · evidence reviewed 28 September 2026
Integer modelCommand-clip accuracyFalse-command windowsModel file
BC-ResNet-8 public reference96.9%5.0%821.6 KiB
reqmote compact model89.2%6.9%45.7 KiB

The compact artifact is about 18× smaller, with lower recognition quality on this evaluation. BC-ResNet-8 is the reference selected for the next stage of target compatibility and footprint-reduction investigation. These results do not establish compiled memory use, board latency or power.

Open results and provenance (JSON) Explore recorded examples →

Data and decision rule

The command set is on, off, stop, go, up and down, with other outputs mapped to unknown or silence. Command accuracy is measured on 1,236 isolated clips from 143 calibration speakers. The non-command evaluation contains 4,561 speech clips plus background sounds and 80 ESC-10 environmental recordings, totalling about 1.39 hours.

Non-command audio is sampled as 19,973 overlapping one-second windows with a 250 ms stride. The headline decision selects the highest score among the eight classes, without a rejection threshold or temporal smoothing. Each model’s two percentages use that same rule.

A window error rate is not false activations per hour. Overlapping windows are correlated, and continuous-listening behavior also depends on event detection, rejection and temporal policy. Do not convert 5.0% into a device activation rate.

A separate stricter continuous-listening policy measured 53.1% command-event recall with one false event in 1.39 hours. That is a different operating setting, not an alternative interpretation of the table above. Homepage replay recordings are illustrative development examples scored separately from the aggregate comparison.

Model identities

BC-ResNet-8 · 841,280 bytes · SHA-256
8ba265270c986c13893417d613e9fa788376dae348bce6f45ab4dfe765b3bbda
reqmote compact model · 46,800 bytes · SHA-256
91ccb54659a56fe3d55b2171fb9c61e801aa01bd889731d44bd42d47243bb069

reqmote developed the compact model and evaluated both artifacts. The public checkpoint is from Andes ModelZoo, based on Qualcomm AI Research’s BC-ResNet; reqmote did not train that checkpoint.

Scope, sources and next evidence

This is development/calibration evidence, not an independent final acceptance test. The experiment does not establish microphone performance, deployment fit, physical-board timing or power. The A10 vision board walkthrough is a separate experiment and does not validate these keyword models.

Data sources: Google Speech Commands v2 (CC BY 4.0) and ESC-10 (CC BY 3.0). Exact attribution, artifact identity and audio provenance are in the linked JSON.

Maintained by reqmote / Qurosphere. Cite this page with the model name, evaluation setting and evidence review date. Read how model size relates to flash and RAM before interpreting the file-size comparison as target fit.